References
Ackley, David H., Geoffrey E. Hinton, and Terrence J. Sejnowski. 1985.
“A Learning Algorithm for Boltzmann Machines.”
Cognitive Science 9 (1): 147–69. https://doi.org/10.1207/s15516709cog0901_7.
Alemi, Alexander A., Ian Fischer, Joshua V. Dillon, and Kevin Murphy.
2017. “Deep Variational Information Bottleneck.” In
ICLR 2017.
Amari, Shun-ichi, and Kenjiro Maginu. 1988. “Statistical
Neurodynamics of Associative Memory.” Neural Networks 1
(1): 63–73. https://doi.org/10.1016/0893-6080(88)90022-6.
Amit, Daniel J. 1989. Modeling Brain Function: The World of
Attractor Neural Networks. Cambridge: Cambridge University Press.
Amit, Daniel J., Hanoch Gutfreund, and Haim Sompolinsky. 1985.
“Storing Infinite Numbers of Patterns in a Spin-Glass Model of
Neural Networks.” Physical Review Letters 55 (14):
1530–33. https://doi.org/10.1103/PhysRevLett.55.1530.
———. 1987. “Statistical Mechanics of Neural Networks Near
Saturation.” Annals of Physics 173 (1): 30–67. https://doi.org/10.1016/0003-4916(87)90092-3.
Anderson, Brian D. O. 1982. “Reverse-Time Diffusion Equation
Models.” Stochastic Processes and Their Applications 12
(3): 313–26. https://doi.org/10.1016/0304-4149(82)90051-5.
Bahdanau, Dzmitry, Kyunghyun Cho, and Yoshua Bengio. 2015. “Neural
Machine Translation by Jointly Learning to Align and Translate.”
In International Conference on Learning Representations (ICLR).
Bahri, Yasaman, Jonathan Kadmon, Jeffrey Pennington, Sam S. Schoenholz,
Jascha Sohl-Dickstein, and Surya Ganguli. 2020. “Statistical
Mechanics of Deep Learning.” Annual Review of Condensed
Matter Physics 11: 501–28. https://doi.org/10.1146/annurev-conmatphys-031119-050745.
Baik, Jinho, Gérard Ben Arous, and Sandrine Péché. 2005. “Phase
Transition of the Largest Eigenvalue for Nonnull Complex Sample
Covariance Matrices.” Annals of Probability 33 (5):
1643–97. https://doi.org/10.1214/009117905000000233.
Baum, Leonard E., Ted Petrie, George Soules, and Norman Weiss. 1970.
“A Maximization Technique Occurring in the Statistical Analysis of
Probabilistic Functions of Markov Chains.” The
Annals of Mathematical Statistics 41 (1): 164–71. https://doi.org/10.1214/aoms/1177697196.
Bennett, Charles H. 1976. “Efficient Estimation of Free Energy
Differences from Monte Carlo Data.”
Journal of Computational Physics 22 (2): 245–68. https://doi.org/10.1016/0021-9991(76)90078-4.
Bethe, Hans A. 1935. “Statistical Theory of Superlattices.”
Proceedings of the Royal Society of London A 150 (871): 552–75.
https://doi.org/10.1098/rspa.1935.0122.
Blei, David M., Alp Kucukelbir, and Jon D. McAuliffe. 2017.
“Variational Inference: A Review for Statisticians.”
Journal of the American Statistical Association 112 (518):
859–77. https://doi.org/10.1080/01621459.2017.1285773.
Carreira-Perpiñán, Miguel Á., and Geoffrey E. Hinton. 2005. “On
Contrastive Divergence Learning.” In Proceedings of the 10th
International Workshop on Artificial Intelligence and Statistics
(AISTATS), R5:33–40. Proceedings of Machine Learning Research.
Černý, Vladimír. 1985. “Thermodynamical Approach to the Travelling
Salesman Problem: An Efficient Simulation Algorithm.” Journal
of Optimization Theory and Applications 45 (1): 41–51. https://doi.org/10.1007/BF00940812.
Chaudhari, Pratik, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo
Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo
Zecchina. 2017. “Entropy-SGD: Biasing Gradient
Descent into Wide Valleys.” In International Conference on
Learning Representations (ICLR).
Chaudhari, Pratik, and Stefano Soatto. 2018. “Stochastic Gradient
Descent Performs Variational Inference, Converges to Limit Cycles for
Deep Networks.” In International Conference on Learning
Representations (ICLR).
Chen, Ricky T. Q., Yulia Rubanova, Jesse Bettencourt, and David
Duvenaud. 2018. “Neural Ordinary Differential Equations.”
In NeurIPS, 31:6572–83.
Choromanska, Anna, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and
Yann LeCun. 2015. “The Loss Surfaces of Multilayer
Networks.” In AISTATS 2015, PMLR, 38:192–204.
Collin, Delphine, Felix Ritort, Christopher Jarzynski, Steven B. Smith,
Jr. Tinoco Ignacio, and Carlos Bustamante. 2005. “Verification of
the Crooks Fluctuation Theorem and Recovery of RNA Folding Free
Energies.” Nature 437 (7056): 231–34. https://doi.org/10.1038/nature04061.
Crooks, Gavin E. 1999. “Entropy Production Fluctuation Theorem and
the Nonequilibrium Work Relation for Free Energy Differences.”
Physical Review E 60 (3): 2721–26. https://doi.org/10.1103/PhysRevE.60.2721.
Dayan, Peter, Geoffrey E. Hinton, Radford M. Neal, and Richard S. Zemel.
1995. “The Helmholtz Machine.” Neural
Computation 7 (5): 889–904. https://doi.org/10.1162/neco.1995.7.5.889.
Decelle, Aurélien, Florent Krzakala, Cristopher Moore, and Lenka
Zdeborová. 2011. “Asymptotic Analysis of the Stochastic Block
Model for Modular Networks and Its Algorithmic Applications.”
Physical Review E 84 (6). https://doi.org/10.1103/PhysRevE.84.066106.
Demircigil, Mete, Judith Heusel, Matthias Löwe, Sven Upgang, and Franck
Vermet. 2017. “On a Model of Associative Memory with Huge Storage
Capacity.” Journal of Statistical Physics 168 (2):
288–99. https://doi.org/10.1007/s10955-017-1806-y.
Dempster, Arthur P., Nan M. Laird, and Donald B. Rubin. 1977.
“Maximum Likelihood from Incomplete Data via the EM
Algorithm.” Journal of the Royal Statistical Society, Series
B 39 (1): 1–38. https://doi.org/10.1111/j.2517-6161.1977.tb01600.x.
Dinh, Laurent, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. 2017.
“Sharp Minima Can Generalize for Deep Nets.” In
Proceedings of the 34th International Conference on Machine Learning
(ICML), 70:1019–28. PMLR.
Donoho, David L., Arian Maleki, and Andrea Montanari. 2009.
“Message-Passing Algorithms for Compressed Sensing.”
Proceedings of the National Academy of Sciences 106 (45):
18914–19. https://doi.org/10.1073/pnas.0909892106.
Edwards, Samuel F., and Philip W. Anderson. 1975. “Theory of Spin
Glasses.” Journal of Physics F: Metal Physics 5 (5):
965–74. https://doi.org/10.1088/0305-4608/5/5/017.
Foret, Pierre, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur.
2021. “Sharpness-Aware Minimization for Efficiently Improving
Generalization.” In International Conference on Learning
Representations (ICLR).
Friston, Karl J. 2010. “The Free-Energy Principle: A Unified Brain
Theory?” Nature Reviews Neuroscience 11 (2): 127–38. https://doi.org/10.1038/nrn2787.
Gallager, Robert G. 1962. “Low-Density Parity-Check Codes.”
IRE Transactions on Information Theory 8 (1): 21–28. https://doi.org/10.1109/TIT.1962.1057683.
Gardner, Elizabeth. 1988. “The Space of Interactions in Neural
Network Models.” Journal of Physics A: Mathematical and
General 21 (1): 257–70. https://doi.org/10.1088/0305-4470/21/1/030.
Geman, Stuart, and Donald Geman. 1984. “Stochastic Relaxation,
Gibbs Distributions, and the Bayesian
Restoration of Images.” IEEE Transactions on Pattern Analysis
and Machine Intelligence 6 (6): 721–41. https://doi.org/10.1109/TPAMI.1984.4767596.
Georges, Antoine, and Jonathan S. Yedidia. 1991. “How to Expand
Around Mean-Field Theory Using High-Temperature Expansions.”
Journal of Physics A: Mathematical and General 24 (9): 2173–92.
https://doi.org/10.1088/0305-4470/24/9/024.
Gibbs, J. Willard. 1902. Elementary Principles in Statistical
Mechanics. New Haven: Yale University Press.
Glauber, Roy J. 1963. “Time-Dependent Statistics of the
Ising Model.” Journal of Mathematical
Physics 4 (2): 294–307. https://doi.org/10.1063/1.1703954.
Goldenfeld, Nigel. 1992. Lectures on Phase Transitions and the
Renormalization Group. Frontiers in Physics. Reading, MA:
Addison-Wesley.
Grathwohl, Will, Ricky T. Q. Chen, Jesse Bettencourt, Ilya Sutskever,
and David Duvenaud. 2019. “FFJORD: Free-Form Continuous Dynamics
for Scalable Reversible Generative Models.” In ICLR.
Hajek, Bruce. 1988. “Cooling Schedules for Optimal
Annealing.” Mathematics of Operations Research 13 (2):
311–29. https://doi.org/10.1287/moor.13.2.311.
Hastings, W. Keith. 1970. “Monte Carlo Sampling Methods Using
Markov Chains and Their Applications.” Biometrika 57
(1): 97–109. https://doi.org/10.1093/biomet/57.1.97.
Hebb, Donald O. 1949. The Organization of Behavior: A
Neuropsychological Theory. New York: Wiley.
Hertz, John, Anders Krogh, and Richard G. Palmer. 1991. Introduction
to the Theory of Neural Computation. Redwood City: Addison-Wesley.
Higgins, Irina, Loic Matthey, Arka Pal, Christopher Burgess, Xavier
Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. 2017.
“β-VAE:
Learning Basic Visual Concepts with a Constrained Variational
Framework.” In International Conference on Learning
Representations (ICLR).
Hinton, Geoffrey E. 2002. “Training Products of Experts by
Minimizing Contrastive Divergence.” Neural Computation
14 (8): 1771–1800. https://doi.org/10.1162/089976602760128018.
———. 2012. “A Practical Guide to Training Restricted Boltzmann
Machines.” In Neural Networks: Tricks of the Trade,
edited by Grégoire Montavon, Genevieve B. Orr, and Klaus-Robert Müller,
2nd ed., 7700:599–619. Lecture Notes in Computer Science. Springer. https://doi.org/10.1007/978-3-642-35289-8_32.
Hinton, Geoffrey E., and Drew van Camp. 1993. “Keeping the Neural
Networks Simple by Minimizing the Description Length of the
Weights.” In Proceedings of the Sixth Annual Conference on
Computational Learning Theory (COLT), 5–13. https://doi.org/10.1145/168304.168306.
Hinton, Geoffrey E., Simon Osindero, and Yee-Whye Teh. 2006. “A
Fast Learning Algorithm for Deep Belief Nets.” Neural
Computation 18 (7): 1527–54. https://doi.org/10.1162/neco.2006.18.7.1527.
Hinton, Geoffrey E., and Ruslan R. Salakhutdinov. 2006. “Reducing
the Dimensionality of Data with Neural Networks.”
Science 313 (5786): 504–7. https://doi.org/10.1126/science.1127647.
Ho, Jonathan, Ajay Jain, and Pieter Abbeel. 2020. “Denoising
Diffusion Probabilistic Models.” In Advances in Neural
Information Processing Systems. Vol. 33. https://arxiv.org/abs/2006.11239.
Hochreiter, Sepp, and Jürgen Schmidhuber. 1997. “Flat
Minima.” Neural Computation 9 (1): 1–42. https://doi.org/10.1162/neco.1997.9.1.1.
Hopfield, John J. 1982. “Neural Networks and Physical Systems with
Emergent Collective Computational Abilities.” Proceedings of
the National Academy of Sciences 79 (8): 2554–58. https://doi.org/10.1073/pnas.79.8.2554.
———. 1984. “Neurons with Graded Response Have Collective
Computational Properties Like Those of Two-State Neurons.”
Proceedings of the National Academy of Sciences 81 (10):
3088–92. https://doi.org/10.1073/pnas.81.10.3088.
Hubbard, John. 1959. “Calculation of Partition Functions.”
Physical Review Letters 3 (2): 77–78. https://doi.org/10.1103/PhysRevLett.3.77.
Hyvärinen, Aapo. 2005. “Estimation of Non-Normalized Statistical
Models by Score Matching.” Journal of Machine Learning
Research 6: 695–709.
Jacot, Arthur, Franck Gabriel, and Clément Hongler. 2018. “Neural
Tangent Kernel: Convergence and Generalization in Neural
Networks.” In NeurIPS. Vol. 31.
Jarzynski, Christopher. 1997. “Nonequilibrium Equality for Free
Energy Differences.” Physical Review Letters 78 (14):
2690–93. https://doi.org/10.1103/PhysRevLett.78.2690.
———. 2011. “Equalities and Inequalities: Irreversibility and the
Second Law of Thermodynamics at the Nanoscale.” Annual Review
of Condensed Matter Physics 2: 329–51. https://doi.org/10.1146/annurev-conmatphys-062910-140506.
Jaynes, Edwin T. 1957. “Information Theory and Statistical
Mechanics.” Physical Review 106 (4): 620–30. https://doi.org/10.1103/PhysRev.106.620.
Jordan, Michael I., Zoubin Ghahramani, Tommi S. Jaakkola, and Lawrence
K. Saul. 1999. “An Introduction to Variational Methods for
Graphical Models.” Machine Learning 37 (2): 183–233. https://doi.org/10.1023/A:1007665907178.
Kappen, Hilbert J., and Francisco B. Rodríguez. 1998. “Efficient
Learning in Boltzmann Machines Using Linear Response Theory.”
Neural Computation 10 (5): 1137–56. https://doi.org/10.1162/089976698300017386.
Karras, Tero, Miika Aittala, Timo Aila, and Samuli Laine. 2022.
“Elucidating the Design Space of Diffusion-Based Generative
Models.” In NeurIPS. Vol. 35.
Keskar, Nitish Shirish, Dheevatsa Mudigere, Jorge Nocedal, Mikhail
Smelyanskiy, and Ping Tak Peter Tang. 2017. “On Large-Batch
Training for Deep Learning: Generalization Gap and Sharp Minima.”
In International Conference on Learning Representations (ICLR).
Kesten, Harry, and Bernt P. Stigum. 1966. “Additional Limit
Theorems for Indecomposable Multidimensional
Galton-Watson Processes.” Annals of
Mathematical Statistics 37 (6): 1463–81. https://doi.org/10.1214/aoms/1177699139.
Kingma, Diederik P., and Max Welling. 2014. “Auto-Encoding
Variational Bayes.” In International Conference
on Learning Representations (ICLR).
———. 2019. “An Introduction to Variational Autoencoders.”
Foundations and Trends in Machine Learning 12 (4): 307–92. https://doi.org/10.1561/2200000056.
Kirkpatrick, Scott, C. Daniel Gelatt, and Mario P. Vecchi. 1983.
“Optimization by Simulated Annealing.” Science 220
(4598): 671–80. https://doi.org/10.1126/science.220.4598.671.
Krotov, Dmitry, and John J. Hopfield. 2016. “Dense Associative
Memory for Pattern Recognition.” In Advances in Neural
Information Processing Systems (NeurIPS). Vol. 29.
———. 2021. “Large Associative Memory Problem in Neurobiology and
Machine Learning.” In International Conference on Learning
Representations (ICLR).
Krzakala, Florent, Cristopher Moore, Elchanan Mossel, Joe Neeman, Allan
Sly, Lenka Zdeborová, and Pan Zhang. 2013. “Spectral Redemption in
Clustering Sparse Networks.” Proceedings of the National
Academy of Sciences 110 (52): 20935–40. https://doi.org/10.1073/pnas.1312486110.
Landau, Lev D. 1937. “On the Theory of Phase Transitions.”
Zhurnal Eksperimentalnoi i Teoreticheskoi Fiziki 7: 19–32.
Langevin, Paul. 1908. “Sur La Théorie Du Mouvement
Brownien.” Comptes Rendus de l’Académie Des Sciences
146: 530–33.
Liphardt, Jan, Sophie Dumont, Steven B. Smith, Jr. Tinoco Ignacio, and
Carlos Bustamante. 2002. “Equilibrium Information from
Nonequilibrium Measurements in an Experimental Test of Jarzynski’s
Equality.” Science 296 (5574): 1832–35. https://doi.org/10.1126/science.1071152.
Lipman, Yaron, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and
Matt Le. 2023. “Flow Matching for Generative Modeling.” In
ICLR.
Little, William A. 1974. “The Existence of Persistent States in
the Brain.” Mathematical Biosciences 19 (1–2): 101–20.
https://doi.org/10.1016/0025-5564(74)90031-5.
MacKay, David J. C. 2003. Information Theory, Inference, and
Learning Algorithms. Cambridge: Cambridge University Press.
Mandt, Stephan, Matthew D. Hoffman, and David M. Blei. 2017.
“Stochastic Gradient Descent as Approximate Bayesian
Inference.” Journal of Machine Learning Research 18
(134): 1–35.
McCulloch, Warren S., and Walter Pitts. 1943. “A Logical Calculus
of the Ideas Immanent in Nervous Activity.” Bulletin of
Mathematical Biophysics 5: 115–33. https://doi.org/10.1007/BF02478259.
McEliece, Robert J., Edward C. Posner, Eugene R. Rodemich, and Santosh
S. Venkatesh. 1987. “The Capacity of the Hopfield Associative
Memory.” IEEE Transactions on Information Theory 33 (4):
461–82. https://doi.org/10.1109/TIT.1987.1057328.
Mehta, Pankaj, Marin Bukov, Ching-Hao Wang, Alexandre G. R. Day, Clint
Richardson, Charles K. Fisher, and David J. Schwab. 2019. “A
High-Bias, Low-Variance Introduction to Machine Learning for
Physicists.” Physics Reports 810: 1–124. https://doi.org/10.1016/j.physrep.2019.03.001.
Metropolis, Nicholas, Arianna W. Rosenbluth, Marshall N. Rosenbluth,
Augusta H. Teller, and Edward Teller. 1953. “Equation of State
Calculations by Fast Computing Machines.” Journal of Chemical
Physics 21 (6): 1087–92. https://doi.org/10.1063/1.1699114.
Mézard, Marc, and Andrea Montanari. 2009. Information, Physics, and
Computation. Oxford: Oxford University Press.
Mézard, Marc, Giorgio Parisi, and Miguel A. Virasoro. 1987. Spin
Glass Theory and Beyond: An Introduction to the Replica Method and Its
Applications. Singapore: World Scientific.
Neal, Radford M. 2001. “Annealed Importance Sampling.”
Statistics and Computing 11 (2): 125–39. https://doi.org/10.1023/A:1008923215028.
Nichol, Alexander Quinn, and Prafulla Dhariwal. 2021. “Improved
Denoising Diffusion Probabilistic Models.” In Proceedings of
the 38th International Conference on Machine Learning (ICML), PMLR
139, 8162–71.
Nishimori, Hidetoshi. 2001. Statistical Physics of Spin Glasses and
Information Processing: An Introduction. Oxford: Oxford University
Press.
Onsager, Lars. 1936. “Electric Moments of Molecules in
Liquids.” Journal of the American Chemical Society 58
(8): 1486–93. https://doi.org/10.1021/ja01299a050.
———. 1944. “Crystal Statistics. I. A Two-Dimensional
Model with an Order-Disorder Transition.” Physical
Review 65 (3–4): 117–49. https://doi.org/10.1103/PhysRev.65.117.
Opper, Manfred, and David Saad, eds. 2001. Advanced Mean Field
Methods: Theory and Practice. Cambridge, MA: MIT Press.
Parisi, Giorgio. 1979. “Infinite Number of Order Parameters for
Spin-Glasses.” Physical Review Letters 43 (23): 1754–56.
https://doi.org/10.1103/PhysRevLett.43.1754.
Pearl, Judea. 1988. Probabilistic Reasoning in Intelligent Systems:
Networks of Plausible Inference. San Francisco: Morgan Kaufmann.
Plefka, Timm. 1982. “Convergence Condition of the TAP
Equation for the Infinite-Ranged Ising Spin Glass
Model.” Journal of Physics A: Mathematical and General
15 (6): 1971–78. https://doi.org/10.1088/0305-4470/15/6/035.
Ramsauer, Hubert, Bernhard Schäfl, Johannes Lehner, Philipp Seidl,
Michael Widrich, Thomas Adler, Lukas Gruber, et al. 2021.
“Hopfield Networks Is All You Need.” In International
Conference on Learning Representations (ICLR).
Rezende, Danilo Jimenez, Shakir Mohamed, and Daan Wierstra. 2014.
“Stochastic Backpropagation and Approximate Inference in Deep
Generative Models.” In Proceedings of the 31st International
Conference on Machine Learning (ICML), 32:1278–86. PMLR.
Risken, Hannes. 1989. The Fokker–Planck Equation:
Methods of Solution and Applications. 2nd ed. Vol. 18. Springer
Series in Synergetics. Berlin: Springer.
Rose, Kenneth. 1998. “Deterministic Annealing for Clustering,
Compression, Classification, Regression, and Related Optimization
Problems.” Proceedings of the IEEE 86 (11): 2210–39. https://doi.org/10.1109/5.726788.
Saade, Alaa, Florent Krzakala, and Lenka Zdeborová. 2014.
“Spectral Clustering of Graphs with the Bethe Hessian.” In
NeurIPS, 27:406–14.
Salakhutdinov, Ruslan, and Iain Murray. 2008. “On the Quantitative
Analysis of Deep Belief Networks.” In Proceedings of the 25th
International Conference on Machine Learning (ICML), 872–79. ACM.
https://doi.org/10.1145/1390156.1390266.
Saxe, Andrew M., Yamini Bansal, Joel Dapello, Madhu Advani, Artemy
Kolchinsky, Brendan D. Tracey, and David D. Cox. 2018. “On the
Information Bottleneck Theory of Deep Learning.” In
International Conference on Learning Representations (ICLR). https://doi.org/10.1088/1742-5468/ab3985.
Schwarz, Ulrich S. 2023. “Theoretical Statistical Physics.”
Lecture-notes script, Heidelberg University.
Sethna, James P. 2021. Statistical Mechanics: Entropy, Order
Parameters, and Complexity. 2nd ed. Oxford: Oxford University
Press.
Shannon, Claude E. 1948. “A Mathematical Theory of
Communication.” Bell System Technical Journal 27:
379–423, 623–56. https://doi.org/10.1002/j.1538-7305.1948.tb01338.x.
Sherrington, David, and Scott Kirkpatrick. 1975. “Solvable Model
of a Spin-Glass.” Physical Review Letters 35 (26):
1792–96. https://doi.org/10.1103/PhysRevLett.35.1792.
Shwartz-Ziv, Ravid, and Naftali Tishby. 2017. “Opening the Black
Box of Deep Neural Networks via Information.”
Smolensky, Paul. 1986. “Information Processing in Dynamical
Systems: Foundations of Harmony Theory.” In Parallel
Distributed Processing: Explorations in the Microstructure of Cognition,
Vol. 1, edited by David E. Rumelhart and James L. McClelland,
194–281. Cambridge, MA: MIT Press.
Sohl-Dickstein, Jascha, Eric Weiss, Niru Maheswaranathan, and Surya
Ganguli. 2015. “Deep Unsupervised Learning Using Nonequilibrium
Thermodynamics.” In Proceedings of the 32nd International
Conference on Machine Learning, 37:2256–65. Proceedings of Machine
Learning Research. https://arxiv.org/abs/1503.03585.
Song, Yang, Conor Durkan, Iain Murray, and Stefano Ermon. 2021.
“Maximum Likelihood Training of Score-Based Diffusion
Models.” In NeurIPS. Vol. 34.
Song, Yang, and Stefano Ermon. 2019. “Generative Modeling by
Estimating Gradients of the Data Distribution.” In Advances
in Neural Information Processing Systems (NeurIPS). Vol. 32.
Song, Yang, Sahaj Garg, Jiaxin Shi, and Stefano Ermon. 2020.
“Sliced Score Matching: A Scalable Approach to Density and Score
Estimation.” In Proceedings of the 35th Uncertainty in
Artificial Intelligence Conference (UAI), 115:574–84. PMLR.
Song, Yang, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar,
Stefano Ermon, and Ben Poole. 2021. “Score-Based Generative
Modeling Through Stochastic Differential Equations.” In
International Conference on Learning Representations. https://arxiv.org/abs/2011.13456.
Stratonovich, Ruslan L. 1957. “On a Method of Calculating Quantum
Distribution Functions.” Soviet Physics Doklady 2:
416–19.
Sutskever, Ilya, and Tijmen Tieleman. 2010. “On the Convergence
Properties of Contrastive Divergence.” In Proceedings of the
13th International Conference on Artificial Intelligence and Statistics
(AISTATS), 9:789–95. Proceedings of Machine Learning Research.
Thouless, David J., Philip W. Anderson, and Richard G. Palmer. 1977.
“Solution of ’Solvable Model of a Spin Glass’.”
Philosophical Magazine 35 (3): 593–601. https://doi.org/10.1080/14786437708235992.
Tieleman, Tijmen. 2008. “Training Restricted Boltzmann Machines
Using Approximations to the Likelihood Gradient.” In
Proceedings of the 25th International Conference on Machine Learning
(ICML), 1064–71. https://doi.org/10.1145/1390156.1390290.
Tipping, Michael E., and Christopher M. Bishop. 1999.
“Probabilistic Principal Component Analysis.” Journal
of the Royal Statistical Society, Series B 61 (3): 611–22. https://doi.org/10.1111/1467-9868.00196.
Tishby, Naftali, Fernando C. Pereira, and William Bialek. 1999.
“The Information Bottleneck Method.” In Proc. 37th
Allerton Conference on Communication, Control, and Computing,
368–77.
Tishby, Naftali, and Noga Zaslavsky. 2015. “Deep Learning and the
Information Bottleneck Principle.” In 2015 IEEE Information
Theory Workshop (ITW), 1–5. https://doi.org/10.1109/ITW.2015.7133169.
Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion
Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017.
“Attention Is All You Need.” In Advances in Neural
Information Processing Systems (NeurIPS), 30:6000–6010.
Vincent, Pascal. 2011. “A Connection Between Score Matching and
Denoising Autoencoders.” Neural Computation 23 (7):
1661–74. https://doi.org/10.1162/neco_a_00142.
Wainwright, Martin J., and Michael I. Jordan. 2008. “Graphical
Models, Exponential Families, and Variational Inference.”
Foundations and Trends in Machine Learning 1 (1–2): 1–305. https://doi.org/10.1561/2200000001.
Weiss, Pierre. 1907. “L’hypothèse Du Champ
Moléculaire Et La Propriété
Ferromagnétique.” Journal de Physique
Théorique Et Appliquée 6 (1): 661–90. https://doi.org/10.1051/jphystap:019070060066100.
Welling, Max, Sirui Lu, and Lars Holdijk. 2026. Generative AI and
Stochastic Thermodynamics: A Tale of Free Energies. Cambridge:
Cambridge University Press. https://doi.org/10.1017/9781009709071.
Welling, Max, and Yee Whye Teh. 2011. “Bayesian Learning via
Stochastic Gradient Langevin Dynamics.” In
Proceedings of the 28th International Conference on Machine Learning
(ICML), 681–88.
Yedidia, Jonathan S., William T. Freeman, and Yair Weiss. 2003.
“Understanding Belief Propagation and Its Generalizations.”
In Exploring Artificial Intelligence in the New Millennium,
edited by Gerhard Lakemeyer and Bernhard Nebel, 239–69. San Francisco:
Morgan Kaufmann.
Yuille, Alan L., and Anand Rangarajan. 2003. “The Concave-Convex
Procedure.” Neural Computation 15 (4): 915–36. https://doi.org/10.1162/08997660360581958.
Zdeborová, Lenka, and Florent Krzakala. 2016. “Statistical Physics
of Inference: Thresholds and Algorithms.” Advances in
Physics 65 (5): 453–552. https://doi.org/10.1080/00018732.2016.1211393.
Zwanzig, Robert W. 1954. “High-Temperature Equation of State by a
Perturbation Method. I. Nonpolar Gases.” The Journal of
Chemical Physics 22 (8): 1420–26. https://doi.org/10.1063/1.1740409.