References

Ackley, David H., Geoffrey E. Hinton, and Terrence J. Sejnowski. 1985. “A Learning Algorithm for Boltzmann Machines.” Cognitive Science 9 (1): 147–69. https://doi.org/10.1207/s15516709cog0901_7.
Alemi, Alexander A., Ian Fischer, Joshua V. Dillon, and Kevin Murphy. 2017. “Deep Variational Information Bottleneck.” In ICLR 2017.
Amari, Shun-ichi, and Kenjiro Maginu. 1988. “Statistical Neurodynamics of Associative Memory.” Neural Networks 1 (1): 63–73. https://doi.org/10.1016/0893-6080(88)90022-6.
Amit, Daniel J. 1989. Modeling Brain Function: The World of Attractor Neural Networks. Cambridge: Cambridge University Press.
Amit, Daniel J., Hanoch Gutfreund, and Haim Sompolinsky. 1985. “Storing Infinite Numbers of Patterns in a Spin-Glass Model of Neural Networks.” Physical Review Letters 55 (14): 1530–33. https://doi.org/10.1103/PhysRevLett.55.1530.
———. 1987. “Statistical Mechanics of Neural Networks Near Saturation.” Annals of Physics 173 (1): 30–67. https://doi.org/10.1016/0003-4916(87)90092-3.
Anderson, Brian D. O. 1982. “Reverse-Time Diffusion Equation Models.” Stochastic Processes and Their Applications 12 (3): 313–26. https://doi.org/10.1016/0304-4149(82)90051-5.
Bahdanau, Dzmitry, Kyunghyun Cho, and Yoshua Bengio. 2015. “Neural Machine Translation by Jointly Learning to Align and Translate.” In International Conference on Learning Representations (ICLR).
Bahri, Yasaman, Jonathan Kadmon, Jeffrey Pennington, Sam S. Schoenholz, Jascha Sohl-Dickstein, and Surya Ganguli. 2020. “Statistical Mechanics of Deep Learning.” Annual Review of Condensed Matter Physics 11: 501–28. https://doi.org/10.1146/annurev-conmatphys-031119-050745.
Baik, Jinho, Gérard Ben Arous, and Sandrine Péché. 2005. “Phase Transition of the Largest Eigenvalue for Nonnull Complex Sample Covariance Matrices.” Annals of Probability 33 (5): 1643–97. https://doi.org/10.1214/009117905000000233.
Baum, Leonard E., Ted Petrie, George Soules, and Norman Weiss. 1970. “A Maximization Technique Occurring in the Statistical Analysis of Probabilistic Functions of Markov Chains.” The Annals of Mathematical Statistics 41 (1): 164–71. https://doi.org/10.1214/aoms/1177697196.
Bennett, Charles H. 1976. “Efficient Estimation of Free Energy Differences from Monte Carlo Data.” Journal of Computational Physics 22 (2): 245–68. https://doi.org/10.1016/0021-9991(76)90078-4.
Bethe, Hans A. 1935. “Statistical Theory of Superlattices.” Proceedings of the Royal Society of London A 150 (871): 552–75. https://doi.org/10.1098/rspa.1935.0122.
Blei, David M., Alp Kucukelbir, and Jon D. McAuliffe. 2017. “Variational Inference: A Review for Statisticians.” Journal of the American Statistical Association 112 (518): 859–77. https://doi.org/10.1080/01621459.2017.1285773.
Carreira-Perpiñán, Miguel Á., and Geoffrey E. Hinton. 2005. “On Contrastive Divergence Learning.” In Proceedings of the 10th International Workshop on Artificial Intelligence and Statistics (AISTATS), R5:33–40. Proceedings of Machine Learning Research.
Černý, Vladimír. 1985. “Thermodynamical Approach to the Travelling Salesman Problem: An Efficient Simulation Algorithm.” Journal of Optimization Theory and Applications 45 (1): 41–51. https://doi.org/10.1007/BF00940812.
Chaudhari, Pratik, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina. 2017. “Entropy-SGD: Biasing Gradient Descent into Wide Valleys.” In International Conference on Learning Representations (ICLR).
Chaudhari, Pratik, and Stefano Soatto. 2018. “Stochastic Gradient Descent Performs Variational Inference, Converges to Limit Cycles for Deep Networks.” In International Conference on Learning Representations (ICLR).
Chen, Ricky T. Q., Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. 2018. “Neural Ordinary Differential Equations.” In NeurIPS, 31:6572–83.
Choromanska, Anna, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun. 2015. “The Loss Surfaces of Multilayer Networks.” In AISTATS 2015, PMLR, 38:192–204.
Collin, Delphine, Felix Ritort, Christopher Jarzynski, Steven B. Smith, Jr. Tinoco Ignacio, and Carlos Bustamante. 2005. “Verification of the Crooks Fluctuation Theorem and Recovery of RNA Folding Free Energies.” Nature 437 (7056): 231–34. https://doi.org/10.1038/nature04061.
Crooks, Gavin E. 1999. “Entropy Production Fluctuation Theorem and the Nonequilibrium Work Relation for Free Energy Differences.” Physical Review E 60 (3): 2721–26. https://doi.org/10.1103/PhysRevE.60.2721.
Dayan, Peter, Geoffrey E. Hinton, Radford M. Neal, and Richard S. Zemel. 1995. “The Helmholtz Machine.” Neural Computation 7 (5): 889–904. https://doi.org/10.1162/neco.1995.7.5.889.
Decelle, Aurélien, Florent Krzakala, Cristopher Moore, and Lenka Zdeborová. 2011. “Asymptotic Analysis of the Stochastic Block Model for Modular Networks and Its Algorithmic Applications.” Physical Review E 84 (6). https://doi.org/10.1103/PhysRevE.84.066106.
Demircigil, Mete, Judith Heusel, Matthias Löwe, Sven Upgang, and Franck Vermet. 2017. “On a Model of Associative Memory with Huge Storage Capacity.” Journal of Statistical Physics 168 (2): 288–99. https://doi.org/10.1007/s10955-017-1806-y.
Dempster, Arthur P., Nan M. Laird, and Donald B. Rubin. 1977. “Maximum Likelihood from Incomplete Data via the EM Algorithm.” Journal of the Royal Statistical Society, Series B 39 (1): 1–38. https://doi.org/10.1111/j.2517-6161.1977.tb01600.x.
Dinh, Laurent, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. 2017. “Sharp Minima Can Generalize for Deep Nets.” In Proceedings of the 34th International Conference on Machine Learning (ICML), 70:1019–28. PMLR.
Donoho, David L., Arian Maleki, and Andrea Montanari. 2009. “Message-Passing Algorithms for Compressed Sensing.” Proceedings of the National Academy of Sciences 106 (45): 18914–19. https://doi.org/10.1073/pnas.0909892106.
Edwards, Samuel F., and Philip W. Anderson. 1975. “Theory of Spin Glasses.” Journal of Physics F: Metal Physics 5 (5): 965–74. https://doi.org/10.1088/0305-4608/5/5/017.
Foret, Pierre, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. 2021. “Sharpness-Aware Minimization for Efficiently Improving Generalization.” In International Conference on Learning Representations (ICLR).
Friston, Karl J. 2010. “The Free-Energy Principle: A Unified Brain Theory?” Nature Reviews Neuroscience 11 (2): 127–38. https://doi.org/10.1038/nrn2787.
Gallager, Robert G. 1962. “Low-Density Parity-Check Codes.” IRE Transactions on Information Theory 8 (1): 21–28. https://doi.org/10.1109/TIT.1962.1057683.
Gardner, Elizabeth. 1988. “The Space of Interactions in Neural Network Models.” Journal of Physics A: Mathematical and General 21 (1): 257–70. https://doi.org/10.1088/0305-4470/21/1/030.
Geman, Stuart, and Donald Geman. 1984. “Stochastic Relaxation, Gibbs Distributions, and the Bayesian Restoration of Images.” IEEE Transactions on Pattern Analysis and Machine Intelligence 6 (6): 721–41. https://doi.org/10.1109/TPAMI.1984.4767596.
Georges, Antoine, and Jonathan S. Yedidia. 1991. “How to Expand Around Mean-Field Theory Using High-Temperature Expansions.” Journal of Physics A: Mathematical and General 24 (9): 2173–92. https://doi.org/10.1088/0305-4470/24/9/024.
Gibbs, J. Willard. 1902. Elementary Principles in Statistical Mechanics. New Haven: Yale University Press.
Glauber, Roy J. 1963. “Time-Dependent Statistics of the Ising Model.” Journal of Mathematical Physics 4 (2): 294–307. https://doi.org/10.1063/1.1703954.
Goldenfeld, Nigel. 1992. Lectures on Phase Transitions and the Renormalization Group. Frontiers in Physics. Reading, MA: Addison-Wesley.
Grathwohl, Will, Ricky T. Q. Chen, Jesse Bettencourt, Ilya Sutskever, and David Duvenaud. 2019. “FFJORD: Free-Form Continuous Dynamics for Scalable Reversible Generative Models.” In ICLR.
Hajek, Bruce. 1988. “Cooling Schedules for Optimal Annealing.” Mathematics of Operations Research 13 (2): 311–29. https://doi.org/10.1287/moor.13.2.311.
Hastings, W. Keith. 1970. “Monte Carlo Sampling Methods Using Markov Chains and Their Applications.” Biometrika 57 (1): 97–109. https://doi.org/10.1093/biomet/57.1.97.
Hebb, Donald O. 1949. The Organization of Behavior: A Neuropsychological Theory. New York: Wiley.
Hertz, John, Anders Krogh, and Richard G. Palmer. 1991. Introduction to the Theory of Neural Computation. Redwood City: Addison-Wesley.
Higgins, Irina, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. 2017. β-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework.” In International Conference on Learning Representations (ICLR).
Hinton, Geoffrey E. 2002. “Training Products of Experts by Minimizing Contrastive Divergence.” Neural Computation 14 (8): 1771–1800. https://doi.org/10.1162/089976602760128018.
———. 2012. “A Practical Guide to Training Restricted Boltzmann Machines.” In Neural Networks: Tricks of the Trade, edited by Grégoire Montavon, Genevieve B. Orr, and Klaus-Robert Müller, 2nd ed., 7700:599–619. Lecture Notes in Computer Science. Springer. https://doi.org/10.1007/978-3-642-35289-8_32.
Hinton, Geoffrey E., and Drew van Camp. 1993. “Keeping the Neural Networks Simple by Minimizing the Description Length of the Weights.” In Proceedings of the Sixth Annual Conference on Computational Learning Theory (COLT), 5–13. https://doi.org/10.1145/168304.168306.
Hinton, Geoffrey E., Simon Osindero, and Yee-Whye Teh. 2006. “A Fast Learning Algorithm for Deep Belief Nets.” Neural Computation 18 (7): 1527–54. https://doi.org/10.1162/neco.2006.18.7.1527.
Hinton, Geoffrey E., and Ruslan R. Salakhutdinov. 2006. “Reducing the Dimensionality of Data with Neural Networks.” Science 313 (5786): 504–7. https://doi.org/10.1126/science.1127647.
Ho, Jonathan, Ajay Jain, and Pieter Abbeel. 2020. “Denoising Diffusion Probabilistic Models.” In Advances in Neural Information Processing Systems. Vol. 33. https://arxiv.org/abs/2006.11239.
Hochreiter, Sepp, and Jürgen Schmidhuber. 1997. “Flat Minima.” Neural Computation 9 (1): 1–42. https://doi.org/10.1162/neco.1997.9.1.1.
Hopfield, John J. 1982. “Neural Networks and Physical Systems with Emergent Collective Computational Abilities.” Proceedings of the National Academy of Sciences 79 (8): 2554–58. https://doi.org/10.1073/pnas.79.8.2554.
———. 1984. “Neurons with Graded Response Have Collective Computational Properties Like Those of Two-State Neurons.” Proceedings of the National Academy of Sciences 81 (10): 3088–92. https://doi.org/10.1073/pnas.81.10.3088.
Hubbard, John. 1959. “Calculation of Partition Functions.” Physical Review Letters 3 (2): 77–78. https://doi.org/10.1103/PhysRevLett.3.77.
Hyvärinen, Aapo. 2005. “Estimation of Non-Normalized Statistical Models by Score Matching.” Journal of Machine Learning Research 6: 695–709.
Jacot, Arthur, Franck Gabriel, and Clément Hongler. 2018. “Neural Tangent Kernel: Convergence and Generalization in Neural Networks.” In NeurIPS. Vol. 31.
Jarzynski, Christopher. 1997. “Nonequilibrium Equality for Free Energy Differences.” Physical Review Letters 78 (14): 2690–93. https://doi.org/10.1103/PhysRevLett.78.2690.
———. 2011. “Equalities and Inequalities: Irreversibility and the Second Law of Thermodynamics at the Nanoscale.” Annual Review of Condensed Matter Physics 2: 329–51. https://doi.org/10.1146/annurev-conmatphys-062910-140506.
Jaynes, Edwin T. 1957. “Information Theory and Statistical Mechanics.” Physical Review 106 (4): 620–30. https://doi.org/10.1103/PhysRev.106.620.
Jordan, Michael I., Zoubin Ghahramani, Tommi S. Jaakkola, and Lawrence K. Saul. 1999. “An Introduction to Variational Methods for Graphical Models.” Machine Learning 37 (2): 183–233. https://doi.org/10.1023/A:1007665907178.
Kappen, Hilbert J., and Francisco B. Rodríguez. 1998. “Efficient Learning in Boltzmann Machines Using Linear Response Theory.” Neural Computation 10 (5): 1137–56. https://doi.org/10.1162/089976698300017386.
Karras, Tero, Miika Aittala, Timo Aila, and Samuli Laine. 2022. “Elucidating the Design Space of Diffusion-Based Generative Models.” In NeurIPS. Vol. 35.
Keskar, Nitish Shirish, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2017. “On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima.” In International Conference on Learning Representations (ICLR).
Kesten, Harry, and Bernt P. Stigum. 1966. “Additional Limit Theorems for Indecomposable Multidimensional Galton-Watson Processes.” Annals of Mathematical Statistics 37 (6): 1463–81. https://doi.org/10.1214/aoms/1177699139.
Kingma, Diederik P., and Max Welling. 2014. “Auto-Encoding Variational Bayes.” In International Conference on Learning Representations (ICLR).
———. 2019. “An Introduction to Variational Autoencoders.” Foundations and Trends in Machine Learning 12 (4): 307–92. https://doi.org/10.1561/2200000056.
Kirkpatrick, Scott, C. Daniel Gelatt, and Mario P. Vecchi. 1983. “Optimization by Simulated Annealing.” Science 220 (4598): 671–80. https://doi.org/10.1126/science.220.4598.671.
Krotov, Dmitry, and John J. Hopfield. 2016. “Dense Associative Memory for Pattern Recognition.” In Advances in Neural Information Processing Systems (NeurIPS). Vol. 29.
———. 2021. “Large Associative Memory Problem in Neurobiology and Machine Learning.” In International Conference on Learning Representations (ICLR).
Krzakala, Florent, Cristopher Moore, Elchanan Mossel, Joe Neeman, Allan Sly, Lenka Zdeborová, and Pan Zhang. 2013. “Spectral Redemption in Clustering Sparse Networks.” Proceedings of the National Academy of Sciences 110 (52): 20935–40. https://doi.org/10.1073/pnas.1312486110.
Landau, Lev D. 1937. “On the Theory of Phase Transitions.” Zhurnal Eksperimentalnoi i Teoreticheskoi Fiziki 7: 19–32.
Langevin, Paul. 1908. “Sur La Théorie Du Mouvement Brownien.” Comptes Rendus de l’Académie Des Sciences 146: 530–33.
Liphardt, Jan, Sophie Dumont, Steven B. Smith, Jr. Tinoco Ignacio, and Carlos Bustamante. 2002. “Equilibrium Information from Nonequilibrium Measurements in an Experimental Test of Jarzynski’s Equality.” Science 296 (5574): 1832–35. https://doi.org/10.1126/science.1071152.
Lipman, Yaron, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. 2023. “Flow Matching for Generative Modeling.” In ICLR.
Little, William A. 1974. “The Existence of Persistent States in the Brain.” Mathematical Biosciences 19 (1–2): 101–20. https://doi.org/10.1016/0025-5564(74)90031-5.
MacKay, David J. C. 2003. Information Theory, Inference, and Learning Algorithms. Cambridge: Cambridge University Press.
Mandt, Stephan, Matthew D. Hoffman, and David M. Blei. 2017. “Stochastic Gradient Descent as Approximate Bayesian Inference.” Journal of Machine Learning Research 18 (134): 1–35.
McCulloch, Warren S., and Walter Pitts. 1943. “A Logical Calculus of the Ideas Immanent in Nervous Activity.” Bulletin of Mathematical Biophysics 5: 115–33. https://doi.org/10.1007/BF02478259.
McEliece, Robert J., Edward C. Posner, Eugene R. Rodemich, and Santosh S. Venkatesh. 1987. “The Capacity of the Hopfield Associative Memory.” IEEE Transactions on Information Theory 33 (4): 461–82. https://doi.org/10.1109/TIT.1987.1057328.
Mehta, Pankaj, Marin Bukov, Ching-Hao Wang, Alexandre G. R. Day, Clint Richardson, Charles K. Fisher, and David J. Schwab. 2019. “A High-Bias, Low-Variance Introduction to Machine Learning for Physicists.” Physics Reports 810: 1–124. https://doi.org/10.1016/j.physrep.2019.03.001.
Metropolis, Nicholas, Arianna W. Rosenbluth, Marshall N. Rosenbluth, Augusta H. Teller, and Edward Teller. 1953. “Equation of State Calculations by Fast Computing Machines.” Journal of Chemical Physics 21 (6): 1087–92. https://doi.org/10.1063/1.1699114.
Mézard, Marc, and Andrea Montanari. 2009. Information, Physics, and Computation. Oxford: Oxford University Press.
Mézard, Marc, Giorgio Parisi, and Miguel A. Virasoro. 1987. Spin Glass Theory and Beyond: An Introduction to the Replica Method and Its Applications. Singapore: World Scientific.
Neal, Radford M. 2001. “Annealed Importance Sampling.” Statistics and Computing 11 (2): 125–39. https://doi.org/10.1023/A:1008923215028.
Nichol, Alexander Quinn, and Prafulla Dhariwal. 2021. “Improved Denoising Diffusion Probabilistic Models.” In Proceedings of the 38th International Conference on Machine Learning (ICML), PMLR 139, 8162–71.
Nishimori, Hidetoshi. 2001. Statistical Physics of Spin Glasses and Information Processing: An Introduction. Oxford: Oxford University Press.
Onsager, Lars. 1936. “Electric Moments of Molecules in Liquids.” Journal of the American Chemical Society 58 (8): 1486–93. https://doi.org/10.1021/ja01299a050.
———. 1944. “Crystal Statistics. I. A Two-Dimensional Model with an Order-Disorder Transition.” Physical Review 65 (3–4): 117–49. https://doi.org/10.1103/PhysRev.65.117.
Opper, Manfred, and David Saad, eds. 2001. Advanced Mean Field Methods: Theory and Practice. Cambridge, MA: MIT Press.
Parisi, Giorgio. 1979. “Infinite Number of Order Parameters for Spin-Glasses.” Physical Review Letters 43 (23): 1754–56. https://doi.org/10.1103/PhysRevLett.43.1754.
Pearl, Judea. 1988. Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference. San Francisco: Morgan Kaufmann.
Plefka, Timm. 1982. “Convergence Condition of the TAP Equation for the Infinite-Ranged Ising Spin Glass Model.” Journal of Physics A: Mathematical and General 15 (6): 1971–78. https://doi.org/10.1088/0305-4470/15/6/035.
Ramsauer, Hubert, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, et al. 2021. “Hopfield Networks Is All You Need.” In International Conference on Learning Representations (ICLR).
Rezende, Danilo Jimenez, Shakir Mohamed, and Daan Wierstra. 2014. “Stochastic Backpropagation and Approximate Inference in Deep Generative Models.” In Proceedings of the 31st International Conference on Machine Learning (ICML), 32:1278–86. PMLR.
Risken, Hannes. 1989. The Fokker–Planck Equation: Methods of Solution and Applications. 2nd ed. Vol. 18. Springer Series in Synergetics. Berlin: Springer.
Rose, Kenneth. 1998. “Deterministic Annealing for Clustering, Compression, Classification, Regression, and Related Optimization Problems.” Proceedings of the IEEE 86 (11): 2210–39. https://doi.org/10.1109/5.726788.
Saade, Alaa, Florent Krzakala, and Lenka Zdeborová. 2014. “Spectral Clustering of Graphs with the Bethe Hessian.” In NeurIPS, 27:406–14.
Salakhutdinov, Ruslan, and Iain Murray. 2008. “On the Quantitative Analysis of Deep Belief Networks.” In Proceedings of the 25th International Conference on Machine Learning (ICML), 872–79. ACM. https://doi.org/10.1145/1390156.1390266.
Saxe, Andrew M., Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan D. Tracey, and David D. Cox. 2018. “On the Information Bottleneck Theory of Deep Learning.” In International Conference on Learning Representations (ICLR). https://doi.org/10.1088/1742-5468/ab3985.
Schwarz, Ulrich S. 2023. “Theoretical Statistical Physics.” Lecture-notes script, Heidelberg University.
Sethna, James P. 2021. Statistical Mechanics: Entropy, Order Parameters, and Complexity. 2nd ed. Oxford: Oxford University Press.
Shannon, Claude E. 1948. “A Mathematical Theory of Communication.” Bell System Technical Journal 27: 379–423, 623–56. https://doi.org/10.1002/j.1538-7305.1948.tb01338.x.
Sherrington, David, and Scott Kirkpatrick. 1975. “Solvable Model of a Spin-Glass.” Physical Review Letters 35 (26): 1792–96. https://doi.org/10.1103/PhysRevLett.35.1792.
Shwartz-Ziv, Ravid, and Naftali Tishby. 2017. “Opening the Black Box of Deep Neural Networks via Information.”
Smolensky, Paul. 1986. “Information Processing in Dynamical Systems: Foundations of Harmony Theory.” In Parallel Distributed Processing: Explorations in the Microstructure of Cognition, Vol. 1, edited by David E. Rumelhart and James L. McClelland, 194–281. Cambridge, MA: MIT Press.
Sohl-Dickstein, Jascha, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015. “Deep Unsupervised Learning Using Nonequilibrium Thermodynamics.” In Proceedings of the 32nd International Conference on Machine Learning, 37:2256–65. Proceedings of Machine Learning Research. https://arxiv.org/abs/1503.03585.
Song, Yang, Conor Durkan, Iain Murray, and Stefano Ermon. 2021. “Maximum Likelihood Training of Score-Based Diffusion Models.” In NeurIPS. Vol. 34.
Song, Yang, and Stefano Ermon. 2019. “Generative Modeling by Estimating Gradients of the Data Distribution.” In Advances in Neural Information Processing Systems (NeurIPS). Vol. 32.
Song, Yang, Sahaj Garg, Jiaxin Shi, and Stefano Ermon. 2020. “Sliced Score Matching: A Scalable Approach to Density and Score Estimation.” In Proceedings of the 35th Uncertainty in Artificial Intelligence Conference (UAI), 115:574–84. PMLR.
Song, Yang, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2021. “Score-Based Generative Modeling Through Stochastic Differential Equations.” In International Conference on Learning Representations. https://arxiv.org/abs/2011.13456.
Stratonovich, Ruslan L. 1957. “On a Method of Calculating Quantum Distribution Functions.” Soviet Physics Doklady 2: 416–19.
Sutskever, Ilya, and Tijmen Tieleman. 2010. “On the Convergence Properties of Contrastive Divergence.” In Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS), 9:789–95. Proceedings of Machine Learning Research.
Thouless, David J., Philip W. Anderson, and Richard G. Palmer. 1977. “Solution of ’Solvable Model of a Spin Glass’.” Philosophical Magazine 35 (3): 593–601. https://doi.org/10.1080/14786437708235992.
Tieleman, Tijmen. 2008. “Training Restricted Boltzmann Machines Using Approximations to the Likelihood Gradient.” In Proceedings of the 25th International Conference on Machine Learning (ICML), 1064–71. https://doi.org/10.1145/1390156.1390290.
Tipping, Michael E., and Christopher M. Bishop. 1999. “Probabilistic Principal Component Analysis.” Journal of the Royal Statistical Society, Series B 61 (3): 611–22. https://doi.org/10.1111/1467-9868.00196.
Tishby, Naftali, Fernando C. Pereira, and William Bialek. 1999. “The Information Bottleneck Method.” In Proc. 37th Allerton Conference on Communication, Control, and Computing, 368–77.
Tishby, Naftali, and Noga Zaslavsky. 2015. “Deep Learning and the Information Bottleneck Principle.” In 2015 IEEE Information Theory Workshop (ITW), 1–5. https://doi.org/10.1109/ITW.2015.7133169.
Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. “Attention Is All You Need.” In Advances in Neural Information Processing Systems (NeurIPS), 30:6000–6010.
Vincent, Pascal. 2011. “A Connection Between Score Matching and Denoising Autoencoders.” Neural Computation 23 (7): 1661–74. https://doi.org/10.1162/neco_a_00142.
Wainwright, Martin J., and Michael I. Jordan. 2008. “Graphical Models, Exponential Families, and Variational Inference.” Foundations and Trends in Machine Learning 1 (1–2): 1–305. https://doi.org/10.1561/2200000001.
Weiss, Pierre. 1907. “L’hypothèse Du Champ Moléculaire Et La Propriété Ferromagnétique.” Journal de Physique Théorique Et Appliquée 6 (1): 661–90. https://doi.org/10.1051/jphystap:019070060066100.
Welling, Max, Sirui Lu, and Lars Holdijk. 2026. Generative AI and Stochastic Thermodynamics: A Tale of Free Energies. Cambridge: Cambridge University Press. https://doi.org/10.1017/9781009709071.
Welling, Max, and Yee Whye Teh. 2011. “Bayesian Learning via Stochastic Gradient Langevin Dynamics.” In Proceedings of the 28th International Conference on Machine Learning (ICML), 681–88.
Yedidia, Jonathan S., William T. Freeman, and Yair Weiss. 2003. “Understanding Belief Propagation and Its Generalizations.” In Exploring Artificial Intelligence in the New Millennium, edited by Gerhard Lakemeyer and Bernhard Nebel, 239–69. San Francisco: Morgan Kaufmann.
Yuille, Alan L., and Anand Rangarajan. 2003. “The Concave-Convex Procedure.” Neural Computation 15 (4): 915–36. https://doi.org/10.1162/08997660360581958.
Zdeborová, Lenka, and Florent Krzakala. 2016. “Statistical Physics of Inference: Thresholds and Algorithms.” Advances in Physics 65 (5): 453–552. https://doi.org/10.1080/00018732.2016.1211393.
Zwanzig, Robert W. 1954. “High-Temperature Equation of State by a Perturbation Method. I. Nonpolar Gases.” The Journal of Chemical Physics 22 (8): 1420–26. https://doi.org/10.1063/1.1740409.