Skip to content

Sources

Every factual claim in The Math Beneath traces back to one of these 14 sources.

  1. MacKay, D. J. C. (2003). Information Theory, Inference, and Learning Algorithms. Cambridge University Press. — pt. 1, 6
  2. Shannon, C. E. (1948). A Mathematical Theory of Communication. Bell System Technical Journal, 27(3), 379–423. — pt. 6, 7
  3. Kullback, S., & Leibler, R. A. (1951). On Information and Sufficiency. Annals of Mathematical Statistics, 22(1), 79–86. — pt. 7
  4. Cover, T. M., & Thomas, J. A. (2006). Elements of Information Theory (2nd ed.). Wiley-Interscience. — pt. 6, 7
  5. Jaynes, E. T. (2003). Probability Theory: The Logic of Science. Cambridge University Press. Ch. 4 (Elementary Hypothesis Testing — evidence in decibels). — pt. 2, 3, 4
  6. Berkson, J. (1944). Application of the Logistic Function to Bio-Assay. Journal of the American Statistical Association, 39(227), 357–365. (Coins the term “logit”.) — pt. 2, 3
  7. Bradley, R. A., & Terry, M. E. (1952). Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons. Biometrika, 39(3/4), 324–345. — pt. 3
  8. Bayes, T. (1763). An Essay towards solving a Problem in the Doctrine of Chances. Philosophical Transactions of the Royal Society of London, 53, 370–418. (Communicated by R. Price.) — pt. 4
  9. Gigerenzer, G., Gaissmaier, W., Kurz-Milcke, E., Schwartz, L. M., & Woloshin, S. (2007). Helping Doctors and Patients Make Sense of Health Statistics. Psychological Science in the Public Interest, 8(2), 53–96. (Physicians misreading screening-test positives.) — pt. 4
  10. Grinstead, C. M., & Snell, J. L. (1997). Introduction to Probability (2nd rev. ed.). American Mathematical Society. Ch. 6 (Expected Value), Ch. 8 (Law of Large Numbers). — pt. 5
  11. Robbins, H., & Monro, S. (1951). A Stochastic Approximation Method. Annals of Mathematical Statistics, 22(3), 400–407. — pt. 5, 8
  12. Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. MIT Press. Ch. 3 (Probability and Information Theory), Ch. 4 (Numerical Computation), Ch. 8 (Optimization). — pt. 1, 7, 8
  13. Lemaréchal, C. (2012). Cauchy and the Gradient Method. Documenta Mathematica, Extra Volume ISMP, 251–254. (On Cauchy’s 1847 note introducing gradient descent.) — pt. 8
  14. Ruder, S. (2016). An Overview of Gradient Descent Optimization Algorithms. arXiv:1609.04747. — pt. 8

Definition

Read the full glossary entry →