Skip to content

Sources

Every factual claim in Transformers, ELI5 traces back to one of these 38 sources.

  1. Jurafsky, Daniel, and James H. Martin. Speech and Language Processing (3rd ed. draft, 2026), Chapter 8: “Transformers.” — pt. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10
  2. Vaswani, Ashish, et al. “Attention Is All You Need.” NeurIPS, 2017. arXiv:1706.03762. — pt. 2, 3, 4, 9
  3. Mikolov, Tomas, Kai Chen, Greg Corrado, and Jeffrey Dean. “Efficient Estimation of Word Representations in Vector Space.” 2013. arXiv:1301.3781. — pt. 2
  4. Mikolov, Tomas, Wen-tau Yih, and Geoffrey Zweig. “Linguistic Regularities in Continuous Space Word Representations.” NAACL-HLT, 2013. — pt. 2
  5. Sennrich, Rico, Barry Haddow, and Alexandra Birch. “Neural Machine Translation of Rare Words with Subword Units.” ACL, 2016. arXiv:1508.07909. — pt. 2
  6. Su, Jianlin, et al. “RoFormer: Enhanced Transformer with Rotary Position Embedding.” 2021. arXiv:2104.09864. — pt. 2
  7. Levesque, Hector J., Ernest Davis, and Leora Morgenstern. “The Winograd Schema Challenge.” KR, 2012. — pt. 3
  8. Hendrycks, Dan, and Kevin Gimpel. “Gaussian Error Linear Units (GELUs).” 2016. arXiv:1606.08415. — pt. 4
  9. Shazeer, Noam. “GLU Variants Improve Transformer.” 2020. arXiv:2002.05202. — pt. 4
  10. Touvron, Hugo, et al. “LLaMA: Open and Efficient Foundation Language Models.” 2023. arXiv:2302.13971. — pt. 4, 6
  11. Geva, Mor, Roei Schuster, Jonathan Berant, and Omer Levy. “Transformer Feed-Forward Layers Are Key-Value Memories.” EMNLP, 2021. arXiv:2012.14913. — pt. 4
  12. Elhage, Nelson, et al. “Toy Models of Superposition.” Transformer Circuits Thread, 2022. — pt. 4, 10
  13. He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. “Deep Residual Learning for Image Recognition.” CVPR, 2016. arXiv:1512.03385. — pt. 5
  14. Elhage, Nelson, et al. “A Mathematical Framework for Transformer Circuits.” Transformer Circuits Thread, 2021. — pt. 5
  15. Ba, Jimmy Lei, Jamie Ryan Kiros, and Geoffrey E. Hinton. “Layer Normalization.” 2016. arXiv:1607.06450. — pt. 5
  16. Xiong, Ruibin, et al. “On Layer Normalization in the Transformer Architecture.” ICML, 2020. arXiv:2002.04745. — pt. 5
  17. Zhang, Biao, and Rico Sennrich. “Root Mean Square Layer Normalization.” NeurIPS, 2019. arXiv:1910.07467. — pt. 5
  18. Brown, Tom B., et al. “Language Models are Few-Shot Learners.” NeurIPS, 2020. arXiv:2005.14165. — pts. 5, 9
  19. Radford, Alec, et al. “Language Models are Unsupervised Multitask Learners.” OpenAI Technical Report, 2019. — pt. 6
  20. Press, Ofir, and Lior Wolf. “Using the Output Embedding to Improve Language Models.” EACL, 2017. arXiv:1608.05859. — pt. 6
  21. Holtzman, Ari, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. “The Curious Case of Neural Text Degeneration.” ICLR, 2020. arXiv:1904.09751. — pt. 7
  22. Nguyen, Minh Nhat, et al. “Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs.” 2024. arXiv:2407.01082. — pt. 7
  23. Keskar, Nitish Shirish, et al. “CTRL: A Conditional Transformer Language Model for Controllable Generation.” 2019. arXiv:1909.05858. — pt. 7
  24. Rumelhart, David E., Geoffrey E. Hinton, and Ronald J. Williams. “Learning representations by back-propagating errors.” Nature 323 (1986): 533–536. — pt. 8
  25. Ouyang, Long, et al. “Training Language Models to Follow Instructions with Human Feedback.” NeurIPS, 2022. arXiv:2203.02155. — pt. 8
  26. Rafailov, Rafael, et al. “Direct Preference Optimization: Your Language Model is Secretly a Reward Model.” NeurIPS, 2023. arXiv:2305.18290. — pt. 8
  27. DeepSeek-AI. “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.” 2025. arXiv:2501.12948. — pt. 8
  28. OpenAI. “Learning to Reason with LLMs.” 2024. — pt. 8
  29. Kaplan, Jared, et al. “Scaling Laws for Neural Language Models.” 2020. arXiv:2001.08361. — pt. 9
  30. Hoffmann, Jordan, et al. “Training Compute-Optimal Large Language Models.” NeurIPS, 2022. arXiv:2203.15556. — pt. 9
  31. Ainslie, Joshua, et al. “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.” EMNLP, 2023. arXiv:2305.13245. — pt. 9
  32. Gemini Team, Google. “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.” 2024. arXiv:2403.05530. — pt. 9
  33. Hu, Edward J., et al. “LoRA: Low-Rank Adaptation of Large Language Models.” 2021. arXiv:2106.09685. — pt. 9
  34. Dettmers, Tim, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. “LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.” NeurIPS, 2022. arXiv:2208.07339. — pt. 9
  35. Fedus, William, Barret Zoph, and Noam Shazeer. “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.” 2021. arXiv:2101.03961. — pt. 9
  36. Wei, Jason, et al. “Emergent Abilities of Large Language Models.” Transactions on Machine Learning Research, 2022. arXiv:2206.07682. — pt. 9
  37. Schaeffer, Rylan, Brando Miranda, and Sanmi Koyejo. “Are Emergent Abilities of Large Language Models a Mirage?” NeurIPS, 2023. arXiv:2304.15004. — pt. 9
  38. Olsson, Catherine, et al. “In-context Learning and Induction Heads.” Transformer Circuits Thread, 2022. — pt. 10
  39. nostalgebraist. “interpreting GPT: the logit lens.” LessWrong, 2020. — pt. 10
  40. Bricken, Trenton, et al. “Towards Monosemanticity: Decomposing Language Models With Dictionary Learning.” Transformer Circuits Thread, 2023. — pt. 10
  41. Dehghani, Mostafa, et al. “Universal Transformers.” ICLR, 2019. arXiv:1807.03819. — recurrence article
  42. Geiping, Jonas, et al. “Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach.” 2025. arXiv:2502.05171. — recurrence article
  43. OpenAI. “What Parameter Golf taught us.” May 12, 2026. — recurrence article
  44. Pachocki, Jakub. “An Alien Mind.” OpenAI, September 6, 2026. — recurrence article

Definition

Read the full glossary entry →