Sources
Every factual claim in Agents and Images traces back to one of these 30 sources.
- Nakano, Reiichiro, et al. “WebGPT: Browser-assisted question-answering with human feedback.” arXiv:2112.09332 (2021). — pt. 1, 4, 5
- Yao, Shunyu, et al. “ReAct: Synergizing Reasoning and Acting in Language Models.” ICLR 2023; arXiv:2210.03629. — pt. 1, 2, 3
- Shridhar, Mohit, et al. “ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.” ICLR 2021; arXiv:1912.01734. — pt. 2
- Anthropic. “Model Context Protocol — specification.” modelcontextprotocol.io (accessed 2026). — pt. 3, 4
- Greshake, Kai, et al. “Not what you’ve signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.” arXiv:2302.12173 (2023). — pt. 7
- Liu, Nelson F., et al. “Lost in the Middle: How Language Models Use Long Contexts.” TACL 2024; arXiv:2307.03172. — pt. 5
- Shinn, Noah, et al. “Reflexion: Language Agents with Verbal Reinforcement Learning.” NeurIPS 2023; arXiv:2303.11366. — pt. 5
- Wei, Jason, et al. “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.” NeurIPS 2022; arXiv:2201.11903. — pt. 2
- Wang, Xuezhi, et al. “Self-Consistency Improves Chain of Thought Reasoning in Language Models.” ICLR 2023; arXiv:2203.11171. — pt. 6
- Yao, Shunyu, et al. “Tree of Thoughts: Deliberate Problem Solving with Large Language Models.” NeurIPS 2023; arXiv:2305.10601. — pt. 6
- Lightman, Hunter, et al. “Let’s Verify Step by Step.” arXiv:2305.20050 (2023). Introduces the PRM800K step-label dataset. — pt. 6
- Zelikman, Eric, et al. “STaR: Bootstrapping Reasoning With Reasoning.” NeurIPS 2022; arXiv:2203.14465. — pt. 6
- OpenAI. “Learning to Reason with LLMs” (o1 announcement), September 2024. — pt. 6
- DeepSeek-AI. “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.” arXiv:2501.12948 (2025). — pt. 6
- Chen, Mark, et al. “Evaluating Large Language Models Trained on Code.” arXiv:2107.03374 (2021). Introduces HumanEval and pass@k. — pt. 8
- Cobbe, Karl, et al. “Training Verifiers to Solve Math Word Problems.” arXiv:2110.14168 (2021). Introduces GSM8K. — pt. 8
- Jimenez, Carlos E., et al. “SWE-bench: Can Language Models Resolve Real-World GitHub Issues?” ICLR 2024; arXiv:2310.06770. — pt. 7, 8
- Aleithan, Reem, et al. “SWE-Bench+: Enhanced Coding Benchmark for LLMs.” arXiv:2410.06992 (2024). Audit finding a fraction of “passing” patches to be spurious. — pt. 8
- Zhou, Shuyan, et al. “WebArena: A Realistic Web Environment for Building Autonomous Agents.” ICLR 2024; arXiv:2307.13854. — pt. 7, 8
- Dosovitskiy, Alexey, et al. “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.” ICLR 2021; arXiv:2010.11929. — pt. 9
- Radford, Alec, et al. “Learning Transferable Visual Models From Natural Language Supervision.” ICML 2021; arXiv:2103.00020. — pt. 10
- Liu, Haotian, et al. “Visual Instruction Tuning.” NeurIPS 2023; arXiv:2304.08485. — pt. 10
- Li, Junnan, et al. “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.” ICML 2023; arXiv:2301.12597. Introduces the Q-Former. — pt. 10
- Li, Yifan, et al. “Evaluating Object Hallucination in Large Vision-Language Models.” EMNLP 2023; arXiv:2305.10355. Introduces POPE. — pt. 10
- Ho, Jonathan, Ajay Jain, and Pieter Abbeel. “Denoising Diffusion Probabilistic Models.” NeurIPS 2020; arXiv:2006.11239. — pt. 11
- Rombach, Robin, et al. “High-Resolution Image Synthesis with Latent Diffusion Models.” CVPR 2022; arXiv:2112.10752. — pt. 11
- Peebles, William, and Saining Xie. “Scalable Diffusion Models with Transformers.” ICCV 2023; arXiv:2212.09748. — pt. 11, 12
- Esser, Patrick, et al. “Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.” ICML 2024; arXiv:2403.03206. The Stable Diffusion 3 / MM-DiT paper. — pt. 12
- Heusel, Martin, et al. “GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium.” NeurIPS 2017; arXiv:1706.08500. Introduces FID. — pt. 11
- Ho, Jonathan, and Tim Salimans. “Classifier-Free Diffusion Guidance.” arXiv:2207.12598 (2022). — pt. 12