principles.fyi · the brain · concept
MMLU
A giant multiple-choice exam across 57 subjects used to score what a model knows.
score = % of 57-subject multiple-choice questions answered correctly
MMLU (Massive Multitask Language Understanding) is a benchmark of about 16,000 multiple-choice questions spanning 57 subjects — from high-school math to law, medicine, and history. To 'take' it, the model is shown a question with options A–D and is scored by which option it rates most likely; the headline number is simply the percentage answered correctly. It became a standard yardstick because one score sweeps across many domains, but it's still only a proxy: it rewards exam-style recall, can be inflated if the questions leaked into pretraining (data contamination), and says nothing about reasoning shown, honesty, efficiency, or real-world usefulness.
Appears in
- How we grade them LLMs in the Wild · pt 4