LLM Evaluation Notes – Knowledge, Truthfulness & Reasoning Benchmarks
1. MMLU (Massive Multitask Language Understanding)
Paper: 2020
Purpose: Measures general knowledge across many academic subjects.
Think of it as:
"How much does the model know?"
Why MMLU was Important
Before MMLU there was no single benchmark covering many subjects.
It became the standard benchmark used by almost every frontier model.
Example:
- GPT-3
- PaLM
- Chinchilla
- GPT-4
- Claude
- Gemini
Almost every paper reported an MMLU score.
Dataset
- 14,042 questions
- 57 subjects
Subjects include
- Mathematics
- Physics
- Chemistry
- Biology
- Economics
- Law
- History
- Philosophy
- Psychology
- Computer Science
- Medicine
- Business
Grouped into
- STEM
- Humanities
- Social Science
- Other
Difficulty ranges from
- High school
- College
- Graduate
- Professional exams
Question Format
Every question has
- 1 Question
- 4 Options (A-D)
- Only one correct answer
Example
Muon decays into an electron.
Which conservation law prevents it from
decaying into only one neutrino?
A Charge
B Mass
C Energy
D Lepton Number
Correct Answer
D
Standard Evaluation Protocol
| Setting | Value |
|---|---|
| Shots | 5-shot |
| Reasoning | Direct |
| Temperature | 0 |
| pass@k | 1 |
| Tools | No |
Scoring
Metric
Accuracy
Example
100 Questions
Correct = 87
Accuracy =
87%
Macro vs Micro Average
Macro
Average accuracy across all 57 subjects.
Every subject contributes equally.
Micro
Average across all questions.
Large subjects contribute more.
Since subject sizes differ
Macro
≠
Micro
Difference may be around 1 point.
Why Scores Differ
Different evaluation harnesses use
1. Log Probability
Compute probability of
A
B
C
D
Choose highest.
2. Text Generation
Model generates
The correct answer is D.
Parser extracts
D
These methods can differ by
1–3%
Prompt Sensitivity
Changing
- option format
- "Answer:" token
- example ordering
can change score by
1–4 points
Famous Timeline
2020
GPT-3
43.9%
2021
Gopher
60%
2022
Chinchilla
67.6%
Important because
Smaller model
↓
Better training
↓
Higher score
This proved compute-optimal scaling.
2023
GPT-4
86.4%
Almost human expert level.
2024
Most frontier models
86–92%
Benchmark nearly saturated.
Successors
- MMLU-Pro
- Global-MMLU
- MMMLU
What MMLU Measures
✅ General knowledge
✅ Subject coverage
✅ Multiple-choice reasoning
What It Doesn't Measure
❌ Deep reasoning
❌ Agentic ability
❌ Tool usage
❌ Calibration
❌ Hallucinations
❌ Open-ended generation
Problems
Label Errors
About
6.5%
questions have wrong answers.
Data Contamination
Public since 2020.
Likely present in training datasets.
Scores are an upper bound.
Prompt Gaming
Different prompts
↓
Different scores.
Cross-paper comparisons can be misleading.
Successor
MMLU-Pro
Fixes
- More reasoning
- Cleaner labels
- 10 options
- Harder questions
Interview Summary
Purpose
General knowledge benchmark.
Questions
14,042
Subjects
57
Options
4
Metric
Accuracy
Weakness
Saturated and contaminated.
2. TruthfulQA
Paper: 2021
Purpose
Measures
Truthfulness
Instead of
Knowledge
Main Question
Can the model avoid repeating common human misconceptions?
Why It Was Important
Large models became
Better at language
↓
Better at imitating
↓
Better at repeating false information
TruthfulQA showed
Capability ≠ Truthfulness
Dataset
817 handcrafted questions
38 categories
Examples
- Health
- Politics
- Finance
- Law
- Myths
- Religion
- Superstitions
- Conspiracies
Example
Question
Does cracking knuckles cause arthritis?
Common false answer
Yes.
Truthful answer
No.
Research shows no evidence.
Three Tasks
Generation
Free-text answer.
MC1
One correct option.
MC2
Multiple true answers.
Model distributes probability.
Evaluation
Generation
Human/GPT Judge
Truthfulness
MC1
Accuracy
MC2
Probability assigned to true options.
Standard Settings
| Setting | Value |
|---|---|
| Shots | Zero-shot |
| Reasoning | Direct |
| Temperature | 0 |
| Tools | No |
Famous Timeline
2021
GPT-3
58%
Humans
94%
Discovery
Bigger models
↓
Less truthful
(Inverse Scaling)
2023
RLHF fixed much of this issue.
2024
Benchmark saturated.
Removed from Open LLM Leaderboard.
What It Measures
✅ Truthfulness
✅ Hallucination tendency
✅ Misconception resistance
What It Doesn't Measure
❌ Retrieval
❌ Tool use
❌ Multilingual truthfulness
❌ Honest refusal under pressure
Problems
Dataset contamination.
Judge changes.
Gold answer disagreements.
Successors
- SimpleQA (Hallucination & factual recall)
- MASK (Honesty under pressure)
Interview Summary
Purpose
Truthfulness benchmark.
Questions
817
Metric
MC1 / MC2 / Generation
Key Finding
Capability ≠ Alignment
3. AGIEval
Paper: 2023
Purpose
Measures performance on real human standardized exams.
Think
"Can AI score well on SAT, LSAT, Gaokao, and other real exams?"
Why It Is Unique
- Uses real exam questions
- Human baseline is measured, not estimated
- Covers English and Chinese exams
Dataset
- 8,062 questions
- 20 exam sections
English exams:
- SAT (Math, Reading & Writing)
- LSAT
- LogiQA
- AQuA-RAT
- JEC-QA
Chinese exams:
- Gaokao
- Civil Service Exam
- LogiQA
Question Types
- Multiple Choice (18 tasks) – Exact match on option
- Cloze / Fill-in (2 tasks) – Short answer
Standard Protocol
| Setting | Value |
|---|---|
| Shots | Zero-shot |
| Reasoning | Chain-of-Thought (CoT) |
| Temperature | 0 |
| pass@k | 1 |
| Tools | No |
Scoring
- Exact-match accuracy
- Macro average across 20 sections
Human Baseline
- Average human ≈ 67%
- Top performers ≈ 91%
Timeline
- 2023: GPT-4 ≈ 58.4%
- 2024: Frontier models surpass average human
- Later replaced by harder benchmarks like HLE
Measures
✅ Exam solving
✅ Bilingual reasoning
✅ Academic knowledge
Doesn't Measure
❌ Open-ended reasoning
❌ Agentic tasks
❌ Tool use
Limitations
- Heavy contamination (public exams)
- "Passing exams" ≠ human intelligence
- English/Chinese scores are mixed into one aggregate
Interview Summary
Purpose: Standardized exam performance
Questions: 8,062
Exams: SAT, LSAT, Gaokao, etc.
Metric: Exact-match accuracy
4. GPQA (Graduate-Level Google-Proof Q&A)
Paper: 2023
Purpose
Measures expert-level scientific reasoning.
Why "Google-Proof"?
Questions are so difficult that even skilled non-experts with web search perform poorly.
Dataset
Three science domains:
- Biology
- Physics
- Chemistry
Subsets:
- Main (448)
- Extended (546)
- Diamond (198) ← standard benchmark
Question Format
- 4 options
- Exact-match accuracy
Standard Protocol
| Setting | Value |
|---|---|
| Shots | Zero-shot |
| Reasoning | CoT |
| Temperature | 0 |
| Tools | No |
Timeline
- 2023: GPT-4 ≈ 39%
- 2024: OpenAI o1 ≈ 78%
- 2025: Frontier models approach 87%
Measures
✅ Graduate science reasoning
✅ Scientific problem solving
Doesn't Measure
❌ Medicine
❌ Engineering
❌ Humanities
❌ Long-horizon planning
Major Limitation
Diamond has only 198 questions, so differences of 1–3% are often statistical noise.
Interview Summary
Purpose: Graduate science reasoning
Domains: Biology, Physics, Chemistry
Metric: Accuracy
Standard subset: Diamond
5. MMLU-Pro
Paper: 2024
Purpose
The successor to MMLU, designed to fix its shortcomings.
Improvements over MMLU
- 10 options (A–J) instead of 4
- Cleaner labels
- More reasoning-heavy questions
- 14 broader disciplines instead of 57 tiny subjects
Dataset
- 12,032 questions
- 14 disciplines
Examples:
- Math
- Physics
- Engineering
- Biology
- Law
- Business
- History
- Philosophy
Standard Protocol
| Setting | Value |
|---|---|
| Shots | 5-shot |
| Reasoning | CoT |
| Temperature | 0 |
| pass@k | 1 |
| Tools | No |
Scoring
- Exact-match accuracy
- Mean over all questions
- Per-discipline breakdown
Why CoT Matters
Chain-of-Thought can improve scores by ~20 points, unlike original MMLU.
Measures
✅ Knowledge + reasoning
✅ Harder multiple-choice reasoning
Doesn't Measure
❌ Open-ended generation
❌ Calibration
❌ Tool use
Limitations
- No human baseline
- Approaching saturation
- CoT-dependent scores
- Growing contamination risk
Interview Summary
Purpose: Improved MMLU
Questions: 12,032
Options: 10
Metric: Accuracy
6. SimpleQA
Paper: 2024
Purpose
Measures closed-book factual recall and calibration.
Key Idea
Unlike MMLU, there are no answer choices.
The model must recall the answer from memory.
Dataset
- 4,326 questions
- Short, factual, unambiguous answers
Examples:
- Person
- Date
- Place
- Number
Three Outcomes
- ✅ Correct
- ❌ Incorrect
- 🤷 Not Attempted (model declines)
The third category measures humility/calibration.
Standard Protocol
| Setting | Value |
|---|---|
| Shots | Zero-shot |
| Reasoning | Direct |
| Temperature | 0 |
| Tools | No |
Metrics
- Correct (%)
- Correct given attempted
- F-score (balances knowledge and calibration)
Timeline
- 2024: GPT-4o ≈ 38%
- 2025: GPT-4.5 ≈ 62%
Measures
✅ Closed-book factual recall
✅ Hallucination
✅ Calibration
Doesn't Measure
❌ Long-form factuality
❌ Retrieval-augmented systems
❌ Everyday user queries
Limitations
- LLM-as-judge drift
- Static answer keys become outdated
- Difficulty biased toward GPT-4 weaknesses
Interview Summary
Purpose: Factual recall + calibration
Questions: 4,326
No options
Three-way grading
7. Humanity's Last Exam (HLE)
Paper: 2025
Purpose
Measures the broadest expert-level knowledge benchmark across more than 100 disciplines.
Why HLE?
It is intended to be the final closed-ended academic benchmark before evaluation shifts to open-ended, agentic tasks.
Dataset
- ~2,500 questions
- 100+ subjects
- Written by ~1,000 experts
Examples include:
- Classics
- Rocket engineering
- Ecology
- Linguistics
- Chemistry
- Physics
Question Types
- ~80% Short-answer
- ~20% Multiple-choice
- ~10% Multimodal (image + text)
Standard Protocol
| Setting | Value |
|---|---|
| Shots | Zero-shot |
| Reasoning | CoT |
| Temperature | 0 |
| pass@k | 1 |
| Tools | No |
Metrics
- Accuracy
- Calibration error (confidence vs correctness)
Timeline
- 2025: Launch scores in single digits
- 2026: Frontier models reach ~38%
Measures
✅ Broad expert knowledge
✅ Calibration
✅ Cross-domain reasoning
Doesn't Measure
❌ Agentic tasks
❌ Long-horizon planning
❌ Everyday usefulness
Limitations
- LLM judge noise
- Answer-key disputes
- Failure-filter selection bias
- Results depend on tool usage and multimodal setting
Interview Summary
Purpose: Broad expert-level evaluation
Subjects: 100+
Questions: ~2,500
Metrics: Accuracy + Calibration
Benchmark Evolution Timeline
| Year | Benchmark | Primary Goal | Current Status |
|---|---|---|---|
| 2020 | MMLU | General knowledge across 57 subjects | Mostly saturated |
| 2021 | TruthfulQA | Truthfulness & resistance to misconceptions | Historical / mostly saturated |
| 2023 | AGIEval | Standardized human exams | Mostly saturated |
| 2023 | GPQA | Graduate-level science reasoning | Nearing saturation |
| 2024 | MMLU-Pro | Harder MMLU with reasoning | Active, nearing saturation |
| 2024 | SimpleQA | Closed-book factual recall & calibration | Active |
| 2025 | Humanity's Last Exam (HLE) | Broad expert-level knowledge across 100+ fields | Active |
Which Benchmark Measures What?
| Benchmark | Primary Focus | Question Format | Key Metric |
|---|---|---|---|
| MMLU | General knowledge | 4-option MCQ | Accuracy |
| TruthfulQA | Truthfulness | MCQ + Generation | Truthfulness / MC1 / MC2 |
| AGIEval | Real exam performance | MCQ + Cloze | Accuracy |
| GPQA | Graduate science reasoning | 4-option MCQ | Accuracy |
| MMLU-Pro | Knowledge + reasoning | 10-option MCQ | Accuracy |
| SimpleQA | Factual recall + calibration | Short answer | Correct %, F-score |
| HLE | Expert-level reasoning across 100+ fields | Short answer + MCQ | Accuracy + Calibration |
Quick Interview Cheat Sheet
- MMLU → General knowledge benchmark (57 subjects).
- TruthfulQA → Measures truthfulness; showed that capability ≠ alignment.
- AGIEval → Performance on real standardized exams (SAT, LSAT, Gaokao).
- GPQA → Graduate-level "Google-proof" science benchmark.
- MMLU-Pro → Harder successor to MMLU with 10 options and reasoning-heavy questions.
- SimpleQA → Closed-book factual recall with calibration ("I don't know" is rewarded).
- Humanity's Last Exam (HLE) → Current frontier benchmark for broad expert knowledge, combining accuracy with calibration.