Building an LLM application is only half the job. This course covers the other half: how AI engineers systematically measure whether a model or application is accurate, safe, fast, and cheap enough to put in front of real users — before something like Air Canada's chatbot happens to you.
Real incidents, failure modes, and the case for moving beyond vibe testing.
What LLM evaluation is, why it exists, and why 'vibe testing' isn't enough.
Air Canada, Chevrolet, and a lawyer sanctioned by a court — three case studies in what goes wrong without evaluation.
Deterministic vs. probabilistic systems, and why traditional software testing doesn't transfer.
The full evaluation landscape, pipelines, and revision-ready summaries.
Model vs. application evaluation, golden datasets, rubrics, and the full learning roadmap.
A single interview-ready reference distilling every concept covered so far.
A second, deeper pass on the model vs. application evaluation distinction.
The 12-step lifecycle from defining a task to monitoring it in production.
Multiple failure points and risk categories mean one app needs several eval pipelines.
Programmatic, human, and LLM-as-judge techniques, plus offline vs online testing.
How foundation models themselves are benchmarked, scored, and compared.
Why AI engineers evaluate models directly, not just the apps built on top of them.
The four-step evaluation flow and the eight core capabilities every model is measured on.
The four components of every benchmark: dataset, run config, scoring, and aggregation.
How a benchmark is actually executed end-to-end, and why harnesses automate it.
Frontier labs, third-party evaluators, and in-house teams — and how to read each one's results.
A complete anatomy of LLM benchmarks, plus common problems like contamination and saturation.
MMLU, TruthfulQA, AGIEval, GPQA, MMLU-Pro, SimpleQA, and Humanity's Last Exam, side by side.