LLM Evaluation
17 modules Cheat sheet github.com/Raka7317
Course dashboard

Ship it, then prove it works.

Building an LLM application is only half the job. This course covers the other half: how AI engineers systematically measure whether a model or application is accurate, safe, fast, and cheap enough to put in front of real users — before something like Air Canada's chatbot happens to you.

17
Modules
4
Sections
7
Benchmarks decoded
3
Case studies

↓ One-page cheat sheet (all 17 modules)   Download as PDF

Vibe test
"Looks right to me"
Define
Task & success criteria
Build
Golden dataset + rubric
Score
Programmatic · Human · Judge
Ship & watch
Offline gate → online signal

Why Evaluation Matters

Real incidents, failure modes, and the case for moving beyond vibe testing.

Building an Evaluation Practice

The full evaluation landscape, pipelines, and revision-ready summaries.

Evaluation Methods

Programmatic, human, and LLM-as-judge techniques, plus offline vs online testing.

Model Evaluation & Benchmarks

How foundation models themselves are benchmarked, scored, and compared.