Watch
How to Automate AI Evals (Correctly)
Argues for reading and labelling real traces before writing any rubric, so the eval set is built from the failures that actually happen instead of cases the author picked.
Engineering Fundamentals for the Agent Era
Contents Section 9, Directing Agents
A prompt change shipped because it looked better on three examples.
A model-graded score that disagrees with human judgment and was never checked against it.
An eval set so easy that every version scores near perfect, so it cannot detect a regression.
Measuring nondeterministic systems with datasets, graders and statistics, so a prompt or model change is judged on evidence.
An agent reports that its new summarization prompt scores 4.6 out of 5 from a model judge on ten cases it chose itself, against 4.2 for the old prompt. Find every reason this does not show the new prompt is better, and design an eval that could.
Watch
Argues for reading and labelling real traces before writing any rubric, so the eval set is built from the failures that actually happen instead of cases the author picked.
A university lecture covering inter-rater agreement metrics, rule-based metrics, LLM-as-a-judge with its position, verbosity and self-enhancement biases, and agent evaluation.
A step-by-step walkthrough of building an eval from real traces: error analysis and open and axial coding, when a code check beats an LLM judge, and the most common mistakes.
Two chapters on evaluation cover exact-match and similarity metrics, AI-as-a-judge and its failure modes, and how to design an evaluation pipeline for an open-ended model feature.
The standard text on variance, statistical power and false positives when comparing two versions, which is exactly the discipline a prompt A/B comparison needs.
Paper
Shows how to put confidence intervals on eval scores, use paired comparisons, and resample each question to account for run-to-run variance before deciding one prompt beats another.
Paper
Measures how often LLM graders agree with human labels, and finds that the grading criteria themselves shift as people review outputs, so a judge has to be checked against labels before anyone trusts it.