Evals are the new unit tests: build the harness before you trust the model
Evaluation sets, graders, regression and what "better" means for an AI workflow
How do you know an AI system is good — on your work, in production, and for the people using it?
The discipline that retires "trust me": build the eval set before you trust the model, read benchmarks as claims to reproduce on your own tasks, monitor what production actually does, and measure the human–AI team rather than the tool.
Evaluation sets, graders, regression and what "better" means for an AI workflow
Saturation, contamination, the jagged frontier and the task-horizon curve
Input and output drift, per-group outcomes, heartbeats and the retrain trigger
AI as tool, teammate or gate; who gains, who deskills, and how to pre-register the metric