Applied AI Academy

Benchmarks lie gently, so measure on your own work

Saturation, contamination, the jagged frontier and the task-horizon curve

The question: When a vendor says its model "leads the benchmark", what exactly has been shown — and what has it got to do with your tasks?

Public benchmarks drive buying decisions, and almost none of them measure your work. Saturation, contamination and jaggedness make headline scores weaker evidence every year; the 2025 research gives you the tools to read them.

What the lesson covers

A benchmark is a fixed test set with a score. Three things make headline scores weak evidence. **Saturation**: once models near the ceiling, differences stop being meaningful and the test stops discriminating. **Contamination**: test items leak into training data, so the model has seen the exam. **Construct drift**: the benchmark measures something adjacent to what you care about — multiple-choice recall is not your contract review.

The **jagged frontier** (Lesson 4) is why a leaderboard rank does not transfer: a model can be superhuman on one task and poor on a neighbouring one that looks the same to you. Apple's 2025 study of reasoning models found performance collapsing beyond a complexity threshold on puzzles the models "reasoned" about — strong on the benchmark regime, fragile just past it. The only way to know where the jag falls for your work is to measure your work.

Some 2025 measurement does transfer, and is worth knowing. METR's task-horizon study found the length of software tasks models can complete at 50% reliability doubling roughly every seven months — a trend line, not a guarantee, and one to watch with a decision trigger (Lesson 30). GDPval measures real deliverables across occupations rather than exam items, which is closer to work — and still not your work.

The discipline: treat every benchmark claim as a hypothesis to **reproduce on twenty of your own tasks**. Write the gap memo — where the public score and your score agree, where they diverge, and what that implies for the decision in front of you. Vendors rarely object to being measured; the ones who do are telling you something.

Honesty runs both ways. Your own twenty-task test is small and noisy; say so, size it up if the decision is large, and keep it as the seed of the eval harness from Lesson 39. The goal is not to distrust all measurement but to trust measurement in proportion to how much it looks like the work.

Key points

Framework — Claim → Reproduce on your tasks → Gap memo → Decide

Every benchmark claim is a hypothesis. Twenty of your own tasks, the gap between public and private scores, and what the gap means for this decision.

The lab

Turn three benchmark claims into evidence about your own work.

Deliverable: The gap memo (one page) with the 20-task table and a purchase or no-purchase recommendation.

Open this lesson, its lab and its quiz

Sources and further reading