Evaluation That Survives Production

An aggregate bar above a set of per-slice bars, one of them markedly shorter and marked

The change was small. The evaluation suite went green, the reviewer approved it, the model shipped on a Tuesday. By Thursday the support queue had a new shape in it — not a flood, just a recurring complaint about one kind of request that used to work. The suite is still green. It was green through the whole incident, and it will be green tomorrow.

Nothing malfunctioned. The suite measured what it was built to measure, which turned out not to be the thing the product depended on. That gap is the normal state of affairs for evaluation, and closing it is mostly a design problem rather than a tooling one.

An evaluation set is a bet about the future

Every held-out set carries an implicit claim: the requests you will receive tomorrow resemble the examples in this file closely enough that performance on one predicts performance on the other. When that claim is true, offline numbers transfer. When it is false, they transfer inconsistently, which is worse than not transferring at all, because it takes several successful launches before anybody starts distrusting the number.

The claim is usually false in one specific direction. Evaluation sets are built early, from whatever data existed at the time, and then the product goes and changes who uses it. New customers arrive with different phrasing, different document formats, different edge cases. The set does not notice. It keeps reporting on a population that stopped being your population sometime last quarter.

So the first property of a durable suite is that somebody owns the question is this still representative, and answers it on a schedule rather than when something breaks.

Averages are a management summary, not an evaluation

A single headline number over a mixed set is almost designed to hide the failure that matters. A change that helps the common case slightly and destroys a small category will move the aggregate in the right direction. If that small category happens to be a customer segment, a language, a document type or a device class, the aggregate has actively misled you.

The fix is to fix the slices before you need them. Decide which partitions of your traffic are meaningful — by tenant size, by input length, by locale, by source system, by whichever axis your users would notice being different — and report every one of them on every run. A change that improves the mean while regressing a named slice should require somebody to say out loud that the trade is acceptable, rather than passing silently because nobody computed it.

Slices also make regressions legible. “Quality dropped” starts an argument; “long documents from the ingestion path regressed while everything else held” starts an investigation.

The set has to grow, and part of it must not

These two requirements sound contradictory and are not. Split the suite in two.

A stable core stays fixed for long stretches so that numbers are comparable across months. If you continuously alter the set, you lose the ability to say whether the system is better than it was in spring, because every comparison is confounded by a change in the ruler.

A living set grows every time production surprises you. Each genuine incident, each complaint that turned out to be real, each internally discovered failure gets converted into a case and added. This is the highest-value data any team has and the most routinely discarded: the ticket closes, the fix ships, and the example that exposed the weakness is never captured, so nothing prevents its return.

Adding cases makes the aggregate look worse over time, since you are deliberately accumulating hard examples. That is fine, and it needs to be explained once to whoever reads the dashboard, or somebody will eventually propose cleaning out the difficult cases to make the trend look better.

Where the labels come from decides what is measurable

If two competent people, given your instructions, would label a case differently, then no model can be scored reliably on that case and no improvement on it can be trusted. Disagreement among careful humans is a ceiling on measurement, not a nuisance to be averaged away.

This has a practical consequence: writing the labelling instructions is part of building the evaluation, not a preliminary chore. Where the instruction is ambiguous, either sharpen it or remove the case. Where a case is genuinely ambiguous in the world — the correct answer really does depend on unstated context — keep it in a separate bucket and stop treating movement there as signal.

Automated graders need to be evaluated too

Using a model to grade another model’s output is a reasonable engineering choice and an unreasonable act of faith. The grader has its own biases: toward longer answers, toward its own phrasing habits, toward outputs whose form matches its instructions even when the content is wrong.

Treat the grader as a component under test. Hold out a sample of items with human labels, measure how often the grader agrees, and re-measure whenever you change the grading prompt or the grading model. A grader that quietly changed underneath you will move every downstream metric at once, and the resulting graph looks exactly like a real quality shift.

Evaluate the system, not the checkpoint

In production, users interact with a pipeline: retrieval, pre-processing, whatever the model does, post-processing, validation, fallback behaviour when something times out. Model-only evaluation measures one stage of that and attributes the result to the whole.

The consequence is a familiar and demoralising pattern — a model change that tests well and delivers nothing, because the retrieval step never surfaced the information the improved model would have used, or because the parser downstream discards the better-formatted output. Run the end-to-end path at least as often as the model-only path, and when the two disagree, believe the end-to-end one.

Making it operational

The suite should run on every change that could plausibly affect behaviour, including changes nobody thinks of as model changes: prompt edits, dependency upgrades, retrieval index rebuilds, configuration flags. Cheap checks run on every commit; expensive ones run before release.

Results belong with the artefact. When you look at a model that has been serving for six weeks, you should be able to retrieve the exact evaluation output that approved it, on the exact set version that was current at the time. Without that, the archaeology after an incident is guesswork, and the suite’s most important function — telling you what you knew, and when — is lost.

An evaluation suite that never blocks a release is not a safety net. It is a formality with a build step.