Why Offline Gains Vanish Online

Four descending stage columns with items falling away between them

The pattern is familiar enough to be a running joke in most teams. A candidate model beats the incumbent offline by a comfortable margin. It goes into a controlled online test. The result is flat, or slightly negative, and the experiment is quietly ended.

The reflex explanation is that online tests are noisy and the effect was too small to detect. Sometimes that is true. More often the offline number was measuring something that was never going to transfer, and the mechanism is identifiable in advance.

The metric is a stand-in for something else

Nobody optimises revenue directly during training. They optimise a loss, and then report a metric that is easier to compute than the outcome anyone cares about: an accuracy, a ranking measure, an overlap score.

Each substitution is a bet that movement in the proxy implies movement in the goal. That bet is defensible over small changes and unreliable over large ones, because optimisation pressure finds the region where the two come apart. A ranking model that gets better at ordering items it was already surfacing may have gained nothing at all in the position where users actually look. A classifier whose overall accuracy rose by becoming more decisive on easy cases has improved a number and changed no decision.

The check is to state, before the experiment, the causal chain from the metric to the outcome — this improves, therefore this behaviour changes, therefore this result moves. If the chain has a step nobody can articulate, the offline result is not evidence about the launch.

The evaluation set is not your traffic

Held-out sets are usually sampled from historical logs, and logs are shaped by whatever was serving at the time. They over-represent the requests the old system handled, the customers who were present, the products in the catalogue that season.

Live traffic contains what the log does not: novel queries, new tenants, the long tail that never accumulates enough volume to matter in a sample but sums to a substantial share of requests. If a candidate model’s advantage is concentrated in the well-represented middle, it will show up strongly offline and be diluted online, because the population weights are different.

Comparing the composition of the evaluation set against a recent slice of production traffic — by segment, by input length, by source, by whatever you slice on — takes an afternoon and heads off a whole category of surprise.

Logged feedback records what the old system chose

This is the deepest of the mechanisms and the easiest to overlook.

When outcomes are observed only for items the previous system surfaced or actions it took, the log is not a sample of the world. It is a sample of the incumbent’s decisions. Evaluating a new model against it rewards agreement with the incumbent, which is precisely the property you were trying to change.

The distortions compound. Users click things at the top because they are at the top. Items that were never shown have no recorded outcome, and treating that absence as a negative outcome punishes the exact candidates a better model would promote. A model that would have made different choices looks worse offline in proportion to how different its choices are.

Counterfactual estimation exists for this and works within limits: reweight the logged outcomes by how likely the old system was to make each choice, which requires having recorded those probabilities at the time. Nobody records them retroactively. If there is any prospect of offline policy evaluation, log the propensities from the start — it is a small change made once and impossible to add later.

Aggregates hide compositional damage

An overall improvement is consistent with a serious regression somewhere specific. If the offline set is dominated by one segment and the candidate trades performance in a smaller segment for gains in the larger one, the mean moves in the right direction and a group of users has a worse experience.

Online, that group may be the one that complains, churns, or accounts for a disproportionate share of value. The aggregate never mentioned them. Slice-level reporting is the entire remedy, and it costs almost nothing once the slices are defined.

The system, not the model

Finally, and mundanely: the model is one component. A better model behind an unchanged retrieval step, a truncating context window, a strict output parser or a conservative fallback may have no room to express its improvement.

This is worth testing early because it is cheap to test. Run the full pipeline on the evaluation set, not just the model. If the end-to-end improvement is much smaller than the model-level improvement, you have found the bottleneck, and it is not the thing you were about to spend a quarter on.

An ordered path instead of a leap

Given all of the above, the reasonable process is not “evaluate offline, then launch”. It is a sequence of increasingly expensive checks, each of which can stop the candidate.

Offline evaluation, sliced, on a set you have recently examined for representativeness. Then replay: run the candidate over a window of real traffic and inspect where it disagrees with the incumbent — disagreement rate and its shape tell you far more about what a launch will feel like than a scalar does. Then shadow, serving real requests without acting on the output, which exposes latency, error rates and operational surprises. Then a small online test with a pre-registered metric and a pre-agreed duration. Then staged rollout.

Each stage kills a different class of candidate, and the early ones are cheap. The purpose is not to be cautious for its own sake; it is that a flat online result is an expensive way to learn something a replay would have shown in a day.

What to do with a flat result

When it happens anyway, the useful response is to find out which mechanism applied rather than to move on. Was the improvement concentrated in a segment that is rare in production? Did the disagreement rate turn out to be tiny — a better model that behaves almost identically? Did the pipeline discard the difference?

Each of those has a different implication for the next attempt, and answering the question is usually a day’s work. Not answering it means the next candidate is built on the same misunderstanding, tested the same way, and ends the same way a quarter later.