Everything around the model decides whether the model works
A trained model is a small artefact surrounded by a large amount of machinery: the pipeline that fed it, the evaluation that approved it, the server that batches its requests, the monitoring that decides when it has gone stale. Almost every production failure lives in that machinery rather than in the weights. These essays are about the parts that are unglamorous to build and expensive to get wrong.
All essays
- Why Offline Gains Vanish Online
The metric improved, the launch did nothing. There are a small number of mechanisms that cause this, and each of them can be checked before you spend a release cycle finding out.
- Feature Stores Solve a Skew Problem
The case for a feature store is not storage or reuse. It is that training and serving compute the same feature twice, in different code, at different moments — and the two answers diverge.
- Drift Is a Question, Not an Alarm
Input distributions move constantly and most of the movement is harmless. Wiring a retraining trigger to a drift statistic buys expensive churn; tying it to outcomes buys something worth having.
- Versioning the Model Means Versioning the Data
A model checkpoint is one member of a tuple. Storing it without the data, code and configuration that produced it gives you an artefact you can serve and cannot explain.
- Batching Is a Scheduling Policy
Grouping requests is usually treated as a throughput trick. It is really a scheduling decision that determines who waits, for how long, and what happens when memory runs out.
- Latency and Cost Are the Same Dial
Serving cheaply and serving fast are opposing settings of one control. Understanding why makes capacity planning tractable and stops teams chasing a target that only exists on an idle machine.
- Evaluation That Survives Production
An evaluation suite earns its keep by failing when the product would fail. Most suites fail somewhere else entirely — on a frozen set, against an average that hides the cases anyone would complain about.