Versioning the Model Means Versioning the Data
Reproducibility · July 1, 2026 · 9 min read
Somebody asks why the system made a particular decision three months ago. The question is reasonable, and it might come from a customer, an auditor, or an engineer trying to work out whether a bug is new.
Answering it requires knowing which model was serving at the time, what that model was trained on, what the code did, and what configuration it ran under. Most teams can produce the first of those four. The rest were reconstructed by memory or not at all, and the honest answer becomes an estimate delivered with a shrug.
The artefact is not the unit
Treating the checkpoint as the versioned thing is the original mistake. A checkpoint is a snapshot of parameters. It does not carry the training data, the preprocessing logic, the hyperparameters, the library versions, the random seeds, or the feature definitions used at inference. Change any of those and you have a different system with the same file.
The unit worth versioning is the whole tuple: data, code, configuration,
artefact. Every one of the four must be identifiable after the fact, and the
identifiers must be recorded together, because knowing that model v14 was
serving is useless if nobody can say what v14 was trained on.
This sounds like ceremony until the first time you need to reproduce a result and discover that the training script now points at a table that has been rewritten twice since.
Data is the hard one, and it is hard for a specific reason
Code has a natural identity: a commit hash. Configuration can be a file next to the code. Artefacts are immutable blobs with checksums. Data resists all of this because data is usually stored in systems designed to be updated in place.
A training set defined as “everything in this table where the flag is set” is not a version. It is a query whose answer changes every time somebody runs a backfill, corrects a record, or lets a late-arriving event land. Rerun the same pipeline a month later and you get a different dataset with the same name, which means a reproduction that fails tells you nothing — you cannot distinguish a broken pipeline from a moved input.
There are three workable answers, in ascending order of effort.
Snapshot the extract. Materialise the exact rows used, store them immutably, record the checksum. Costly in storage, trivially correct, and appropriate whenever the set is small enough.
Use a store with time travel. Table formats that retain versions let you pin a snapshot identifier instead of copying rows. The dataset becomes “table X at version N”, which is a real identity as long as retention outlives the models trained from it — a detail that is worth checking before relying on it.
Make the pipeline deterministic and pin its inputs. Record the source version, the filter, the seed and the transformation code, and require that rerunning them yields a byte-identical result. This is the cheapest to store and the most demanding to maintain, because a single non-deterministic step — an unordered join, a timestamp, a floating-point reduction in an unspecified order — quietly voids the guarantee.
Whichever you choose, the requirement is the same: given a model, produce the exact data it saw.
Lineage is what makes an incident survivable
The reason this matters operationally is not intellectual tidiness, it is recovery time.
Suppose a labelling error is discovered in an upstream source. The question that follows within minutes is: which models were trained on the affected period, and which of those are in production anywhere, including in the batch job nobody remembers? With lineage recorded, that is a query. Without it, it is an investigation, and the investigation happens under time pressure while the suspect models continue serving.
The same structure answers the friendlier version of the question — a model performs unexpectedly well and somebody wants to know why. Often the answer is that its training window happened to include an unusual period, which is only discoverable if the window is recorded.
Log the version with the prediction
Every prediction served should carry the identifier of the tuple that produced it. Not the model name, which is reused; not the endpoint, which is repointed; the specific version.
This one habit converts a whole class of unanswerable questions into straightforward ones. It lets you compare outcomes between versions without running an experiment framework. It tells you exactly when a rollout reached a given customer. And when a deployment is rolled back at three in the morning, it distinguishes predictions made before the rollback from those made after, which is otherwise guesswork based on timestamps and a deployment log with minute resolution.
Rollback is only real if it restores everything
Teams that version carefully still get caught here. The model is rolled back and the behaviour does not return, because the feature transformation code shipped separately and is still on the new version, or because a threshold was tuned for the new model and left in place.
If a model version’s inputs are computed by code that is deployed independently, the two can drift apart, and the combination running in production may be one that has never been evaluated. Either deploy them as a unit, or record their compatibility explicitly and refuse to start when the pairing is unknown.
The test is simple to state and uncomfortable to run: take a model that has been in production for a month, roll the whole system back to its exact configuration, and check that a set of recorded inputs produces the recorded outputs. Teams that try this the first time usually find one dependency they did not know they had.
The cheap version
None of this requires a platform. A registry can begin as a table with one row per trained model: an identifier, the data snapshot reference, the code commit, the configuration hash, the evaluation results, who approved it and when. Write to it from the training job so it cannot be skipped, and make the serving layer refuse to load an artefact that has no row.
That last constraint is what makes the whole thing hold. Documentation that depends on discipline decays; a deployment path that fails without the record does not.