Feature Stores Solve a Skew Problem
Data · July 18, 2026 · 9 min read
A model performs well in evaluation and disappointingly in production. Nobody changed the weights. The inputs look plausible. Somewhere between the notebook and the endpoint, a feature is being computed differently, and finding out which one takes a week.
This is training–serving skew, and it is the reason feature stores exist. Not reuse, not cataloguing, not a place to put things — those are consequences. The core problem is that the same feature is defined twice, in two codebases, under two different sets of constraints, and there is no mechanism forcing the two definitions to agree.
Why the same feature gets written twice
The training version runs in a batch context. It has the whole history available, it can join freely, it is written in whatever the data team uses, and it is optimised for throughput over a large table.
The serving version runs inside a request. It has milliseconds, it sees one entity at a time, it is written in the service’s language, and it can only reach data that a low-latency store already holds. Nobody chose to have two implementations; the constraints made it inevitable, and once there are two, they drift apart at the speed of ordinary maintenance.
The divergences are rarely dramatic. A rounding difference. A time window that is inclusive on one side and exclusive on the other. A null handled as zero in one implementation and as a category in the other. A default that appears when an upstream call times out in production but never appears in a batch job that had the data waiting. Each is small; each shifts a feature just enough to move the model off the region it was fitted on.
Point-in-time correctness is the subtler half
The second failure is worse because it inflates evaluation results rather than depressing production ones, so nothing looks wrong until launch.
Building a training set means assembling, for each labelled event, the feature values as they were at the moment of that event. The naive join takes the current value from a table that has since been updated — the customer’s present lifetime spend, the item’s present rating, the account’s present status. Every one of those has absorbed information from after the event, including information caused by the outcome you are trying to predict.
The model learns from it, evaluation confirms it works beautifully, and production has none of it. A serving path that is scrupulously correct will still underperform, because the model was fitted on a world in which the future was visible.
Doing this properly requires either storing values with valid-time ranges and joining on them, or keeping an append-only log of feature updates and reconstructing state at any timestamp. It is fiddly, easy to get subtly wrong, and the single most valuable thing a feature platform does on a team’s behalf.
What a feature store actually provides
Strip the marketing and there are four things:
A single definition per feature, executed for both training and serving, so the two cannot diverge without somebody editing one file.
Point-in-time joins, so training sets are assembled from values that existed when the event occurred.
An online store, holding current values at request latency, kept in sync with the offline history by the same pipeline that computes it.
Metadata — ownership, freshness, lineage, which models consume which feature — which sounds administrative until you need to deprecate something and cannot find out who would break.
Reuse across teams is a real benefit, but it is the last one to materialise and the first one quoted in a business case.
When not to build one
The costs are not hypothetical. You acquire a distributed system in the request path, whose availability now bounds your model’s availability. You acquire a sync process whose staleness is a new source of skew, replacing the one you removed. You acquire a definition language between the people who write features and the data they come from.
For a single model with a handful of features computed inside one service, this trade is plainly bad. For a batch-scored model with no online path at all, the online half is dead weight. For a team of three, the coordination benefit is zero because coordination is a conversation.
The threshold is roughly: several models, several teams, online serving, and features that depend on aggregations over history rather than on fields present in the request. Below that, the cheaper pattern wins.
The cheaper pattern
Most of the value can be had without adopting a platform, and it is worth knowing what the minimum viable version looks like.
Put every transformation in one library, imported by both the training pipeline and the serving code, with the entity’s raw attributes as input and the feature vector as output. One implementation, one place to fix a bug. This alone removes the majority of skew.
Log the features at serving time, exactly as they were computed, alongside the prediction. This gives you two things at once: a training set that is guaranteed point-in-time correct because it was assembled at the point in time, and a diagnostic record that answers “what did the model actually see” without reconstruction.
Then assert on the pair. Sample production requests, recompute their features through the batch path, and compare. Any divergence beyond tolerance is a bug with a name and a location. Teams that run this check consistently find that skew is not an occasional mystery but a steady trickle of small breakages, each individually trivial to fix once visible.
The point of the whole exercise
Whatever the implementation, the property being bought is that a feature means
one thing. When a model sees a value called orders_last_thirty_days, that
number was produced by the same logic, over the same window, with the same
null-handling, as the number the model was fitted on — and it reflects only
information that existed before the moment of prediction.
Everything else in this area is machinery for enforcing those two sentences. The machinery is worth buying when you have enough models and teams that discipline will not hold on its own, and worth skipping when it will.