Glossary
Definitions as they are used on this site. Machine learning operations borrowed its vocabulary from statistics, from distributed systems and from product analytics, and the three do not always agree. Where a term is genuinely used inconsistently, the entry says so rather than pretending otherwise.
Batch inference
Scoring a large set of records in one scheduled job rather than one at a time in response to requests. Latency is irrelevant, throughput is everything, and the operational problems are those of a data pipeline rather than of a service.
Checkpoint
A saved snapshot of a model’s parameters. Frequently spoken of as though it were the model, which it is not: without the code, configuration and data that produced it, a checkpoint can be served but not explained or reproduced.
Concept drift
A change in the relationship between inputs and correct outputs — the same input now warrants a different answer. The only form of drift that reliably degrades a model, and the only one that cannot be detected without labels.
Continuous batching
A serving scheduler in which the set of sequences being processed changes at every step: finished ones leave immediately and waiting ones join, rather than the batch being assembled, run to completion and replaced. Sometimes called in-flight batching or iteration-level scheduling; the three names describe the same mechanism.
Data drift
A change in the distribution of incoming feature values. Used inconsistently: some teams mean it strictly, as a statistical property of the inputs, while others use it loosely for any deterioration in a model, including concept drift and outright pipeline bugs. Worth establishing which sense is meant before agreeing that data drift is the cause of anything.
Evaluation set
Data held aside to measure a model rather than to fit it. The naming is notoriously unstable — “validation” and “test” are swapped freely between teams, and in some workflows a third split exists that is only touched at release. What matters is whether the set influenced training, not what it is called.
Feature
An input value the model consumes. Ambiguous in practice: a data engineer often means a column in a table, while a modeller means the transformed quantity that reaches the model. The distinction matters precisely because the transformation is where training and serving diverge.
Feature store
A system holding feature definitions plus both a historical record and a low-latency copy of current values, so that training and serving compute the same quantity from one definition. Its central function is preventing skew and supplying point-in-time correct joins; the catalogue and reuse benefits are real but secondary.
Golden set
A small, carefully curated set of examples treated as authoritative. Used inconsistently: for some teams it is a stable regression suite that must never change, for others it is a growing collection of known-hard cases. The two purposes conflict, and a team that has not decided which it has will eventually argue about whether a number is comparable to last quarter’s.
Ground truth
The label treated as correct for evaluation. Often less solid than the term suggests — it may be a human judgement with genuine disagreement in it, a proxy event that only correlates with the outcome, or a record of what the previous system decided.
Inference
Running a trained model to produce an output. Sometimes used to mean one forward pass, sometimes the entire request path including retrieval and post-processing. Cost and latency claims are frequently made in the first sense and read in the second.
Key-value cache
The intermediate state a generative model keeps for tokens already processed, so that each new token does not require recomputing the whole sequence. It grows with sequence length, and its memory footprint — not the batch size limit — is usually what actually constrains how many requests a replica can hold at once.
Label leakage
Information about the outcome finding its way into the features. Usually accidental, most often through joining a feature value as it stands now rather than as it stood when the event occurred. Produces excellent offline results and no production benefit.
Latency budget
The maximum acceptable response time, stated as a target on a specific percentile rather than an average. A budget that names no percentile is not a budget, because a system can satisfy it at the median while failing a substantial share of requests.
Lineage
The recorded chain connecting a prediction to the model that made it, that model to its training data and code, and that data to its sources. The property that turns “which models are affected by this upstream error” from an investigation into a query.
MLOps
An umbrella term for the operational practice around machine learning systems. Used very inconsistently — sometimes for deployment tooling specifically, sometimes for the whole lifecycle including data engineering and governance, sometimes as a job title covering whichever of those a particular organisation found it lacked.
Model
Depending on the speaker: a set of trained parameters, an artefact in a registry, or the entire deployed system including its pre- and post-processing. Most disputes about whether “the model” got worse are really disputes about which of these three is under discussion.
Model registry
The record of trained models and what each one is: its data snapshot, code commit, configuration, evaluation results and approval history. Most useful when the serving path refuses to load an artefact that has no entry, since a registry maintained by discipline alone decays.
Offline evaluation
Measuring a model against stored data rather than live traffic. Cheap, fast, repeatable, and dependent on the assumption that the stored data resembles what will actually arrive — an assumption that fails quietly rather than loudly.
Point-in-time correctness
The property that each training example’s features reflect only information available at the moment of the event being predicted. Easy to describe, easy to violate with an ordinary join, and the usual explanation for a model that evaluated far better than it performed.
Prefill and decode
The two phases of a generative request: an initial pass over the whole input, which is compute-heavy and scales with input length, and the token-by-token generation that follows, which is memory-bandwidth-bound. They respond to batching so differently that a single “tokens per second” figure covering both is close to meaningless.
Propensity
The probability with which the system in production chose a given action. Must be logged at the time; it cannot be reconstructed afterwards. Without it, offline comparison of a new policy against logged outcomes has no principled correction for the fact that the old policy chose what got observed.
Quantisation
Representing model weights, and sometimes activations, at reduced numerical precision to cut memory use and increase throughput. The quality cost is workload-specific rather than universal, which is why a quantised model needs re-evaluating on your own slices rather than trusting a general claim.
Replay
Running a candidate model over a window of recorded production traffic to see where it disagrees with the incumbent, without acting on its outputs. The cheapest way to find out what a launch would feel like, and the stage most often skipped between offline evaluation and a live experiment.
Shadow deployment
Serving real production traffic to a candidate model while discarding its outputs, so that latency, error rates and resource use are exercised under real conditions. Sometimes conflated with canary deployment, which is different: a canary’s output is actually used, by a small fraction of users.
Slice
A named subset of traffic — by segment, locale, input length, source system or any axis your users would notice. Reporting metrics per slice is what prevents an aggregate improvement from concealing a regression that affects a specific group of people.
Throughput
Work completed per unit of time, usually requests or generated units per second per replica. Only comparable when the sequence lengths, batch size, precision and hardware are all stated, which they frequently are not.
Training–serving skew
Any difference between how a feature is computed during training and how it is computed at serving time. Arises naturally because the two run under different constraints in different code, and it is diagnosed by recomputing logged production features through the batch path and comparing.
Utilisation
The fraction of available capacity actually in use. The dial that sets both cost per request and queueing delay: high utilisation makes each request cheap and makes waiting time rise sharply, and there is no configuration that avoids the trade.
Warm-up
The period after a replica starts during which it is not yet able to serve at full speed — weights loading, caches empty, compilation or autotuning still running. It sets a floor on how fast autoscaling can respond, which is why scaling is a cost mechanism rather than a defence against sudden spikes.