Batching Is a Scheduling Policy

A grid of serving slots across successive steps, with freed slots refilled immediately

Batching gets introduced to a system as an optimisation. Requests arrive one at a time, the accelerator is underused on each of them, so they are grouped and sent together. Throughput improves, the change is declared a success, and the batch size becomes a configuration value that nobody revisits.

What has actually been installed is a scheduler. It decides which requests run together, which ones wait, how long they wait, and what happens when there is not enough memory for all of them. Those are policy questions, and leaving them implicit is how serving systems acquire behaviour their operators cannot explain.

Why grouping helps at all

Modern accelerators are far better at arithmetic than at fetching the operands for it. When a single request is processed alone, most of the work is loading model weights out of memory to apply them to a very small amount of activation data. The device is waiting on memory, not computing.

Adding a second request to the same pass reuses that weight load. So does a third. The marginal cost of an extra sequence in the batch is genuinely small until some other resource runs out — which is why batching feels like free throughput right up to the point where it abruptly is not.

The two phases do not batch alike

A generative request has an initial pass over its whole input, which processes many tokens at once and saturates the device’s compute quite readily. Then it has a generation phase producing one token at a time per sequence, which is memory-bound and leaves compute idle.

Batching helps enormously in the second phase and much less in the first. A long input already keeps the device busy on its own; adding more long inputs to the same pass mostly just makes everyone wait. Meanwhile a batch of sequences in generation shares the weight loads and multiplies output for nearly no extra time.

This asymmetry is the origin of most serving pathologies. A single very long input arriving in the middle of a batch of short generations will stall all of them, because the pass containing it takes far longer than the passes around it. The victims are requests that did nothing wrong and are indistinguishable, from the client’s side, from an outage.

Static batching wastes what it collects

The naive implementation collects requests until the batch is full or a timer expires, runs the batch to completion, then collects the next one. It has two compounding flaws.

Every request pays the timer. Under light traffic, a request that could have been served immediately sits waiting for companions who never arrive. Under heavy traffic, the timer never matters and the queue does — so the delay you tuned for the quiet case is not the delay you get in the loud one.

And the batch runs at the speed of its slowest member. Sequences that finish early leave their slot occupied and idle while the longest generation continues. The device is nominally busy at a batch size it is not really achieving.

Continuous scheduling fixes both by making the batch a mutable set rather than a fixed group: finished sequences leave at each step, waiting ones join immediately, and nobody waits for a timer to fill a group. It is strictly better and it is now the default in serious serving stacks, which means the interesting questions have moved on to what constrains it.

The constraint is memory, and it is dynamic

Each sequence in flight holds intermediate state proportional to the tokens it has seen so far. That state grows as generation proceeds. The batch’s memory demand therefore increases over time even if no new request is admitted, and it depends on the lengths of the sequences, which you do not know in advance because you cannot see the future of an output that has not been generated yet.

This is why maximum batch size is a poor control. It is a proxy for the thing that actually runs out. A batch of short sequences may fit at a size several times larger than a batch of long ones, and a fleet configured for the average case will either run below capacity most of the time or hit memory exhaustion when several long requests coincide.

Admission control based on tokens rather than requests is more honest: estimate the state each candidate will need, admit while there is room, hold the rest in a queue you can see. The failure mode then becomes queueing — visible, bounded, reportable — rather than the ugly alternative, which is evicting sequences mid-flight and recomputing them later while the client waits.

Fairness is a decision you make by not making it

Once you have a scheduler, request ordering matters. First-come-first-served is simple and lets one enormous job delay a hundred small ones. Shortest-first maximises throughput and can starve long requests indefinitely under sustained load. Neither is right by default; the appropriate policy depends on whether your callers are people watching a screen or pipelines that do not care.

The practical middle ground is a small number of priority classes with separate admission limits, plus a cap on how much of the batch any single sequence may occupy. That prevents both classic pathologies without needing a sophisticated policy, and it makes the system’s behaviour under contention explainable to the team that gets paged.

What to measure

Batch size is not a health metric. Two numbers say much more.

Occupancy — how much of the theoretically available batch capacity is actually in use per step — tells you whether your fleet is doing the work you are paying for. Low occupancy under high queue depth means the scheduler is refusing work it could accept, usually because a limit is set on the wrong quantity.

Queue time as a share of total latency tells you whether your problem is the model or the policy. Teams routinely spend months on kernel-level optimisation for a system in which most of the elapsed time is spent waiting to start. A faster model reduces compute time; it does nothing about a queue that is deep because admission is over-conservative.

Instrument both, per priority class. The moment you can see them side by side, the batching configuration stops being folklore and becomes a set of trades you can argue about with evidence.