Inference Cost & Throughput Estimator

Fill in your own numbers. Everything recalculates as you type, entirely in the browser: no prices are baked in, no request is made, nothing is stored.

—cost per 1,000 requests
Per request—
Per day—
Per 30 days—
Peak output units/sec—
Replicas for peak—
Utilisation at average load—
Per-stream rate—
Generation time——

—

Why there are no prices in it

Every accelerator price, every per-token rate and every committed-use discount is specific to a vendor, a region, a contract and a date. A calculator that shipped with those numbers inside it would be quietly wrong within months, and the people most likely to be misled are the ones who did not know the figures were there.

So the tool asks for yours. If you are buying capacity by the token, use the rate on your own invoice. If you are running your own servers, work out what an hour of one replica costs you — instance price, or amortised hardware plus power and the fraction of a platform team that keeps it alive — and put that in instead. The arithmetic is the same either way; only the source of the unit price changes.

“Units” rather than “tokens” throughout, because the same sums apply to images, audio seconds, embeddings or rows scored. Use whatever your provider bills and your throughput measurement counts, as long as it is the same thing in both boxes.

What the throughput fields mean

Output units/sec per replica is a measurement, not a specification. Run your model, at your sequence lengths, at the batch size you intend to use, and record the sustained rate. A figure quoted for a different model shape, a different quantisation or an idle machine will produce a fleet estimate that is wrong in the expensive direction.

Concurrent requests per replica is how many sequences your server keeps in flight at once. It divides the replica’s throughput into per-stream throughput: raising it makes the fleet cheaper and every individual response slower. That trade is the whole substance of serving configuration, and the two numbers move together here so you can see it happen.

On the per-replica-hour basis, the daily figure assumes the peak-sized fleet stays up all day, because that is what most teams actually do. If you can genuinely scale down between peaks — and your warm-up time allows it — the real bill sits between that figure and the same sum computed at average load.

What the latency figure covers

The generation time shown is output units divided by the per-stream rate. That is the part you control by choosing hardware and batch size, and it is only part of a real response time.

It excludes the initial pass over the input, which grows with input length and can dominate for long prompts and short answers. It excludes queue time, which is near zero on an idle fleet and unbounded on a saturated one. It excludes network transit, authentication, retrieval, and any post-processing. Treat the number as a floor: your true tail latency is this plus everything the estimate leaves out, and the gap widens as utilisation rises.

The utilisation figure is a warning light for exactly that reason. It shows the fraction of your peak-sized fleet that is busy at average load, so a low number means you are paying for headroom — which may be precisely what you want if a person is waiting, and pure waste if the work is a nightly batch.

Reading the result

Three checks usually decide the design.

Does the monthly figure fit the budget you have? If not, the levers in descending order of effect are output length, batch size, and unit price — and output length is almost always the one nobody has looked at.

Does the fleet size hold at peak rather than at average? Sizing to average guarantees a queue during every busy hour, and queueing is the component that turns a slow system into a failing one.

Does the generation time leave room underneath the budget for the parts this tool does not model? If the floor already consumes the whole allowance, no amount of infrastructure work will make the target reachable, and the honest move is to renegotiate the target or shorten the output.