Ayush Gupta

Your GPUs are probably idle

Notes from production · ~7 min read

If your inference bill is growing faster than your traffic, the instinct is to assume you need more GPUs. Usually you don't. You need the ones you have to stop waiting.

I spent the last two years running inference for a real-time video platform serving 500k+ monthly active users, where it was comfortably the most expensive part of the stack. We ended up sustaining 1,280+ requests/sec on hardware we never upgraded. These are the things that actually mattered, roughly in the order they became obvious.

first, the architecture that gets everyone

Almost everyone builds the same thing first, because on day one it works: each model in its own container behind a small HTTP server, scaled out behind a load balancer. One request in, one forward pass, one response out.

It stays reasonable right up until you look at a GPU utilization graph and find it sitting far below what you're paying for. The cause is structural, not a bug. A GPU is a throughput device wearing a latency device's clothes: feed it one input at a time and the kernel launch overhead, the host-to-device transfer, and every fixed per-call cost get amortized over exactly one sample. Between requests it does nothing at all.

Which is why the obvious fix fails. Adding replicas multiplies the bill without touching the waste, because every new replica is idle in precisely the same way as the old ones. If your GPUs are under about 40% busy, more of them will not help you.

dynamic batching, and the trade it forces

Triton Inference Server fixes this with dynamic batching. Instead of executing each request on arrival, it holds a brief queue and merges whatever accumulated into one batched forward pass. The per-call overhead is now amortized across the whole batch.

The parameter that decides everything is how long you'll wait for that queue to fill — max_queue_delay_microseconds. It is a direct and honest trade: every microsecond spent waiting for a fuller batch is a microsecond added to the request already queued.

This is where it goes wrong. Tuning for throughput alone gets you an impressive number for a slide and a p99 your users experience as the product feeling sluggish. The discipline that works is the reverse: fix your acceptable p99 first, then take as much throughput as that budget allows. Decide the ceiling before you start, because if you decide it afterwards you will talk yourself into raising it, one reasonable-looking increment at a time.

The right value is workload-specific and there's no shortcut — it depends on your arrival rate, your model's batch-scaling curve, and what your product can tolerate. Sweep it, plot throughput against p99, and pick the knee of the curve.

concurrent execution, and its ceiling

Batching alone still leaves gaps. While one batch copies memory across the bus, the compute units idle. Triton's instance groups let you run several copies of a model on one GPU so their phases interleave — one instance computes while another transfers.

More is emphatically not better. Each instance holds its own copy of the weights, so you're trading VRAM for overlap. Push too far and you get one of two outcomes. The good one is an out-of-memory error at load time, because you find out instantly. The bad one is landing just under the ceiling, where it benchmarks fine and then degrades under real traffic as memory pressure bites. That second failure mode is worth knowing about in advance; it does not look like a memory problem while you're chasing it.

you cannot tune what you cannot see

The thing I would insist on doing first, before touching a single config value: instrument.

We ran Prometheus against Triton's metrics endpoint and OpenTelemetry traces through the entire request path — gateway, queue, inference, response. That distinction carries more weight than it sounds like. Aggregate throughput tells you a number moved. A distributed trace tells you where the time went, and the answer is regularly somewhere other than the model.

The signals worth watching, in rough order of how often they changed my mind:

  • Queue time versus compute time, per request. The ratio tells you immediately whether you have a batching problem or a model problem. They have completely different fixes.
  • Realized batch size, not configured batch size. What you actually get under production traffic is usually well below the maximum you set. If the gap is large, your queue delay is too short or your arrival rate too thin for batching to do anything.
  • GPU utilization and memory as a pair. Either alone will mislead you.
  • p50, p95 and p99 tracked separately. Averages hide exactly the behaviour users complain about.

A meaningful share of what teams call "inference latency" turns out not to be inference. That is the single most common thing tracing exposes, and you cannot find it from throughput graphs.

match the server to the workload

Triton with dynamic batching earns its keep on vision, classification, and embedding traffic — workloads where every request costs about the same and a batch finishes together.

Autoregressive LLM serving is a different problem and mostly wants a different tool. Generation lengths vary enormously per request, so a fixed batch stalls on its slowest member while finished sequences hold their slots. That's what continuous batching in vLLM exists to solve: it evicts completed sequences and admits new ones mid-flight, so the batch refills continuously instead of draining to its slowest member. Paired with paged attention for KV-cache memory, it's a substantially better fit for generation than request-level batching.

Running both, for the workloads each suits, is a perfectly respectable architecture.

the short version

  • Measure utilization before you buy anything. It's the cheapest diagnostic available and it's the one teams skip.
  • Fix your latency ceiling first. Throughput tuning without a stated p99 budget always ends the same way.
  • Trace the whole path. Much of your "inference latency" isn't inference.
  • Watch realized batch sizes. The gap against configured size tells you whether batching is working at all.
  • Pick the server that matches the workload. Dynamic batching for vision, continuous batching for generation.

None of this made a single model faster. It stopped the hardware from waiting, and that turned out to be worth considerably more.