Home / Guides / Continuous Batching Explained for Local LLM Servers

guide

Continuous Batching Explained for Local LLM Servers

Updated 2026-09-13

Continuous batching keeps a local LLM server processing new and active requests together instead of waiting for a fixed batch to finish. This guide explains how it improves throughput, what it costs in VRAM and latency, and how to tune it safely.

Continuous batching is a serving technique that lets an LLM runtime add, remove, and schedule requests while other requests are still generating tokens. Instead of waiting for every request in a fixed batch to finish, the server keeps the GPU supplied with work as requests enter and leave.

The practical result is usually better throughput under concurrent load. The trade-off is that higher concurrency consumes more VRAM, can increase queueing and latency, and requires limits on active requests, context length, and token budgets.

Continuous batching helps most when:

  • Multiple users or applications send requests at the same time.
  • Requests have different prompt lengths or output lengths.
  • The GPU would otherwise sit idle while a small batch waits for its slowest request.
  • You care about total generated tokens per second, not just the latency of one isolated request.

Start with the serving objective

Before changing a server setting, decide which result matters most:

ObjectiveUseful metricsTypical priority
Fast response to one userTime to first token (TTFT), time per output tokenLow concurrency
Interactive multi-user serviceTTFT, tail latency, queue timeBounded concurrency
Maximum total outputAggregate tokens per second, requests per secondHigher concurrency
Predictable production behaviorp95/p99 latency, error rate, VRAM headroomConservative limits

A useful throughput definition is:

Aggregate output throughput = total generated tokens / elapsed serving time

A server that generates more total tokens per second may still feel worse if requests spend too long waiting in a queue. Conversely, optimizing only for the fastest single request can leave substantial GPU capacity unused when several requests arrive together.

Prefill and decode behave differently

LLM serving has two broad phases:

  • Prefill: The runtime processes the input prompt. This is relatively compute-intensive and can process many prompt tokens together.
  • Decode: The runtime generates output tokens, usually one new token at a time for each active sequence.

A server may be able to process a large prompt efficiently during prefill but become limited by memory bandwidth, KV-cache capacity, or scheduling overhead during decode. Good continuous-batching implementations schedule both phases while trying to prevent long prompts from making active users wait indefinitely.

What continuous batching changes

Static batching

In static, or fixed, batching, the server groups requests together and runs them as a batch. The batch often remains active until all requests finish, or until the runtime applies padding and other handling for different sequence lengths.

This creates a synchronization problem. If one request produces a long answer and the others finish quickly, the shorter requests may leave the batch while the remaining work continues. New requests may have to wait for the next batch.

Continuous or iteration-level batching

Continuous batching makes scheduling decisions repeatedly during generation. At each scheduling step, the runtime can generally:

  1. Continue decoding for active sequences.
  2. Remove sequences that have finished or reached a limit.
  3. Admit queued requests when capacity is available.
  4. Run prompt-prefill work, sometimes in chunks.
  5. Allocate the next decode work across the current set of sequences.

The exact policy depends on the serving runtime. Some systems use token budgets, request limits, priority rules, or separate controls for prompt processing and generation. “Continuous batching” describes the scheduling approach, not one universal configuration.

A small scheduling example

Imagine four requests with different prompt and output lengths:

  • Request A has a short prompt and produces a short answer.
  • Request B has a long prompt and produces a long answer.
  • Request C arrives while A and B are decoding.
  • Request D arrives after A finishes.

With a fixed batch, C and D may wait for the current batch boundary. With continuous batching, the scheduler can remove A when it finishes and admit C or D without waiting for B to complete. The active set changes over time, while the GPU continues processing available work.

This does not make each token free. It reduces idle gaps and avoids tying new work to the slowest request in an older batch.

How concurrency, context, and batch affect VRAM

Continuous batching increases the number of sequences that may be resident at once. The main memory costs are:

  • Model weights.
  • KV cache for active sequences.
  • Temporary activations and attention workspaces.
  • Runtime allocations, CUDA graphs, communication buffers, and allocator fragmentation.
  • Any features such as speculative decoding, adapters, or multimodal inputs enabled by the server.

The KV cache is especially important because it grows with the number of active tokens, not just the number of requests.

A useful KV-cache estimate

A simplified estimate for KV-cache storage is:

KV-cache bytes ≈ layers × 2 × total cached tokens × KV heads × head dimension × bytes per element

Where:

  • 2 represents keys and values.
  • total cached tokens is the sum of tokens currently retained across active sequences.
  • KV heads may be lower than the number of attention heads when the model uses grouped-query or multi-query attention.
  • bytes per element depends on the cache data type and runtime.

This is an estimate, not a complete VRAM calculation. The runtime may use different layouts, quantization formats, block allocation, or temporary buffers.

For a simple planning model:

Total cached tokens ≈ active requests × average tokens per active request

A request with a 4,000-token prompt and 1,000 generated tokens can occupy roughly 5,000 cached tokens if the runtime retains the full sequence. Eight such requests would represent roughly 40,000 cached tokens. The exact allocation depends on the runtime and model architecture, but the scaling relationship is the important part.

Context length is a capacity limit, not a target

A maximum context setting defines how long a sequence is allowed to become. It does not mean every request will use that much memory. However, reserving for a high maximum can affect how the runtime plans memory, and long prompts or conversations can quickly consume KV-cache capacity.

If you increase both:

  • Maximum concurrent requests, and
  • Maximum context or total tokens per request,

then the worst-case KV-cache requirement can rise substantially.

A practical planning expression is:

Worst-case active tokens = concurrency limit × per-request token limit

This is intentionally conservative. Real workloads may use less, but configuring a limit that the GPU cannot sustain under a plausible workload can lead to out-of-memory errors or aggressive queueing.

Batch size is not always a single number

In local LLM servers, “batch size” can refer to several different limits:

  • Number of active sequences.
  • Number of prompt tokens processed in one step.
  • Number of decode tokens processed in one iteration.
  • Maximum total tokens in the scheduler.
  • Maximum number of requests admitted at once.

Check the runtime’s terminology before comparing settings. A token budget may be more informative than a request count when prompts vary widely in length.

Latency trade-offs

Continuous batching improves utilization by combining work, but more work in an iteration can affect responsiveness.

Time to first token

TTFT includes request queueing and prompt processing. A new request may wait while the scheduler processes other prompt or decode work. If the server aggressively prioritizes throughput, TTFT can increase during bursts.

Time per output token

Once generation begins, users often experience latency as the interval between streamed tokens. More active sequences can improve aggregate throughput, but the per-sequence rate may fall if the GPU or memory subsystem becomes saturated.

Tail latency

Average latency can hide problems. Measure p95 or p99 TTFT and completion time when testing concurrency. A configuration that looks good on average may produce unacceptable delays for requests that arrive during a long prefill or near a full KV cache.

The central trade-off is:

Higher concurrency → potentially higher aggregate throughput, but also higher memory use and queueing

There is no universal concurrency value that is optimal for every model, GPU, prompt distribution, or latency target.

A small deployment example

Suppose you are configuring a local server for a small internal application with intermittent concurrent use. The goal is to support several simultaneous requests while preserving interactive streaming.

Start with a deliberately conservative profile:

  • Maximum active sequences: 4
  • Maximum per-request context: a limit appropriate for the application
  • A bounded total-token or scheduler budget
  • A request queue rather than unlimited admission
  • Streaming enabled if users need early output
  • A reserved VRAM margin for runtime overhead

Then test with representative prompts:

  1. One short request.
  2. Several short requests at once.
  3. One long prompt combined with short prompts.
  4. Requests with long generated answers.
  5. A burst larger than the concurrency limit.

Record:

  • TTFT.
  • Time per output token.
  • Total output tokens per second.
  • Queue time.
  • Peak VRAM allocation.
  • Whether requests are rejected, delayed, or terminated.
  • p95 latency during bursts.

If VRAM remains comfortable and latency is acceptable, increase the active-sequence or token limit gradually. If output throughput improves only slightly while tail latency and memory use rise sharply, the previous setting may be the better operating point.

These values are a deployment example, not a recommendation for every model or GPU. The correct limits depend on the model’s weights, quantization, context behavior, runtime, and workload.

Tuning continuous batching safely

1. Establish a single-request baseline

Measure the server with one request before increasing concurrency. This gives you a reference for:

  • Single-request TTFT.
  • Single-request generation speed.
  • Peak VRAM use.
  • Any model or runtime errors.

Without a baseline, it is difficult to tell whether batching improved throughput or merely added queueing.

2. Increase concurrency in small steps

Test a sequence such as 1, 2, 4, and 8 active requests only if the hardware and application justify those levels. At each step, monitor both throughput and tail latency.

Stop increasing concurrency when one of these occurs:

  • VRAM approaches the safe operating limit.
  • Out-of-memory errors appear.
  • TTFT or p95 latency exceeds the application target.
  • Per-request token speed becomes too low.
  • Aggregate throughput stops improving.

3. Control total tokens, not only request count

Four short requests and four long-context requests are not equivalent. If the runtime supports a total-token budget, it can provide more predictable memory control than a request-count limit alone.

Use separate limits where available for:

  • Maximum active sequences.
  • Maximum prompt tokens admitted in one iteration.
  • Maximum total cached tokens.
  • Maximum generated tokens per request.

4. Leave VRAM headroom

Do not plan to consume every reported free megabyte. Memory fragmentation, temporary workspaces, longer-than-average prompts, and runtime behavior can make the peak higher than a simple estimate.

A safe margin is a deployment policy, not a fixed universal percentage. Determine it through repeated tests using worst-case prompts and expected concurrency.

5. Protect interactive requests from long prefills

A very long prompt can monopolize scheduler capacity or delay other users. Depending on the runtime, useful controls may include:

  • Chunked prefill.
  • Prompt-token limits.
  • Request priorities.
  • Separate queues.
  • Admission control.
  • Per-user rate limits.

These features are runtime-specific, so verify their behavior rather than assuming that a setting with a familiar name works the same way everywhere.

6. Benchmark realistic traffic

A benchmark with identical prompts and identical output lengths can overstate the benefit of batching. Include variation in:

  • Prompt length.
  • Output length.
  • Arrival timing.
  • Number of simultaneous users.
  • Conversation history.
  • Requests that stop early.
  • Requests that reach their output limit.

Measure both steady-state traffic and short bursts. Many local deployments perform well at steady load but develop long queues when several prompts arrive at once.

Common failure modes and checks

Out-of-memory errors

Likely causes:

  • Too many active sequences.
  • Excessive context or output limits.
  • KV cache growth.
  • A large prompt burst.
  • Runtime overhead or memory fragmentation.

Checks:

  • Reduce concurrency.
  • Reduce the total-token budget.
  • Lower maximum context or output length.
  • Confirm the model and cache data types.
  • Leave more VRAM headroom.
  • Test long prompts separately from short prompts.

Throughput does not improve

Likely causes:

  • The workload is mostly single-request.
  • The GPU is already saturated at low concurrency.
  • Requests are too short for batching gains to matter.
  • The bottleneck is CPU preprocessing, storage, networking, or tokenization.
  • Scheduler overhead offsets the additional work.

Checks:

  • Compare one request with several simultaneous requests.
  • Inspect GPU utilization and memory activity.
  • Measure input and output token rates separately.
  • Verify that requests are actually admitted concurrently.
  • Check CPU and host-memory utilization.

Latency becomes unacceptable

Likely causes:

  • The scheduler favors throughput over prompt responsiveness.
  • A long prefill delays shorter requests.
  • The queue is unbounded.
  • Concurrency is above the useful operating point.

Checks:

  • Measure queue time separately from model execution time.
  • Reduce active concurrency.
  • Add request timeouts or queue limits.
  • Test chunked prefill or priority controls if supported.
  • Track p95 and p99 latency rather than only averages.

The server rejects requests unexpectedly

Likely causes:

  • A maximum sequence or token limit was reached.
  • The runtime predicts insufficient KV-cache capacity.
  • The request exceeds the configured context window.
  • Admission control is working as configured.

Checks:

  • Log the request’s input and requested output token counts.
  • Compare them with the runtime’s active-token and context limits.
  • Confirm whether limits apply per request or across the entire batch.
  • Decide whether to queue, reject, truncate, or route oversized requests.

Results vary between runtimes

Continuous batching is not a guarantee of identical performance across servers. Different runtimes can use different attention kernels, cache allocation strategies, scheduling policies, quantization implementations, and prefill behavior.

Compare the settings and metrics using the same:

  • Model files and quantization.
  • Prompt set.
  • Output limits.
  • Concurrency pattern.
  • Hardware and software environment.

Planning hardware for continuous batching

Hardware selection should account for more than whether the model weights fit. For concurrent serving, ask:

  • Is there enough VRAM for weights, KV cache, and runtime overhead?
  • Does the GPU provide enough memory bandwidth for the intended decode workload?
  • Can the host CPU and system memory feed requests without becoming bottlenecks?
  • Does the system have sufficient cooling and power capacity for sustained load?
  • Is one larger GPU preferable to multiple smaller GPUs for the selected runtime?
  • What latency target must the system maintain at the expected concurrency?

The right GPU depends on the model, quantization, context distribution, concurrency target, and serving runtime. Use RigForAI’s Find a GPU tool to narrow hardware candidates around your VRAM and performance requirements. For a more complete multi-GPU or rack-style deployment, compare options in the GPU servers guide.

Continuous batching is most valuable when the workload can keep the GPU busy. If you only serve one short request at a time, a simpler configuration may deliver nearly the same user experience with less memory pressure. If several users share the server, controlled batching can turn otherwise idle gaps into useful work—but only when concurrency and token limits are sized for the available VRAM and latency target.

Related guides