Home / Guides / LLM Capacity Planning: Requests, Context, Batch and VRAM

guide

LLM Capacity Planning: Requests, Context, Batch and VRAM

Updated 2026-09-02

Estimate the GPU memory, concurrency, and throughput an LLM inference service needs by connecting request rate, context length, batching, and KV-cache growth. Includes a practical sizing method and failure-mode checklist.

The short answer

Size an LLM serving system from the workload outward:

  1. Define the request rate, latency target, and acceptable queueing delay.
  2. Measure or estimate tokens per second for the selected model, quantization, runtime, and GPU.
  3. Calculate the required concurrent sequences from request rate and service time.
  4. Budget VRAM for model weights, KV cache, temporary activations, runtime overhead, and safety headroom.
  5. Load-test at the intended context length and concurrency. Increase hardware only when the measured bottleneck requires it.

The model's weight size is only the starting point. Long contexts and multiple simultaneous requests can consume substantial additional VRAM through the KV cache. A GPU that can load a model may still be unsuitable for serving it at the required latency or concurrency.

1. Define the serving objective

“Can this GPU run the model?” is not a complete capacity question. First write down what the service must do.

Record at least:

  • Model and quantization: Include the exact model variant, weight format, and runtime.
  • Prompt length: The number of input tokens, including system prompts, retrieved documents, chat history, and tool instructions.
  • Output length: The maximum and typical number of generated tokens.
  • Requests per second: Use average rate and peak rate separately.
  • Concurrency: The maximum number of requests that may be active at once.
  • Latency target: Separate time to first token (TTFT) from time per output token or total completion time.
  • Context window: The maximum total input-plus-output token capacity the service must support.
  • Reliability target: Decide whether occasional out-of-memory errors are acceptable. For most services, they are not.
  • Hardware constraints: Available VRAM, system RAM, power, noise, and whether multiple GPUs are practical.

A local assistant used by one person has a different objective from a shared API. The first may prioritize responsiveness for one request. The second may accept a small increase in individual latency to process several requests efficiently through batching.

Capacity versus performance

These terms are related but not interchangeable:

  • Capacity is how many requests or tokens the system can handle within resource limits.
  • Throughput is commonly measured in output tokens per second or requests per second.
  • Latency is how long a request waits and runs.
  • Concurrency is how many requests are active simultaneously.

A system can have high peak throughput but poor interactive latency if requests wait in a queue. Conversely, a low-concurrency system may feel fast for one user but fail when several users connect.

2. Understand what the runtime is doing

LLM generation usually has two distinct phases.

Prefill

During prefill, the runtime processes the input prompt and builds the attention state for it. Longer prompts generally require more computation and temporarily increase memory pressure.

Prefill strongly affects:

  • Time to first token
  • Prompt-processing throughput
  • Initial GPU utilization
  • The amount of KV cache allocated for the request

Decode

During decode, the model generates output one token at a time. Each new token attends to the existing context, so the KV cache must remain available until the request finishes.

Decode strongly affects:

  • Sustained output-token throughput
  • Completion time
  • Concurrent-request capacity
  • KV-cache growth as responses become longer

Continuous batching

Modern serving runtimes often use continuous, or dynamic, batching. Instead of waiting for an entire batch to finish, the scheduler adds and removes sequences as requests arrive and complete.

This can improve GPU utilization, especially when many requests are generating tokens at the same time. It does not make VRAM unlimited. Every active sequence still needs its own KV-cache allocation, and larger batches can increase temporary workspace and scheduling overhead.

Static batching is simpler but may leave the GPU idle when requests have different lengths. Continuous batching is usually more useful for an API with variable request arrivals, provided the runtime supports it and the workload is tested.

Paged attention and cache management

Some runtimes allocate KV cache in blocks or pages rather than reserving one large contiguous region per request. This can reduce fragmentation and make sharing VRAM among variable-length requests more practical.

Paging does not eliminate the underlying memory requirement. The cache still grows with the number of active sequences, their token counts, and the model's attention dimensions.

3. Connect workload inputs to VRAM

A useful planning model is:

Total VRAM = model weights + KV cache + temporary activations/workspace + runtime overhead + safety headroom

Each term depends on the exact model and runtime.

Model weights

A rough lower-bound estimate for uncompressed weights is:

Weight memory in bytes ≈ parameter count × bytes per parameter

Quantized weights use fewer bits per parameter, but the practical allocation also includes scales, metadata, alignment, and runtime-specific buffers. Therefore, a nominal bit-per-parameter calculation is an estimate, not a guaranteed VRAM requirement.

Do not size a GPU by comparing only its VRAM with the advertised size of a model file. The runtime may need additional memory, and some model components may use a different precision.

KV cache

The KV cache stores attention keys and values for each token already processed by each active sequence. A simplified relationship is:

KV cache memory ∝ active sequences × tokens per sequence × model attention dimensions × bytes per KV element

A more explicit approximation for a transformer with grouped-query or multi-query attention is:

KV bytes ≈ 2 × layers × KV heads × head dimension × tokens × bytes per element

Multiply that result by the number of active sequences. The exact formula and allocation behavior vary by architecture and runtime, so treat it as a planning estimate.

The important operational consequences are:

  • Doubling active context tokens approximately doubles KV-cache usage.
  • Doubling concurrent sequences approximately doubles KV-cache usage.
  • A longer maximum context can reduce the number of simultaneous requests that fit.
  • KV-cache precision may differ from weight precision.
  • Prefix caching can reduce repeated prefill work, but it does not always remove all per-request cache allocation.

Temporary memory and overhead

Leave room for:

  • Attention and matrix-multiplication workspaces
  • CUDA or other accelerator runtime allocations
  • Graph-capture or kernel buffers
  • Tokenizer and scheduler state
  • Communication buffers for multi-GPU inference
  • Fragmentation and allocator behavior

A practical planning rule is to reserve headroom rather than targeting the exact reported VRAM capacity. The appropriate margin depends on the runtime and workload; validate it with load testing instead of treating any single percentage as universal.

4. Estimate concurrency and throughput

Use Little's Law as a first-order estimate:

Average active requests ≈ arrival rate × average time in system

For example, if a service receives 0.5 requests per second and each request occupies the system for an average of 8 seconds:

0.5 requests/second × 8 seconds = 4 average active requests

This is an average, not a safe peak concurrency setting. Bursty traffic, long-tail prompts, retries, and slow generations require additional capacity or queueing controls.

For token-oriented planning:

Required output-token throughput ≈ requests per second × average output tokens per request

Suppose a workload produces 0.25 requests per second with 400 output tokens per request:

0.25 × 400 = 100 output tokens/second

That calculation says nothing about TTFT or whether the runtime can process the prompt phase quickly enough. It is only the sustained output requirement.

A service also needs to account for input processing:

Required input-token throughput ≈ requests per second × average input tokens per request

A workload with short answers and very long prompts can be limited by prefill even when its output-token requirement appears modest.

A small deployment example

Assume a private service has this target workload:

  • Peak arrival rate: 0.2 requests per second
  • Typical input: 1,500 tokens
  • Typical output: 300 tokens
  • Maximum supported total context: 8,192 tokens
  • Desired average completion time: no more than 10 seconds
  • Several users may submit requests at the same time

First, estimate the average active-request requirement:

0.2 requests/second × 10 seconds = 2 average active requests

Next, estimate output throughput:

0.2 × 300 = 60 output tokens/second

And input-token throughput:

0.2 × 1,500 = 300 input tokens/second

This service should not be sized for exactly two cache entries. Test at a higher concurrency, such as the expected burst level, while reserving VRAM for the full configured context and runtime overhead.

The next step is to measure the selected model and runtime on candidate hardware. If the measured output rate is below 60 tokens per second, requests will accumulate unless the service lowers its target, reduces output length, uses more hardware, or accepts queueing. If output throughput is adequate but TTFT is too high, investigate prompt length and prefill performance rather than only adding decode capacity.

The example intentionally avoids claiming a particular GPU is sufficient. Model architecture, quantization, runtime, and measured workload performance determine that answer.

5. Choose a capacity target, not just a maximum batch

Batch size is a control setting, not a universal performance rating.

A larger batch can:

  • Improve accelerator utilization
  • Increase aggregate token throughput
  • Increase VRAM use
  • Increase queueing and individual-request latency
  • Make long and short requests interfere with one another

A smaller batch can:

  • Improve interactive responsiveness
  • Waste available compute under multi-user load
  • Reduce the number of active KV caches
  • Limit aggregate throughput

For interactive services, optimize for a TTFT and per-token latency target. For offline jobs, optimize for total token throughput. For mixed workloads, use separate queues or service instances when the runtime and hardware budget allow it.

6. A repeatable sizing workflow

Step 1: Fix the model configuration

Specify the model revision, quantization, context limit, reasoning or tool-use settings, and runtime. A change in any of these can invalidate the estimate.

Step 2: Define representative traffic

Create test cases for:

  • Short and long prompts
  • Short and long outputs
  • One request
  • Expected average concurrency
  • Peak concurrency
  • Repeated prefixes
  • Worst-case context length

Use a distribution, not only one convenient prompt.

Step 3: Calculate first-order resource needs

Estimate:

  • Weight memory
  • KV-cache memory at expected and maximum concurrency
  • Input and output token throughput
  • Required system RAM for offload, loading, and the runtime
  • Storage for model files and temporary artifacts

Step 4: Benchmark one request and concurrent requests

Measure:

  • TTFT
  • Prefill tokens per second
  • Decode tokens per second
  • Total latency
  • Peak VRAM
  • Queue time
  • Rejected or failed requests

Test with the actual serving runtime. A benchmark from a different backend or model format may not predict production behavior.

Step 5: Add operational headroom

Do not run continuously at the edge of VRAM. Leave room for longer-than-average inputs, allocator variation, model reloads, and normal traffic bursts.

Step 6: Select hardware

Use Find a GPU to narrow candidates by VRAM and other system constraints. If the service requires multiple GPUs, chassis, power, cooling, or expansion planning, compare complete GPU servers rather than evaluating the card in isolation.

7. Tuning levers

If the system is too slow or runs out of memory, change one variable at a time.

Reduce context growth

  • Limit retained chat history.
  • Trim or summarize retrieved documents.
  • Set a realistic maximum output length.
  • Reject requests that exceed a defined token budget.
  • Avoid configuring an unnecessarily large context window for every request.

This often reduces KV-cache pressure, but aggressive truncation can reduce answer quality.

Control concurrency and queueing

Set a maximum number of active sequences and a queue limit. A bounded queue protects the GPU from unbounded memory growth and gives callers a clear overload response.

Admission control is usually better than allowing every request to start and failing unpredictably with an out-of-memory error.

Adjust batching

Increase batch capacity when aggregate throughput is poor and latency has room to increase. Reduce it when TTFT, tail latency, or VRAM usage is unacceptable.

Measure p95 or p99 latency where possible; averages can hide overloaded periods.

Use quantization or a smaller model

Lower-precision weights can reduce model memory and sometimes improve memory bandwidth efficiency, but quality and runtime support vary. A smaller model may provide a better user experience than a larger model that constantly queues or offloads.

Consider offload and multi-GPU operation carefully

CPU offload can make a model fit, but transfers across the CPU-GPU link may reduce performance substantially. Splitting a model across GPUs introduces communication and may require coordinated hardware and runtime support. Treat these as architectural choices to benchmark, not automatic fixes.

8. Failure-mode checks

The model loads, then generation fails

Likely causes include insufficient KV-cache space, temporary workspace, fragmentation, or a larger-than-expected request. Check peak VRAM during concurrent generation, not only memory after model loading.

Single-user tests pass, multi-user tests fail

The model weights fit, but per-request KV cache or batching overhead does not. Test the intended concurrency and maximum token budget.

Throughput is acceptable, but users wait too long

Aggregate throughput can hide queueing. Inspect TTFT, queue time, and tail latency. Reduce batch size or concurrency, shorten prompts, or dedicate capacity to interactive requests.

VRAM is available, but performance is poor

The bottleneck may be memory bandwidth, CPU tokenization, PCIe transfers, synchronization between GPUs, thermal throttling, or an inefficient runtime configuration. VRAM capacity alone does not predict tokens per second.

Performance degrades as conversations get longer

The growing KV cache and longer prefill phase are expected causes. Enforce history limits, summarize old turns, or route long-context requests separately.

Capacity changes after a model update

Recheck the model architecture, tokenizer, quantization, context settings, and runtime version. “Same parameter count” does not guarantee the same memory or performance behavior.

9. The practical decision rule

A candidate system is suitable only when all of these conditions hold under representative peak testing:

  • The model and runtime fit in the available memory.
  • KV cache fits at the configured concurrency and token limits.
  • Input and output throughput meet the workload target.
  • TTFT and completion latency meet the user-facing target.
  • Queue limits and overload behavior are predictable.
  • Adequate VRAM, system RAM, power, cooling, and operational headroom remain.

Start with a conservative workload estimate, benchmark the exact stack, and then tune concurrency and context together. Treat VRAM as a capacity budget and latency as a service-level objective; optimizing only one usually produces an unreliable inference service.

Related guides