How Much VRAM Does a 128K Context Window Use?
A 128K context window can consume anywhere from several gigabytes to dozens of gigabytes of VRAM just for the KV cache. The exact cost depends more on layer count, KV-head count, head dimension, cache precision, and batch size than on parameter count alone.
A 128K context window does not have one fixed VRAM requirement. The memory cost depends on the model architecture and inference settings.
As a practical illustration, a single sequence at 128K tokens might require approximately:
- 16 GiB of KV-cache memory for one illustrative smaller model architecture
- 20 GiB for one illustrative mid-sized architecture
- 40 GiB for one illustrative larger architecture
These figures are for the KV cache only. You must still add model weights, temporary workspace, activations, runtime overhead, and a safety margin. A quantized model with modest weight memory can therefore still require a large GPU—or system RAM offload—for a 128K prompt.
The short answer
The KV-cache memory for a transformer decoder can be estimated as:
KV-cache bytes = batch size × sequence length × number of layers × 2 × KV heads × head dimension × bytes per cache value
The 2 accounts for the key and value cache.
At 128K tokens, the sequence-length term is large:
128K cache memory = per-token KV memory × 128,000
In practice, use the model's actual architecture and inference engine settings. Do not estimate long-context VRAM from parameter count alone.
Weight memory and context memory are different
A local inference system usually needs memory for several separate components:
- Model weights — the parameters loaded from the model file.
- KV cache — stored keys and values for tokens already processed.
- Activations and temporary workspace — memory used during prompt processing and token generation.
- Runtime and allocator overhead — buffers, metadata, CUDA graphs, kernels, and fragmentation.
- Optional offload buffers — CPU RAM or other memory used when the model does not fit entirely in VRAM.
Weight quantization primarily reduces the first category. It does not automatically make the KV cache smaller.
A rough lower-bound estimate for weight storage is:
Weight bytes ≈ parameter count × bits per weight / 8
For example, a hypothetical 7-billion-parameter model at 4 bits per weight has a raw parameter-storage estimate of:
7,000,000,000 × 4 / 8 ≈ 3.5 GB
That is not the complete runtime requirement. Quantization scales, metadata, tensor alignment, non-quantized layers, buffers, and the inference engine add overhead. The exact model format and runtime matter.
The KV cache is separate. A 4-bit model can still need many gigabytes of KV memory at 128K context.
How KV-cache memory grows with context length
During generation, the model reuses the attention keys and values for previous tokens instead of recomputing them from scratch. Those tensors form the KV cache.
For a decoder-only transformer:
KV-cache bytes = batch × tokens × layers × 2 × KV heads × head dimension × bytes per value
Where:
batchis the number of simultaneous sequences.tokensis the number of cached tokens, including the prompt and generated tokens.layersis the number of transformer layers.KV headsis the number of key/value heads.head dimensionis the size of each attention head.bytes per valueis usually 2 for FP16 or BF16, and approximately 1 for an 8-bit cache.
The cache grows linearly with tokens:
- 32K tokens use one quarter as much KV memory as 128K.
- 64K tokens use half as much.
- 128K tokens use twice as much as 64K.
- 256K tokens use twice as much as 128K, if the model and runtime support it.
This linear growth is why long-context support can change the hardware requirement even when the model weights remain unchanged.
Why GQA and MQA matter
Not every attention head stores a separate key and value. Models using grouped-query attention (GQA) or multi-query attention (MQA) use fewer KV heads than query heads.
Fewer KV heads produce a smaller cache:
- Multi-head attention (MHA): KV heads generally equal query heads.
- Grouped-query attention (GQA): several query heads share each KV head.
- Multi-query attention (MQA): all query heads may share one KV head.
Two models with similar parameter counts can therefore have very different 128K context memory requirements.
Illustrative 128K KV-cache examples
The following examples use deliberately stated architectural assumptions. They are calculation examples, not universal requirements for every model in a given size class.
These calculations use:
- One sequence
- 128,000 cached tokens
- FP16 or BF16 KV values at 2 bytes each
- No allowance for runtime overhead
| Illustrative architecture | Layers | KV heads | Head dimension | Approx. KV memory at 128K |
|---|---|---|---|---|
| Smaller model-style example | 32 | 8 | 128 | 15.6 GiB |
| Mid-sized model-style example | 40 | 8 | 128 | 19.5 GiB |
| Larger model-style example | 80 | 8 | 128 | 39.1 GiB |
The calculation for the first row is:
32 × 128,000 × 2 × 8 × 128 × 2 bytes = 16,777,216,000 bytes
That is approximately 15.6 GiB.
Using 131,072 tokens instead of 128,000 produces a neat binary result of approximately 16 GiB for the same architecture. Real applications may also cache system prompts, chat templates, tool definitions, and generated output, so the usable context limit may be reached before the nominal 128K number.
A comparison with full multi-head attention
To show why KV-head count matters, consider an otherwise identical illustrative architecture with 32 KV heads instead of 8:
32 × 128,000 × 2 × 32 × 128 × 2 bytes
This requires approximately 62.5 GiB of KV memory—four times the 8-KV-head example.
This is why parameter count is an incomplete way to estimate long-context memory. The attention layout can dominate the result.
Batch size multiplies the cache
The KV cache is normally stored separately for each active sequence. If the batch size doubles, the cache requirement approximately doubles:
Total KV cache = single-sequence KV cache × batch size
Using the smaller illustrative example above:
| Batch size | Approx. FP16/BF16 KV memory at 128K |
|---|---|
| 1 | 15.6 GiB |
| 2 | 31.3 GiB |
| 4 | 62.5 GiB |
| 8 | 125 GiB |
These values assume every sequence is actually holding 128K cached tokens. A dynamic batching system may have a different distribution of sequence lengths, but the same principle applies: total cached tokens across active requests drive memory use.
For a local single-user setup, batch 1 is often the relevant starting point. For a server handling multiple users, batch size and aggregate active tokens can become more important than the nominal context length.
Cache precision can reduce the context penalty
KV caches are commonly stored in FP16 or BF16, using approximately 2 bytes per value. Some inference engines support lower-precision caches, such as 8-bit formats.
If the cache uses approximately 1 byte per value:
8-bit KV memory ≈ FP16/BF16 KV memory ÷ 2
For the illustrative examples:
| Illustrative architecture | FP16/BF16 KV | Approximately 8-bit KV |
|---|---|---|
| Smaller model-style example | 15.6 GiB | 7.8 GiB |
| Mid-sized model-style example | 19.5 GiB | 9.8 GiB |
| Larger model-style example | 39.1 GiB | 19.5 GiB |
These are estimates. Actual memory use can be higher because of scale data, block alignment, padding, and implementation-specific layouts.
Lower-precision KV caching can affect output quality, especially for demanding workloads or very long sequences. Check whether the selected inference engine supports the cache format for your model, GPU, and attention implementation. Weight quantization and KV-cache quantization are separate settings and should be evaluated separately.
A complete VRAM estimate
A useful planning equation is:
Required VRAM ≈ weight memory + KV-cache memory + workspace + runtime overhead + safety margin
For a 128K workload, the KV term is often the largest variable. A model may fit its quantized weights comfortably but fail when the context is extended.
For example, suppose an illustrative model has:
- 6 GiB of loaded weights
- 15.6 GiB of FP16/BF16 KV cache at 128K
- Additional workspace and runtime requirements
The system already needs more than 21.6 GiB before accounting for those additional buffers. A 24 GiB GPU may therefore be too tight even though the model's weight file appears much smaller.
The correct hardware decision should use the model's actual:
- Parameter count and quantization format
- Layer count
- KV-head count
- Head dimension
- Maximum active tokens
- KV-cache precision
- Batch size
- Inference framework
- GPU memory capacity and any CPU-offload behavior
Ways to reduce 128K memory use
Use a lower cache precision
If supported and acceptable for your workload, an 8-bit KV cache can reduce cache memory substantially compared with FP16 or BF16. Test quality with your actual prompts rather than assuming that weight quantization results will predict KV-cache behavior.
Reduce the active context
A nominal 128K context window does not mean every request must use 128K tokens. Setting a lower context limit reduces cache memory in direct proportion.
For example, moving from 128K to 64K reduces the KV-cache requirement by approximately half for the same model, cache precision, and batch size.
Keep batch size low
For personal use, limiting concurrent sequences can save more memory than changing weight quantization. A server workload should budget for the maximum aggregate active tokens, not only the longest individual request.
Use sliding-window or selective attention when supported
Some architectures and runtimes can retain only a recent portion of the attention history or use different attention patterns for different layers. This can reduce the amount of history that must remain active, but it may change the model's ability to use information from the beginning of a long document.
Support is model- and runtime-specific. A context setting of 128K does not by itself prove that a particular sliding-window implementation is available.
Summarize or retrieve older material
For document analysis and chat, you can summarize earlier turns, retrieve only relevant passages, or maintain external memory instead of keeping every token in the active context. This reduces both KV memory and prompt-processing work.
Offload to system RAM
CPU offload can make a model or cache fit when GPU VRAM is insufficient. The trade-off is lower performance and possible transfer bottlenecks. Whether offload works well depends on the runtime, interconnect, memory bandwidth, and which tensors are moved.
Treat offload as a compatibility and performance compromise, not as equivalent to having enough VRAM.
Select an architecture with fewer KV heads
When comparing models, inspect KV-head count and head dimension—not just parameter count. GQA and MQA architectures can have a much lower long-context memory cost than otherwise similar full multi-head designs.
Common planning mistakes
Mistaking the context limit for a fixed allocation
Some runtimes reserve memory for the configured maximum context, while others grow the cache dynamically. The observed allocation can therefore differ from the simple calculation. Configure the limit you actually need and verify behavior in the chosen runtime.
Counting only input tokens
The cache can include generated output as well as the prompt. A request with 120K input tokens and 8K generated tokens may approach a 128K total-token limit, depending on the API and runtime's definition of context length.
Assuming quantized weights solve long-context memory
A 4-bit weight format reduces model-weight storage. It does not necessarily reduce the KV cache, which may still use FP16, BF16, or another independently configured format.
Ignoring multiple users
A single 128K request and four simultaneous 128K requests have very different memory requirements. Batch size and total active tokens must be included in server planning.
Filling the GPU to the last megabyte
Exact allocations vary by framework and workload. Leave headroom for workspace, memory fragmentation, graphics or display use, and changes in prompt length. A configuration that barely fits in a static test can fail in normal use.
How to choose hardware for a 128K workload
Start with the model architecture rather than the advertised parameter count:
- Estimate weight memory from the model's quantization format.
- Obtain the layer count, KV-head count, and head dimension.
- Calculate KV memory at the maximum number of active tokens.
- Adjust for batch size and cache precision.
- Add runtime overhead and a practical safety margin.
- Decide whether CPU or multi-GPU offload is acceptable.
- Validate the exact model and runtime combination before buying hardware.
For a fast first-pass comparison, use the What can I run tool to check how candidate models and context settings fit your available hardware. Treat the result as a planning aid, then confirm the model's architecture and inference-engine settings because KV-cache behavior is implementation-specific.
The main decision is not simply “Can this GPU hold the model?” It is:
Can this system hold the weights + the required active context + the chosen concurrency level?
For 128K inference, that second question is often the one that determines whether a configuration is practical.