Home / Guides / How Much VRAM Does a 128K Context Window Use?

guide

How Much VRAM Does a 128K Context Window Use?

Updated 2026-08-23

A 128K context window can consume anywhere from several gigabytes to dozens of gigabytes of VRAM just for the KV cache. The exact cost depends more on layer count, KV-head count, head dimension, cache precision, and batch size than on parameter count alone.

A 128K context window does not have one fixed VRAM requirement. The memory cost depends on the model architecture and inference settings.

As a practical illustration, a single sequence at 128K tokens might require approximately:

  • 16 GiB of KV-cache memory for one illustrative smaller model architecture
  • 20 GiB for one illustrative mid-sized architecture
  • 40 GiB for one illustrative larger architecture

These figures are for the KV cache only. You must still add model weights, temporary workspace, activations, runtime overhead, and a safety margin. A quantized model with modest weight memory can therefore still require a large GPU—or system RAM offload—for a 128K prompt.

The short answer

The KV-cache memory for a transformer decoder can be estimated as:

KV-cache bytes = batch size × sequence length × number of layers × 2 × KV heads × head dimension × bytes per cache value

The 2 accounts for the key and value cache.

At 128K tokens, the sequence-length term is large:

128K cache memory = per-token KV memory × 128,000

In practice, use the model's actual architecture and inference engine settings. Do not estimate long-context VRAM from parameter count alone.

Weight memory and context memory are different

A local inference system usually needs memory for several separate components:

  1. Model weights — the parameters loaded from the model file.
  2. KV cache — stored keys and values for tokens already processed.
  3. Activations and temporary workspace — memory used during prompt processing and token generation.
  4. Runtime and allocator overhead — buffers, metadata, CUDA graphs, kernels, and fragmentation.
  5. Optional offload buffers — CPU RAM or other memory used when the model does not fit entirely in VRAM.

Weight quantization primarily reduces the first category. It does not automatically make the KV cache smaller.

A rough lower-bound estimate for weight storage is:

Weight bytes ≈ parameter count × bits per weight / 8

For example, a hypothetical 7-billion-parameter model at 4 bits per weight has a raw parameter-storage estimate of:

7,000,000,000 × 4 / 8 ≈ 3.5 GB

That is not the complete runtime requirement. Quantization scales, metadata, tensor alignment, non-quantized layers, buffers, and the inference engine add overhead. The exact model format and runtime matter.

The KV cache is separate. A 4-bit model can still need many gigabytes of KV memory at 128K context.

How KV-cache memory grows with context length

During generation, the model reuses the attention keys and values for previous tokens instead of recomputing them from scratch. Those tensors form the KV cache.

For a decoder-only transformer:

KV-cache bytes = batch × tokens × layers × 2 × KV heads × head dimension × bytes per value

Where:

  • batch is the number of simultaneous sequences.
  • tokens is the number of cached tokens, including the prompt and generated tokens.
  • layers is the number of transformer layers.
  • KV heads is the number of key/value heads.
  • head dimension is the size of each attention head.
  • bytes per value is usually 2 for FP16 or BF16, and approximately 1 for an 8-bit cache.

The cache grows linearly with tokens:

  • 32K tokens use one quarter as much KV memory as 128K.
  • 64K tokens use half as much.
  • 128K tokens use twice as much as 64K.
  • 256K tokens use twice as much as 128K, if the model and runtime support it.

This linear growth is why long-context support can change the hardware requirement even when the model weights remain unchanged.

Why GQA and MQA matter

Not every attention head stores a separate key and value. Models using grouped-query attention (GQA) or multi-query attention (MQA) use fewer KV heads than query heads.

Fewer KV heads produce a smaller cache:

  • Multi-head attention (MHA): KV heads generally equal query heads.
  • Grouped-query attention (GQA): several query heads share each KV head.
  • Multi-query attention (MQA): all query heads may share one KV head.

Two models with similar parameter counts can therefore have very different 128K context memory requirements.

Illustrative 128K KV-cache examples

The following examples use deliberately stated architectural assumptions. They are calculation examples, not universal requirements for every model in a given size class.

These calculations use:

  • One sequence
  • 128,000 cached tokens
  • FP16 or BF16 KV values at 2 bytes each
  • No allowance for runtime overhead
Illustrative architectureLayersKV headsHead dimensionApprox. KV memory at 128K
Smaller model-style example32812815.6 GiB
Mid-sized model-style example40812819.5 GiB
Larger model-style example80812839.1 GiB

The calculation for the first row is:

32 × 128,000 × 2 × 8 × 128 × 2 bytes = 16,777,216,000 bytes

That is approximately 15.6 GiB.

Using 131,072 tokens instead of 128,000 produces a neat binary result of approximately 16 GiB for the same architecture. Real applications may also cache system prompts, chat templates, tool definitions, and generated output, so the usable context limit may be reached before the nominal 128K number.

A comparison with full multi-head attention

To show why KV-head count matters, consider an otherwise identical illustrative architecture with 32 KV heads instead of 8:

32 × 128,000 × 2 × 32 × 128 × 2 bytes

This requires approximately 62.5 GiB of KV memory—four times the 8-KV-head example.

This is why parameter count is an incomplete way to estimate long-context memory. The attention layout can dominate the result.

Batch size multiplies the cache

The KV cache is normally stored separately for each active sequence. If the batch size doubles, the cache requirement approximately doubles:

Total KV cache = single-sequence KV cache × batch size

Using the smaller illustrative example above:

Batch sizeApprox. FP16/BF16 KV memory at 128K
115.6 GiB
231.3 GiB
462.5 GiB
8125 GiB

These values assume every sequence is actually holding 128K cached tokens. A dynamic batching system may have a different distribution of sequence lengths, but the same principle applies: total cached tokens across active requests drive memory use.

For a local single-user setup, batch 1 is often the relevant starting point. For a server handling multiple users, batch size and aggregate active tokens can become more important than the nominal context length.

Cache precision can reduce the context penalty

KV caches are commonly stored in FP16 or BF16, using approximately 2 bytes per value. Some inference engines support lower-precision caches, such as 8-bit formats.

If the cache uses approximately 1 byte per value:

8-bit KV memory ≈ FP16/BF16 KV memory ÷ 2

For the illustrative examples:

Illustrative architectureFP16/BF16 KVApproximately 8-bit KV
Smaller model-style example15.6 GiB7.8 GiB
Mid-sized model-style example19.5 GiB9.8 GiB
Larger model-style example39.1 GiB19.5 GiB

These are estimates. Actual memory use can be higher because of scale data, block alignment, padding, and implementation-specific layouts.

Lower-precision KV caching can affect output quality, especially for demanding workloads or very long sequences. Check whether the selected inference engine supports the cache format for your model, GPU, and attention implementation. Weight quantization and KV-cache quantization are separate settings and should be evaluated separately.

A complete VRAM estimate

A useful planning equation is:

Required VRAM ≈ weight memory + KV-cache memory + workspace + runtime overhead + safety margin

For a 128K workload, the KV term is often the largest variable. A model may fit its quantized weights comfortably but fail when the context is extended.

For example, suppose an illustrative model has:

  • 6 GiB of loaded weights
  • 15.6 GiB of FP16/BF16 KV cache at 128K
  • Additional workspace and runtime requirements

The system already needs more than 21.6 GiB before accounting for those additional buffers. A 24 GiB GPU may therefore be too tight even though the model's weight file appears much smaller.

The correct hardware decision should use the model's actual:

  • Parameter count and quantization format
  • Layer count
  • KV-head count
  • Head dimension
  • Maximum active tokens
  • KV-cache precision
  • Batch size
  • Inference framework
  • GPU memory capacity and any CPU-offload behavior

Ways to reduce 128K memory use

Use a lower cache precision

If supported and acceptable for your workload, an 8-bit KV cache can reduce cache memory substantially compared with FP16 or BF16. Test quality with your actual prompts rather than assuming that weight quantization results will predict KV-cache behavior.

Reduce the active context

A nominal 128K context window does not mean every request must use 128K tokens. Setting a lower context limit reduces cache memory in direct proportion.

For example, moving from 128K to 64K reduces the KV-cache requirement by approximately half for the same model, cache precision, and batch size.

Keep batch size low

For personal use, limiting concurrent sequences can save more memory than changing weight quantization. A server workload should budget for the maximum aggregate active tokens, not only the longest individual request.

Use sliding-window or selective attention when supported

Some architectures and runtimes can retain only a recent portion of the attention history or use different attention patterns for different layers. This can reduce the amount of history that must remain active, but it may change the model's ability to use information from the beginning of a long document.

Support is model- and runtime-specific. A context setting of 128K does not by itself prove that a particular sliding-window implementation is available.

Summarize or retrieve older material

For document analysis and chat, you can summarize earlier turns, retrieve only relevant passages, or maintain external memory instead of keeping every token in the active context. This reduces both KV memory and prompt-processing work.

Offload to system RAM

CPU offload can make a model or cache fit when GPU VRAM is insufficient. The trade-off is lower performance and possible transfer bottlenecks. Whether offload works well depends on the runtime, interconnect, memory bandwidth, and which tensors are moved.

Treat offload as a compatibility and performance compromise, not as equivalent to having enough VRAM.

Select an architecture with fewer KV heads

When comparing models, inspect KV-head count and head dimension—not just parameter count. GQA and MQA architectures can have a much lower long-context memory cost than otherwise similar full multi-head designs.

Common planning mistakes

Mistaking the context limit for a fixed allocation

Some runtimes reserve memory for the configured maximum context, while others grow the cache dynamically. The observed allocation can therefore differ from the simple calculation. Configure the limit you actually need and verify behavior in the chosen runtime.

Counting only input tokens

The cache can include generated output as well as the prompt. A request with 120K input tokens and 8K generated tokens may approach a 128K total-token limit, depending on the API and runtime's definition of context length.

Assuming quantized weights solve long-context memory

A 4-bit weight format reduces model-weight storage. It does not necessarily reduce the KV cache, which may still use FP16, BF16, or another independently configured format.

Ignoring multiple users

A single 128K request and four simultaneous 128K requests have very different memory requirements. Batch size and total active tokens must be included in server planning.

Filling the GPU to the last megabyte

Exact allocations vary by framework and workload. Leave headroom for workspace, memory fragmentation, graphics or display use, and changes in prompt length. A configuration that barely fits in a static test can fail in normal use.

How to choose hardware for a 128K workload

Start with the model architecture rather than the advertised parameter count:

  1. Estimate weight memory from the model's quantization format.
  2. Obtain the layer count, KV-head count, and head dimension.
  3. Calculate KV memory at the maximum number of active tokens.
  4. Adjust for batch size and cache precision.
  5. Add runtime overhead and a practical safety margin.
  6. Decide whether CPU or multi-GPU offload is acceptable.
  7. Validate the exact model and runtime combination before buying hardware.

For a fast first-pass comparison, use the What can I run tool to check how candidate models and context settings fit your available hardware. Treat the result as a planning aid, then confirm the model's architecture and inference-engine settings because KV-cache behavior is implementation-specific.

The main decision is not simply “Can this GPU hold the model?” It is:

Can this system hold the weights + the required active context + the chosen concurrency level?

For 128K inference, that second question is often the one that determines whether a configuration is practical.

Related guides