Why Does Long Context Cause OOM?
Long conversations can exhaust VRAM even when the model and short prompts fit comfortably. This guide explains KV-cache growth, temporary attention memory, batching, and the configuration changes that usually fix context-related OOM errors.
A model can handle short prompts successfully and still run out of memory on a long conversation. The usual reason is not that the model weights suddenly become larger. It is that the runtime must store more information about every token already in the context.
In transformer inference, that stored history is primarily the key-value (KV) cache. As context length grows, the KV cache usually grows approximately linearly. A long prompt can therefore exhaust VRAM even when loading the model itself leaves plenty of memory free.
Other contributors include temporary memory used during prompt processing, multiple simultaneous requests, duplicated chat history, memory fragmentation, and a context setting that is larger than the model or hardware can practically support.
Start with the symptom
The error message and the point at which it occurs help identify the cause:
- OOM while loading the model: The weights, runtime overhead, or initial KV allocation do not fit. Context length may still be part of the initial allocation, but this is primarily a model-fit problem.
- OOM when submitting a long prompt: The KV cache or prompt-processing workspace is the leading suspect.
- OOM only during generation: The cache grows as new tokens are generated, or temporary attention memory peaks during decoding.
- OOM only with several users or requests: The runtime is allocating a separate cache for each active sequence, or batching is increasing the total token count.
- Memory usage rises after many requests and does not fall: The allocator may be caching memory, or a server may be retaining conversations, requests, or model copies.
- The operating system kills the process instead of reporting a CUDA error: This may be system RAM exhaustion, swap pressure, or unified-memory spillover rather than a pure VRAM OOM.
First determine which memory pool is exhausted. On an NVIDIA system, monitor the GPU while reproducing the problem:
nvidia-smi
watch -n 1 nvidia-smi
For a process-level view, use the monitoring facilities provided by your runtime. In PyTorch-based applications, these checks can help distinguish allocated memory from reserved memory:
import torch
print(torch.cuda.memory_summary())
print("allocated:", torch.cuda.memory_allocated() / 1024**3, "GiB")
print("reserved: ", torch.cuda.memory_reserved() / 1024**3, "GiB")
reserved memory can be larger than allocated because PyTorch keeps blocks in its caching allocator. That alone does not prove a leak. A steadily increasing allocated value across completed requests is more concerning.
The most likely causes, in order
1. KV-cache growth with context length
During autoregressive generation, the model computes keys and values for the tokens it has seen and keeps them available for later tokens. The cache grows with:
- Number of transformer layers
- Number of KV heads
- Head dimension
- Number of tokens in the context
- Number of active sequences
- Bytes used for each cached value
A useful estimate for a decoder-only model is:
KV cache bytes ≈
2 × layers × KV heads × head dimension × tokens × batch size × bytes per element
The factor of 2 represents the key and value tensors.
If the model uses grouped-query attention or multi-query attention, the number of KV heads can be much smaller than the number of attention heads. That substantially reduces KV-cache memory. If it uses full multi-head attention, the cache is larger.
The formula is an estimate, not an exact allocation forecast. Runtime metadata, padding, alignment, page tables, temporary buffers, and implementation details add overhead.
Worked estimate
Suppose a model has:
- 32 layers
- 8 KV heads
- 128 dimensions per head
- 16,384 tokens
- FP16 KV values, or 2 bytes per element
- One active sequence
Then:
2 × 32 × 8 × 128 × 16,384 × 2 bytes
= about 2.0 GiB
A second active sequence of the same length would approximately double that cache requirement. A context limit of 32,768 tokens would approximately double it again, assuming the other variables remain unchanged.
This is why reducing the model's context limit can fix an OOM without changing the model weights.
2. The runtime reserves the maximum context up front
Some inference engines allocate a KV cache based on the configured maximum context, not the number of tokens currently in the conversation. If the maximum is set to 32K or 128K, the runtime may reserve a large block before you have used all of it.
Check settings such as:
context sizemax sequence lengthmax model lengthmax tokensnum parallelbatch sizemax batch tokens
Names differ by application. For example, a local command-line runner may expose a context-size option, while a server may expose separate limits for sequence length, concurrent sequences, and batched tokens.
Start with a conservative context limit that matches your actual workload. Do not configure the largest context the model claims to support unless you have measured that your hardware and runtime can sustain it.
3. Multiple requests multiply the cache
Context memory is usually charged per active sequence. A server handling four conversations can need roughly four times the KV cache of one conversation at the same length.
The important quantity is often total active tokens:
Active KV memory ≈ per-token KV memory × total tokens across active sequences
Total tokens may include:
- User prompts
- Assistant history
- System prompts
- Tool calls and tool results
- Tokens currently being generated
- Padding added by a batched implementation
A server with a small model can therefore run out of VRAM under concurrent use even though one interactive conversation works.
For troubleshooting, set concurrency to one and disable or reduce continuous batching if your runtime allows it. If the OOM disappears, the model is probably not the only issue; the serving configuration is exceeding the available cache budget.
4. Prefill uses temporary memory
Processing a long prompt is called prefill. It can require more temporary workspace than generating one token at a time.
During prefill, the runtime may need memory for:
- Attention intermediates
- Activations
- Workspaces selected by the GPU kernels
- Input and output buffers
- Logits or other requested outputs
This can create a peak even if the steady-state KV cache appears to fit. A prompt that fails immediately after submission is especially consistent with a prefill peak.
Efficient attention implementations can reduce intermediate memory, but they do not make the KV cache free. Flash or memory-efficient attention mainly changes temporary attention-memory behavior; the cache still grows with context.
5. The conversation contains more tokens than expected
A chat UI may show a manageable number of messages while sending a much larger request. Common sources of hidden tokens include:
- The full conversation being resent on every turn
- Large system instructions
- Tool schemas
- Tool outputs
- Attached documents
- Retrieved passages
- Repeated summaries or duplicated messages
- Hidden reasoning or application metadata
Count the actual tokens sent to the model, not just the visible words. If your application uses a tokenizer, log the input-token count before inference.
A rough word count is not a reliable substitute. Tokenization varies by language, code, formatting, and model family.
6. Memory fragmentation or allocator behavior
An OOM can occur even when the total amount of apparently free VRAM seems sufficient. The runtime may not have a contiguous block of the required size, or cached allocations may be poorly shaped for the next request.
Signs include:
- The first long request works, but a later request fails
- Restarting the process temporarily fixes the issue
- Reserved memory is much larger than allocated memory
- Failures depend on request order or prompt length
Useful tests are:
- Restart the inference process.
- Run the same prompt first, without other workloads.
- Use a fixed context and batch size.
- Reduce concurrency.
- Compare allocated and reserved memory using your framework's diagnostics.
Emptying a framework cache can release reusable blocks, but it does not reduce the model's true memory requirement. Treat it as a diagnostic or cleanup step, not a solution to insufficient capacity.
7. Model copies, offload, or another GPU workload consume the headroom
Check whether another process is using the GPU:
nvidia-smi
Also verify that the runtime has not loaded:
- Two copies of the model
- A draft model for speculative decoding
- An embedding or reranking model
- A separate vision encoder
- Multiple worker processes
- CPU/GPU offload buffers that consume additional memory
A model that barely fits at startup has little room for context growth. Practical operation requires headroom for the KV cache and temporary workspaces.
What to change first: no-cost configuration fixes
Apply these changes before buying hardware. They are usually the fastest way to confirm that context growth is the cause.
Reduce the maximum context
Lower the context setting to the largest value your real workload needs, with some margin. If a 16K context works and 32K fails, that is valuable diagnostic information even if you eventually need a hardware upgrade.
A lower context limit does not make each token lower quality. It limits how much history the runtime can retain at once.
Trim or summarize old messages
Keep recent turns verbatim and summarize older turns. A practical application pattern is:
- Reserve space for the user's new request.
- Reserve space for the expected answer.
- Reserve space for tools or retrieved documents.
- Use the remaining budget for conversation history.
- Replace older history with a compact, verified summary when needed.
Summarization trades exact historical wording for lower memory use. Keep facts, decisions, constraints, and unresolved tasks in the summary rather than copying every exchange.
Use a sliding-window or rolling context
If the task does not require the entire conversation, retain only the most recent tokens or messages. This prevents the cache from growing indefinitely during a long session.
Sliding-window behavior must be supported by the model and runtime in a useful way. Simply setting a large context and hoping the application drops old messages will not help if the full history is still sent.
Reduce concurrency and batch limits
For a local interactive setup, start with one active sequence. In a server, reduce:
- Maximum concurrent sequences
- Maximum batched tokens
- Request batch size
- Number of worker processes
Throughput may decrease, but the peak memory requirement should become easier to predict.
Avoid unnecessary output and tool payloads
Limit maximum generated tokens where appropriate. Remove duplicated tool schemas, old tool results, and large retrieved passages. This helps both the input context and the generated-token portion of the cache.
Reducing output length does not fix an already oversized prompt, but it can prevent the cache from growing beyond the limit during generation.
Use supported KV-cache compression
Some runtimes support storing the KV cache in a lower-precision format than the model's weights, or using cache quantization. This can reduce memory substantially, but it may involve quality, compatibility, or speed trade-offs.
Do not assume that quantizing model weights also quantizes the KV cache. They are separate settings. Verify the runtime's documentation and inspect memory behavior after enabling the option.
Enable memory-efficient attention when available
If the OOM occurs during prompt ingestion rather than after sustained generation, a memory-efficient or flash-attention backend may reduce temporary workspace requirements.
This does not eliminate linear KV-cache growth. It is most useful when the cache fits but the prefill peak does not.
Check for a stale process
A crashed or abandoned worker may continue holding VRAM. Confirm which process owns the memory, stop it cleanly, and restart the inference server if necessary.
Hardware and model changes
If configuration changes do not provide enough context, the remaining options are to reduce the model's memory footprint or add memory capacity.
Use a model with a smaller KV cache
KV-cache requirements depend heavily on the number of KV heads, layers, and head dimension—not only on parameter count. Two models with similar weight sizes can have different long-context memory behavior.
When comparing models for local use, check:
- Number of layers
- KV-head count or attention configuration
- Supported context length
- Whether the runtime supports the model's attention implementation
- KV-cache data type options
A model with grouped-query attention may be a better fit for long context than a similarly sized model with a larger KV cache.
Quantize the model weights
Weight quantization frees VRAM for context, but it does not automatically reduce KV-cache memory. It helps when the weights themselves consume most of the available memory and the freed space can accommodate the cache.
For a long-context workload, evaluate both separately:
Total memory ≈ model weights
+ KV cache
+ runtime overhead
+ temporary peak memory
Lower-bit weights can introduce quality changes and may alter speed depending on the hardware and backend. Treat quantization as a capacity trade-off, not a guaranteed long-context fix.
Add GPU memory
More VRAM is the most direct hardware fix when the workload is otherwise suitable. The required amount depends on the model, quantization, context, concurrency, and runtime overhead.
Do not size a GPU by model-weight size alone. Leave room for:
- The complete model
- The intended KV cache
- Prefill and generation workspaces
- The operating system and other GPU applications
- Multiple active sequences, if applicable
A GPU that can load the model may still be unsuitable for the desired context length.
Use CPU RAM or unified-memory offload carefully
Offloading layers or cache data to system RAM can make a workload fit, but it often reduces speed and increases latency. Transfers across the CPU-GPU link can become the bottleneck, especially during interactive generation.
This is a useful fallback when fitting the model is more important than response speed. It is not equivalent to having the same capacity in dedicated VRAM.
Use multiple GPUs only when the runtime supports the required strategy
Multi-GPU execution can distribute weights and, depending on the runtime, cache or computation. But support varies by model format and inference engine. Communication overhead, uneven memory sizes, and tensor- or layer-splitting behavior affect the result.
Verify multi-GPU support for the exact runtime and model before treating several smaller GPUs as a replacement for one larger-memory GPU.
A practical memory-budget method
Estimate the major components separately:
Required memory ≈
model weights
+ KV cache at maximum active tokens
+ runtime overhead
+ prefill/generation peak
+ other GPU allocations
For the KV portion:
KV cache bytes ≈
2 × layers × KV heads × head dimension
× active tokens × bytes per KV element
For multiple requests:
Active tokens =
tokens in request 1
+ tokens in request 2
+ ...
+ tokens currently being generated
These formulas are planning estimates. Confirm them with the specific runtime because paged caches, padding, cache quantization, and kernel workspaces change the actual allocation.
A useful test sequence is:
- Load the model with one sequence.
- Record baseline VRAM after loading.
- Test a short prompt.
- Increase prompt length in steps.
- Record peak VRAM for each step.
- Repeat with generation disabled or limited, if the runtime supports that test.
- Repeat with concurrency increased one sequence at a time.
If memory rises roughly with prompt length, the KV cache is the dominant factor. If it jumps at a particular prompt size, investigate prefill workspaces, batching thresholds, or allocator behavior.
Final decision tree
Use this sequence to narrow down the fix:
- Does the model fail before a prompt is submitted?
- Yes: reduce model size or weight precision, remove duplicate processes, or use a GPU with more VRAM.
- No: continue.
- Does the failure occur only when the prompt or conversation becomes long?
- Yes: reduce context, trim history, summarize, or use a smaller KV-cache configuration.
- No: continue.
- Does one short conversation work but multiple conversations fail?
- Yes: reduce concurrency, batch size, or maximum active tokens.
- No: continue.
- Does the failure happen during prompt ingestion and disappear with memory-efficient attention?
- Yes: the prefill workspace is likely the peak; use that backend or reduce prompt size.
- No: continue.
- Does restarting the process fix the problem temporarily?
- Yes: inspect allocator fragmentation, retained requests, worker processes, and memory leaks.
- No: continue.
- Does reducing model weight precision help, but the cache still limits context?
- Yes: weight memory was part of the problem; address the remaining KV-cache requirement separately.
- No: continue.
- Do you need the current model, context, and concurrency together?
- Yes: plan for more VRAM, a model with a smaller KV cache, supported cache compression, or a carefully tested offload/multi-GPU setup.
- No: choose the lowest-cost combination of context trimming, concurrency reduction, and model-size changes.
If you are unsure whether your current system can support a model at the required context and concurrency, use What can I run to evaluate the workload. If the result points to a hardware limit, Build a PC can help plan a system with enough VRAM and supporting components rather than sizing the machine only for model loading.