Optimizer Memory: Why Training Uses So Much More VRAM
A model that fits comfortably in VRAM for inference may need several times more memory to train because training stores gradients, optimizer states, and activations. This guide shows where that memory goes and which hardware and software levers can make training feasible.
A model fitting in VRAM for inference does not mean the same GPU can train it. Inference usually needs the model weights, temporary workspaces, and—during generation—a growing key-value (KV) cache. Training also needs gradients, optimizer states, saved activations, and often higher-precision copies of parameters.
For full-parameter training with Adam or AdamW, the model and optimizer data alone can require roughly 8–10 times the model's parameter count in bytes, before activations and framework overhead. Mixed precision, sharding, quantization, LoRA, and checkpointing can reduce that requirement, but each changes the training setup or its trade-offs.
Inference memory versus training memory
Inference memory
A simple inference memory estimate is:
Inference VRAM ≈ model weights + KV cache + temporary workspace + framework overhead
The largest item is usually the model weights:
Weight memory ≈ parameter count × bytes per parameter
Typical weight-storage sizes are:
| Weight format | Bytes per parameter | Approximate 7B weight size |
|---|---|---|
| FP32 | 4 | 28 GB |
| FP16 or BF16 | 2 | 14 GB |
| INT8 | 1 | 7 GB |
| 4-bit | 0.5, plus quantization overhead | About 3.5 GB plus overhead |
These are estimates using decimal gigabytes and idealized packing. Real files and runtime allocations can be larger because of metadata, scales, padding, CUDA workspaces, and memory fragmentation.
For a single short prompt, the KV cache may be modest. For long context, many simultaneous sequences, or large batch sizes, it can become a major part of inference memory.
Training memory
Training must calculate how each parameter should change and retain information needed to calculate those changes. A more useful estimate is:
Training VRAM ≈ weights + gradients + optimizer states + activations + temporary workspace
For mixed-precision full fine-tuning with AdamW, a common rough accounting per trainable parameter is:
| Training item | Typical storage assumption | Bytes per parameter |
|---|---|---|
| Compute weights | FP16 or BF16 | 2 |
| Gradients | Often FP16/BF16, implementation-dependent | About 2 |
| FP32 master weights | Used by many mixed-precision setups | About 4 |
| Adam first and second moments | Usually FP32 | About 8 |
| Subtotal | Before activations and workspace | About 16 |
Some implementations store gradients differently or use additional buffers. Treat 16 bytes per parameter as a planning estimate, not a universal rule.
The optimizer subtotal is why training can exceed inference memory so dramatically. A 7B-parameter model may require only about 14 GB for FP16 weights during inference, but roughly:
7,000,000,000 parameters × 16 bytes ≈ 112 GB
That 112 GB estimate covers only the parameter-related training state. It does not include activations, temporary buffers, the input batch, CUDA allocations, or memory needed by the training framework.
The main sources of training VRAM use
1. Model weights
The weights are present in both inference and training. Full-precision or half-precision training keeps the trainable parameters available for forward and backward passes.
Quantizing the base model can reduce this component substantially, but quantized training is not the same as ordinary full-precision training. Methods such as QLoRA generally keep the base model quantized while training separate low-rank adapter weights.
2. Gradients
After the forward pass, backpropagation calculates a gradient for each trainable parameter. These gradients normally remain in memory until the optimizer updates the parameters.
Full fine-tuning stores gradients for the entire model. Parameter-efficient fine-tuning, such as LoRA, stores gradients primarily for the adapter parameters, although the base model still participates in the forward and backward computation.
3. Optimizer states
Adam and AdamW track two running statistics for every trainable parameter:
- The first moment, often called the momentum term.
- The second moment, which tracks squared gradients.
These states are commonly stored in FP32. Together, they use approximately 8 bytes per parameter:
Adam moments ≈ parameter count × 2 states × 4 bytes
For a 7B-parameter model:
7B × 8 bytes ≈ 56 GB
This is in addition to the model weights and gradients.
Other optimizers have different memory profiles. An optimizer with fewer or lower-precision states may reduce VRAM use, but it can change convergence behavior, stability, or the training workflow. Do not assume that changing optimizers is a free capacity upgrade.
4. Master weights
Mixed-precision training may use FP16 or BF16 for computation while retaining an FP32 master copy for updates. This adds approximately 4 bytes per trainable parameter.
The exact behavior depends on the framework, optimizer, and precision configuration. Some newer optimizers and distributed-training systems manage this state differently.
5. Activations
Activations are intermediate results from the forward pass that backpropagation needs later. Their memory depends on more than parameter count:
- Sequence length
- Microbatch size
- Number of layers
- Hidden dimension
- Attention implementation
- Precision
- Whether gradient checkpointing is enabled
- Whether the model uses encoder-decoder or decoder-only architecture
This is why two models with similar parameter counts can have different training-memory requirements.
A useful high-level relationship is:
Activation memory generally rises with microbatch size × sequence length
The exact scaling varies by architecture and kernels. Increasing sequence length can be particularly expensive because attention-related intermediates may grow rapidly unless memory-efficient attention implementations are used.
6. Temporary workspace and framework overhead
Kernels may allocate temporary workspaces for matrix multiplication, attention, communication, and fused operations. PyTorch or another framework may also reserve memory in a caching allocator.
As a result, a configuration that theoretically needs 23 GB may still fail on a 24 GB GPU. Leave practical headroom rather than planning to use every last megabyte.
A realistic 7B training example
Consider a 7B-parameter decoder-only model.
Inference in FP16
The idealized weight requirement is:
7B × 2 bytes ≈ 14 GB
The actual GPU requirement is higher after adding the KV cache, runtime buffers, and framework overhead. Short-context, low-concurrency inference may fit on a GPU with substantially more than the weight size, but long contexts and multiple requests require additional KV-cache capacity.
Full FP16/BF16 fine-tuning with AdamW
A rough parameter-state estimate is:
Weights: 7B × 2 bytes = 14 GB
Gradients: 7B × 2 bytes = 14 GB
FP32 master copy: 7B × 4 bytes = 28 GB
Adam moments: 7B × 8 bytes = 56 GB
Parameter total: 112 GB
This is before activations and temporary workspace. A single consumer GPU cannot generally treat this as a 14 GB problem just because inference uses 14 GB of weights.
The model could be trained with multiple GPUs using parameter, optimizer, or data sharding. However, adding GPUs does not automatically pool VRAM: the software must explicitly distribute the model and training state.
4-bit LoRA or QLoRA-style fine-tuning
Now consider keeping the 7B base model in 4-bit form and training only low-rank adapter weights.
An idealized base-weight estimate is:
7B × 0.5 bytes ≈ 3.5 GB
Real usage is higher because quantized formats need scales, metadata, and runtime buffers. The trainable adapter has far fewer parameters than the full model, so its gradients and optimizer states are much smaller than those for full fine-tuning.
The dominant remaining variables are often:
- Activation memory
- Sequence length
- Microbatch size
- Quantization implementation
- Temporary workspace
- Whether gradient checkpointing is enabled
This setup can make fine-tuning possible on hardware that cannot support full-parameter training, but it is not equivalent to updating every base-model parameter. Adapter capacity, target modules, learning rate, and dataset characteristics all affect the result.
The most useful ways to reduce training memory
Reduce microbatch size
Microbatch size is the number of examples processed simultaneously on each device. Reducing it usually lowers activation memory directly.
If a batch of 4 fails, try a microbatch of 1. You can often preserve a larger effective batch using gradient accumulation:
Effective batch size =
microbatch size × gradient accumulation steps × number of data-parallel devices
For example:
1 example × 16 accumulation steps × 1 GPU = effective batch size of 16
Gradient accumulation reduces the peak memory of each forward/backward pass, but it increases the number of steps needed to process an effective batch. It does not make a single long sequence inexpensive.
Shorten the sequence length
Sequence length is one of the strongest activation-memory controls for language-model training. If your task does not require 8,192-token examples, training at 2,048 or 4,096 tokens can substantially reduce memory pressure.
Use realistic data preprocessing:
- Pack short examples when appropriate.
- Remove unnecessary repeated context.
- Truncate only when losing the end of an example is acceptable.
- Separate long-context training from ordinary fine-tuning.
Shortening sequences can change the task distribution, so it is a quality trade-off rather than a purely technical optimization.
Enable gradient checkpointing
Gradient checkpointing, also called activation checkpointing, discards selected forward activations and recomputes them during backpropagation.
The trade-off is:
- Lower activation VRAM
- More computation and longer training time
- Additional recomputation overhead
It is particularly useful when weights and optimizer states fit but activations push the job over the GPU limit.
Use lower-precision computation
FP16 and BF16 reduce the size of many weights, gradients, and activations compared with FP32. BF16 generally offers a wider exponent range, while FP16 may be available on a wider range of older hardware and software stacks.
Lower precision does not necessarily halve total training memory because optimizer states may remain in FP32 and some buffers are not reduced. It also requires compatible hardware, drivers, and training software.
Use parameter-efficient fine-tuning
LoRA and related methods freeze the base model and train small adapter matrices. This reduces:
- Trainable gradients
- Optimizer states
- Update-related memory
The base model still consumes weight memory and contributes to activation usage. LoRA is therefore a major reduction in optimizer memory, not a guarantee that any model will fit on any GPU.
Quantize the base model
Quantization can reduce base-weight memory, especially for adapter training. 8-bit and 4-bit methods have different accuracy, compatibility, and kernel-support trade-offs.
Check all of the following before choosing a quantized workflow:
- Whether the model architecture is supported
- Whether the training library supports the selected quantization format
- Which layers remain in higher precision
- Whether the GPU supports the required kernels
- Whether quantization affects the quality target for your task
A quantized model file's size is not a complete VRAM estimate. Runtime buffers, activations, adapters, and CUDA workspace still matter.
Offload weights or optimizer states
CPU offload moves some data from GPU VRAM to system RAM. NVMe offload can extend this idea to storage, but it is substantially slower and more operationally complex.
Offload is most useful when:
- The GPU is close to fitting the workload.
- The system has sufficient RAM.
- Training speed is secondary to making the run possible.
- The software has a tested offload path.
Offload does not create free performance. Data transfers can become a major bottleneck, and insufficient PCIe bandwidth or system RAM can make training impractical.
Shard the training state across GPUs
Distributed systems can split parameters, gradients, and optimizer states across GPUs. This is different from simply launching the same complete model on every GPU.
Relevant approaches include:
- Data parallelism
- Tensor parallelism
- Pipeline parallelism
- Fully sharded data parallel training
- ZeRO-style optimizer and parameter sharding
Each has communication, configuration, and scaling costs. Multiple GPUs also need suitable interconnects and software support; two cards do not automatically behave like one card with their combined VRAM.
A practical hardware decision path
Step 1: Identify the training method
First decide whether you need:
- Full-parameter fine-tuning
- LoRA or another adapter method
- QLoRA-style quantized adapter training
- Pretraining from scratch
- Continued pretraining
- Distillation or another specialized workflow
Do not size hardware from the model's inference requirement alone. The training method determines whether optimizer states exist for every parameter and whether the base model can remain frozen.
Step 2: Estimate parameter-state memory
For full AdamW fine-tuning, begin with a rough planning estimate:
Parameter-state memory ≈ parameter count × 16 bytes
Then add activations, workspace, and headroom. If your setup stores FP32 gradients or additional buffers, use a higher estimate.
For adapter tuning, calculate the base model's weight memory separately from the adapter's optimizer state. The adapter state may be small, but activations can still dominate.
Step 3: Set sequence and batch targets
Write down:
- Maximum sequence length
- Per-device microbatch size
- Number of GPUs
- Desired effective batch size
- Whether gradient accumulation is acceptable
If the run must support long sequences, prioritize VRAM and memory-efficient kernels over a configuration designed only for short examples.
Step 4: Decide which compromises are acceptable
Common choices are:
| Constraint | Useful lever | Main cost |
|---|---|---|
| Not enough VRAM | Smaller microbatch | More accumulation or lower throughput |
| Activations are too large | Gradient checkpointing | More computation |
| Base weights are too large | 8-bit or 4-bit quantization | Compatibility and possible quality trade-offs |
| Optimizer state is too large | LoRA or another adapter method | Not a full model update |
| One GPU is insufficient | Sharding or offload | Communication or transfer overhead |
| Training is too slow | More GPUs or larger batch | Higher hardware and power cost |
Step 5: Choose the GPU and system around the bottleneck
For a single-GPU adapter workflow, VRAM is usually the first filter. A faster GPU with insufficient VRAM cannot run a job that a slower, larger-memory GPU can run.
For full fine-tuning, a single consumer GPU may not be the appropriate target at all. You may need:
- Multiple GPUs with a supported sharding strategy
- A workstation or server GPU with much more VRAM
- CPU or NVMe offload
- A smaller model
- Adapter training instead of full fine-tuning
System RAM also matters. Offload and large model loading can require substantial RAM beyond the GPU requirement. Adequate cooling, power delivery, PCIe slots, and physical GPU clearance become important in multi-GPU systems.
Before buying, use Find a GPU to filter candidates by VRAM and intended workload. If you need multiple GPUs, high system RAM, or a specific power and cooling configuration, use Build a PC to plan the complete machine rather than selecting a graphics card in isolation.
Bottom line
Inference mostly needs model weights plus runtime memory. Training needs those weights plus gradients, optimizer states, saved activations, and temporary buffers. For full AdamW fine-tuning, the parameter-related state alone can be roughly 16 bytes per parameter in a common mixed-precision setup, before activations.
If full fine-tuning does not fit, the most practical sequence is usually:
- Reduce microbatch size.
- Shorten sequence length.
- Enable gradient checkpointing.
- Use FP16 or BF16 where supported.
- Switch to LoRA or QLoRA if updating every parameter is unnecessary.
- Use offload or sharding when the workload justifies the added complexity.
- Reassess the GPU, system RAM, and number of GPUs together.