Home / Models / Llama 3.2 3B Instruct
Llama 3.2 3B Instruct
3.2B parameters · native context 131,072 · llama
Estimated VRAM requirement
Weight memory uses parameter count × bits per weight, plus a quantization overhead factor. KV cache uses standard GQA formulas. Values are calculated estimates, not measured allocations.
| Quantization | Estimated weights | 2K total | 4K total | 8K total | 16K total | 32K total | 64K total | 128K total |
|---|---|---|---|---|---|---|---|---|
| FP16 | 6.1 GB | 7.6 GB | 7.8 GB | 8.2 GB | 9.1 GB | 10.8 GB | 14.3 GB | 21.3 GB |
| Q8 | 3.4 GB | 4.7 GB | 4.9 GB | 5.3 GB | 6.2 GB | 8.0 GB | 11.5 GB | 18.5 GB |
| Q6 | 2.7 GB | 3.9 GB | 4.1 GB | 4.5 GB | 5.4 GB | 7.1 GB | 10.6 GB | 17.6 GB |
| Q5 | 2.3 GB | 3.4 GB | 3.6 GB | 4.1 GB | 4.9 GB | 6.7 GB | 10.2 GB | 17.2 GB |
| Q4 | 1.9 GB | 3.0 GB | 3.2 GB | 3.7 GB | 4.5 GB | 6.3 GB | 9.8 GB | 16.8 GB |
| Q3 | 1.5 GB | 2.6 GB | 2.8 GB | 3.2 GB | 4.1 GB | 5.9 GB | 9.4 GB | 16.4 GB |
Find a GPU for this model
Run at the quantization and context selected below. Defaults are Q4 and 8K when you do not change the controls. Results are full-VRAM fits only — not a speed ranking.
- Run at
- Q4 · 8K context
- Required VRAM
- 3.7 GB
No fresh matched Amazon price is currently available for a full-VRAM fit. See compatible GPUs · Workstations · Rent a server · Buy vs rent.
Lowest-cost full-VRAM option
Sorted by lowest current Amazon price among GPUs that fit entirely in VRAM. This is not a performance ranking.
No compatible GPU currently has a fresh Amazon price.
More VRAM headroom
Sorted by leftover usable VRAM after the calculated requirement. Headroom is not tokens/sec.
NVIDIA B200 192 GB
- GPU VRAM
- 192 GB advertised
- Required VRAM
- 3.7 GB
- Usable GPU VRAM
- 172.8 GB
- VRAM headroom
- 169.1 GB — leftover usable memory after the calculated requirement, not a speed ranking
- Compatible
- Yes — fits entirely in VRAM
- Product SKUs
- 0 mapped · 0 Amazon-matched
- Amazon
- Check availability
NVIDIA H200 141 GB
- GPU VRAM
- 141 GB advertised
- Required VRAM
- 3.7 GB
- Usable GPU VRAM
- 126.9 GB
- VRAM headroom
- 123.2 GB — leftover usable memory after the calculated requirement, not a speed ranking
- Compatible
- Yes — fits entirely in VRAM
- Product SKUs
- 0 mapped · 0 Amazon-matched
- Amazon
- Check availability
NVIDIA GH200 96 GB
- GPU VRAM
- 96 GB advertised
- Required VRAM
- 3.7 GB
- Usable GPU VRAM
- 86.4 GB
- VRAM headroom
- 82.7 GB — leftover usable memory after the calculated requirement, not a speed ranking
- Compatible
- Yes — fits entirely in VRAM
- Product SKUs
- 0 mapped · 0 Amazon-matched
- Amazon
- Check availability
Other compatible GPUs
AMD Radeon RX 470 8 GB
- GPU VRAM
- 8 GB advertised
- Required VRAM
- 3.7 GB
- Usable GPU VRAM
- 7.2 GB
- VRAM headroom
- 3.5 GB — leftover usable memory after the calculated requirement, not a speed ranking
- Compatible
- Yes — fits entirely in VRAM
- Product SKUs
- 1 mapped · 0 Amazon-matched
- Amazon
- Check availability
NVIDIA Quadro P4000 8 GB
- GPU VRAM
- 8 GB advertised
- Required VRAM
- 3.7 GB
- Usable GPU VRAM
- 7.2 GB
- VRAM headroom
- 3.5 GB — leftover usable memory after the calculated requirement, not a speed ranking
- Compatible
- Yes — fits entirely in VRAM
- Product SKUs
- 6 mapped · 2 Amazon-matched
- Amazon
- Check price on Amazon
AMD Radeon RX 580 8 GB
- GPU VRAM
- 8 GB advertised
- Required VRAM
- 3.7 GB
- Usable GPU VRAM
- 7.2 GB
- VRAM headroom
- 3.5 GB — leftover usable memory after the calculated requirement, not a speed ranking
- Compatible
- Yes — fits entirely in VRAM
- Product SKUs
- 13 mapped · 4 Amazon-matched
- Amazon
- Check price on Amazon
Compatible GPUs
Status is a calculated memory fit against usable VRAM (advertised × 0.9). Open a GPU to see product SKUs and Amazon listings.
Compare GPUs for this model
Published chip comparisons where both GPUs fully fit this model at Q4 / 8K. Opens with this model preselected; that query is not indexed.
How fast can this model run?
llama.cpp CUDA decode at Q4 under llama-bench tg128 conditions. Not a ranking. Open a pair for prefill and methodology.
| GPU | Decode | State | Confidence |
|---|---|---|---|
| NVIDIA RTX 5090 32 GB | ~568 tok/s | Calibrated estimate | MEDIUM |
| NVIDIA RTX 4090 24 GB | ~357 tok/s | Calibrated estimate | MEDIUM |
| NVIDIA RTX 5080 16 GB | ~349 tok/s | Calibrated estimate | MEDIUM |
| NVIDIA RTX 5070 Ti 16 GB | ~345 tok/s | Calibrated estimate | MEDIUM |
| NVIDIA RTX 3090 Ti 24 GB | ~326 tok/s | Calibrated estimate | MEDIUM |
| NVIDIA RTX 3090 24 GB | ~306 tok/s | Calibrated estimate | MEDIUM |
| NVIDIA RTX 4080 SUPER 16 GB | ~279 tok/s | Calibrated estimate | MEDIUM |
| NVIDIA RTX A6000 48 GB | ~274 tok/s | Calibrated estimate | MEDIUM |
| NVIDIA RTX 4080 16 GB | ~271 tok/s | Calibrated estimate | MEDIUM |
| NVIDIA RTX 3080 10 GB | ~265 tok/s | Calibrated estimate | MEDIUM |
Sources
Model specifications: official repository meta-llama/Llama-3.2-3B-Instruct · revision official-fallback-config. Compatibility: RigForAI calculator v1.0.0 (estimated VRAM). GPU products: Icecat. Amazon is a separate affiliate match and is not used as a spec source.