Home / Models / Qwen2.5 72B Instruct
Qwen2.5 72B Instruct
73B parameters · native context 32,768 · qwen2
Estimated VRAM requirement
Weight memory uses parameter count × bits per weight, plus a quantization overhead factor. KV cache uses standard GQA formulas. Values are calculated estimates, not measured allocations.
| Quantization | Estimated weights | 2K total | 4K total | 8K total | 16K total | 32K total |
|---|---|---|---|---|---|---|
| FP16 | 138.1 GB | 150.5 GB | 151.2 GB | 152.4 GB | 154.9 GB | 159.9 GB |
| Q8 | 77.7 GB | 85.3 GB | 85.9 GB | 87.2 GB | 89.7 GB | 94.7 GB |
| Q6 | 60.5 GB | 66.7 GB | 67.4 GB | 68.6 GB | 71.1 GB | 76.1 GB |
| Q5 | 51.2 GB | 56.7 GB | 57.3 GB | 58.5 GB | 61.0 GB | 66.0 GB |
| Q4 | 42.7 GB | 47.4 GB | 48.1 GB | 49.3 GB | 51.8 GB | 56.8 GB |
| Q3 | 34.1 GB | 38.2 GB | 38.8 GB | 40.0 GB | 42.5 GB | 47.5 GB |
Find a GPU for this model
Run at the quantization and context selected below. Defaults are Q4 and 8K when you do not change the controls. Results are full-VRAM fits only — not a speed ranking.
- Run at
- Q4 · 8K context
- Required VRAM
- —
No published GPU calculates as a full VRAM fit for Q4 at 8K.
Compatible GPUs
Status is a calculated memory fit against usable VRAM (advertised × 0.9). Open a GPU to see product SKUs and Amazon listings.
Sources
Model specifications: official repository Qwen/Qwen2.5-72B-Instruct · revision main. Compatibility: RigForAI calculator v1.0.0 (estimated VRAM). GPU products: Icecat. Amazon is a separate affiliate match and is not used as a spec source.