Home / Models / Llama 3.1 70B Instruct
Llama 3.1 70B Instruct
71B parameters · native context 131,072 · llama
Estimated VRAM requirement
Weight memory uses parameter count × bits per weight, plus a quantization overhead factor. KV cache uses standard GQA formulas. Values are calculated estimates, not measured allocations.
| Quantization | Estimated weights | 2K total | 4K total | 8K total | 16K total | 32K total | 64K total | 128K total |
|---|---|---|---|---|---|---|---|---|
| FP16 | 134.1 GB | 146.2 GB | 146.9 GB | 148.1 GB | 150.6 GB | 155.6 GB | 165.6 GB | 185.6 GB |
| Q8 | 75.4 GB | 82.9 GB | 83.5 GB | 84.7 GB | 87.2 GB | 92.2 GB | 102.2 GB | 122.2 GB |
| Q6 | 58.8 GB | 64.8 GB | 65.5 GB | 66.7 GB | 69.2 GB | 74.2 GB | 84.2 GB | 104.2 GB |
| Q5 | 49.7 GB | 55.1 GB | 55.7 GB | 57.0 GB | 59.5 GB | 64.5 GB | 74.5 GB | 94.5 GB |
| Q4 | 41.4 GB | 46.1 GB | 46.7 GB | 48.0 GB | 50.5 GB | 55.5 GB | 65.5 GB | 85.5 GB |
| Q3 | 33.1 GB | 37.1 GB | 37.7 GB | 39.0 GB | 41.5 GB | 46.5 GB | 56.5 GB | 76.5 GB |
Find a GPU for this model
Run at the quantization and context selected below. Defaults are Q4 and 8K when you do not change the controls. Results are full-VRAM fits only — not a speed ranking.
- Run at
- Q4 · 8K context
- Required VRAM
- —
No published GPU calculates as a full VRAM fit for Q4 at 8K.
Compatible GPUs
Status is a calculated memory fit against usable VRAM (advertised × 0.9). Open a GPU to see product SKUs and Amazon listings.
Sources
Model specifications: official repository meta-llama/Llama-3.1-70B-Instruct · revision official-fallback-config. Compatibility: RigForAI calculator v1.0.0 (estimated VRAM). GPU products: Icecat. Amazon is a separate affiliate match and is not used as a spec source.