GPT-OSS 120B
117B parameters · native context 131,072 · gpt_oss
Estimated VRAM requirement
Weight memory uses parameter count × bits per weight, plus a quantization overhead factor. KV cache uses standard GQA formulas. Values are calculated estimates, not measured allocations.
| Quantization | Estimated weights | 2K total | 4K total | 8K total | 16K total | 32K total | 64K total | 128K total |
|---|---|---|---|---|---|---|---|---|
| FP16 | 222.3 GB | 241.0 GB | 241.1 GB | 241.4 GB | 241.9 GB | 243.1 GB | 245.3 GB | 249.8 GB |
| Q8 | 125.0 GB | 135.9 GB | 136.1 GB | 136.4 GB | 136.9 GB | 138.0 GB | 140.3 GB | 144.8 GB |
| Q6 | 97.4 GB | 106.1 GB | 106.2 GB | 106.5 GB | 107.1 GB | 108.2 GB | 110.4 GB | 114.9 GB |
| Q5 | 82.4 GB | 89.9 GB | 90.0 GB | 90.3 GB | 90.9 GB | 92.0 GB | 94.2 GB | 98.7 GB |
| Q4 | 68.6 GB | 75.0 GB | 75.2 GB | 75.5 GB | 76.0 GB | 77.1 GB | 79.4 GB | 83.9 GB |
| Q3 | 54.8 GB | 60.1 GB | 60.2 GB | 60.5 GB | 61.1 GB | 62.2 GB | 64.5 GB | 69.0 GB |
Find a GPU for this model
Run at the quantization and context selected below. Defaults are Q4 and 8K when you do not change the controls. Results are full-VRAM fits only — not a speed ranking.
- Run at
- Q4 · 8K context
- Required VRAM
- —
No published GPU calculates as a full VRAM fit for Q4 at 8K.
Compatible GPUs
Status is a calculated memory fit against usable VRAM (advertised × 0.9). Open a GPU to see product SKUs and Amazon listings.
Sources
Model specifications: official repository openai/gpt-oss-120b · revision main. Compatibility: RigForAI calculator v1.0.0 (estimated VRAM). GPU products: Icecat. Amazon is a separate affiliate match and is not used as a spec source.