GLM-4.5
358B parameters · 32B active · native context 131,072 · glm4_moe
Estimated VRAM requirement
Weight memory uses parameter count × bits per weight, plus a quantization overhead factor. KV cache uses standard GQA formulas. Values are calculated estimates, not measured allocations.
| Quantization | Estimated weights | 2K total | 4K total | 8K total | 16K total | 32K total | 64K total | 128K total |
|---|---|---|---|---|---|---|---|---|
| FP16 | 680.8 GB | 736.7 GB | 737.5 GB | 738.9 GB | 741.8 GB | 747.5 GB | 759.0 GB | 782.0 GB |
| Q8 | 383.0 GB | 415.1 GB | 415.8 GB | 417.2 GB | 420.1 GB | 425.8 GB | 437.3 GB | 460.3 GB |
| Q6 | 298.3 GB | 323.6 GB | 324.3 GB | 325.8 GB | 328.6 GB | 334.4 GB | 345.9 GB | 368.9 GB |
| Q5 | 252.4 GB | 274.0 GB | 274.8 GB | 276.2 GB | 279.1 GB | 284.8 GB | 296.3 GB | 319.3 GB |
| Q4 | 210.2 GB | 228.5 GB | 229.3 GB | 230.7 GB | 233.6 GB | 239.3 GB | 250.8 GB | 273.8 GB |
| Q3 | 167.9 GB | 182.8 GB | 183.5 GB | 185.0 GB | 187.8 GB | 193.6 GB | 205.1 GB | 228.1 GB |
Find a GPU for this model
Run at the quantization and context selected below. Defaults are Q4 and 8K when you do not change the controls. Results are full-VRAM fits only — not a speed ranking.
- Run at
- Q4 · 8K context
- Required VRAM
- —
No fresh matched Amazon price is currently available for a full-VRAM fit. See compatible GPUs · Workstations · Rent a server · Buy vs rent.
No published GPU calculates as a full VRAM fit for Q4 at 8K. Explore multi-GPU options · Build a PC · Buy vs rent.
Compatible GPUs
Status is a calculated memory fit against usable VRAM (advertised × 0.9). Open a GPU to see product SKUs and Amazon listings.
Sources
Model specifications: official repository zai-org/GLM-4.5 · revision main. Compatibility: RigForAI calculator v1.0.0 (estimated VRAM). GPU products: Icecat. Amazon is a separate affiliate match and is not used as a spec source.