Home / AI GPUs / NVIDIA H100 NVL 94 GB
NVIDIA H100 NVL 94 GB
94 GB advertised VRAM. Usable memory for calculations is 90% of advertised capacity. This page is the chip + VRAM entity — not a single cooler SKU.
No fresh matched Amazon price is currently available for this canonical GPU.
What can I runCompare this GPUValue explorerFind a GPU for a model
What AI models can it run?
Calculated memory fit for catalog models on this chip. Defaults are Q4 and 8K when you do not change the controls. This is not a speed ranking.
- GPU
- NVIDIA H100 NVL 94 GB · 94 GB advertised
- Usable VRAM
- 84.6 GB (advertised × 0.90)
- Quantization
- Q4 (default for this form)
- Context
- 8K (default for this form)
79 of 82 catalog models calculate as a full VRAM fit under these assumptions.
Runs fully in GPU VRAM
79 models at the selected quantization and context. Memory fit only — not a speed ranking.
GPT-OSS 120B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 75.5 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 9.1 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 No BF16 No Q8 Offload Q6 Offload Q5 Offload Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
GLM-4.5 Air
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 72.2 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 12.4 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Offload Q5 Offload Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Qwen2.5 72B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 49.3 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 35.3 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
DeepSeek R1 Distill Llama 70B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 48.0 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 36.6 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Limited context Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Hermes 3 Llama 3.1 70B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 48.0 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 36.6 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Limited context Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Llama 3.1 70B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 48.0 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 36.6 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Limited context Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Llama 3.3 70B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 48.0 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 36.6 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Limited context Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Mixtral 8x7B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 31.3 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 53.3 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
Code Llama 34B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 23.6 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 61.0 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 4K Fits 8K Fits
DeepSeek Coder 33B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 23.8 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 60.8 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 4K Fits 8K Fits
Qwen3 32B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 23.5 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 61.1 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits
DeepSeek R1 Distill Qwen 32B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 23.3 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 61.3 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Qwen2.5 32B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 23.3 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 61.3 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
Qwen2.5-Coder 32B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 23.3 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 61.3 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
QwQ 32B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 23.3 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 61.3 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits
Gemma 4 31B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 28.1 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 56.5 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 256K Limited context 32K Fits 4K Fits 64K Fits 8K Fits
Qwen3 30B-A3B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 20.8 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 63.8 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits
Qwen3-Coder 30B-A3B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 20.8 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 63.8 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 256K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Gemma 3 27B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 22.0 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 62.6 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Gemma 2 27B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 20.9 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 63.7 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 2K Fits 4K Fits 8K Fits
Requires CPU/RAM offload
3 models at the selected quantization and context. Memory fit only — not a speed ranking.
GLM-4.5
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 230.7 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 146.1 GB over usable VRAM
- Fit
- Offload Requires CPU/RAM offload
At 8K: FP16 No BF16 No Q8 No Q6 No Q5 No Q4 Offload Q3 Offload
Q4: 128K Offload 16K Offload 2K Offload 32K Offload 4K Offload 64K Offload 8K Offload
Qwen3 235B-A22B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 151.2 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 66.6 GB over usable VRAM
- Fit
- Offload Requires CPU/RAM offload
At 8K: FP16 No BF16 No Q8 No Q6 Offload Q5 Offload Q4 Offload Q3 Offload
Q4: 16K Offload 2K Offload 32K Offload 4K Offload 40K Offload 8K Offload
Mixtral 8x22B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 91.8 GB
- Usable GPU VRAM
- 84.6 GB
- VRAM headroom
- 7.2 GB over usable VRAM
- Fit
- Offload Requires CPU/RAM offload
At 8K: FP16 No BF16 No Q8 Offload Q6 Offload Q5 Offload Q4 Offload Q3 Fits
Q4: 16K Offload 2K Offload 32K Offload 4K Offload 64K Offload 8K Offload
Q4 / Q8 / FP16 overview
Status below is the best calculated result across evaluated context lengths for each quantization. Open a model for the full table.
| Model | Params | Q4 | Q8 | FP16 | Max context in VRAM |
|---|
Available graphics cards
These are Icecat product SKUs mapped to this chip. Different board-partner cards are not assumed to share a price. Amazon CTAs appear only for EXACT/HIGH matches that are not bundled accessories.
No mapped product SKUs yet.
Sources
GPU identity and VRAM: Icecat structured specifications. Compatibility: RigForAI calculated estimate. Amazon: live affiliate match on individual SKUs, via /go/amazon/.