Home / What can I run

Deterministic VRAM fit · reverse of Find a GPU

What AI models can this GPU run?

Choose a canonical GPU chip (for example RTX 3090 24 GB). The calculator returns actual catalog models that fit in usable VRAM at the selected quantization and context. This does not rank speed.

I know my model — find a GPU

Assumptions used

GPU
NVIDIA RTX PRO 6000 Blackwell 96 GB · 96 GB advertised
Usable VRAM
86.4 GB (advertised × 0.90)
Quantization
Q4
Context
8K

79 of 82 catalog models calculate as a full VRAM fit under these assumptions.

Runs fully in GPU VRAM

79 models at the selected quantization and context. Memory fit only — not a speed ranking.

OpenAI · 117B

GPT-OSS 120B

Quantization
Q4
Context
8K
Required VRAM
75.5 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
10.9 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 No BF16 No Q8 Offload Q6 Offload Q5 Offload Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Zhipu AI · 110B

GLM-4.5 Air

Quantization
Q4
Context
8K
Required VRAM
72.2 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
14.2 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Offload Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Limited context 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 73B

Qwen2.5 72B Instruct

Quantization
Q4
Context
8K
Required VRAM
49.3 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
37.1 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Offload BF16 Offload Q8 Limited context Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

DeepSeek · 71B

DeepSeek R1 Distill Llama 70B

Quantization
Q4
Context
8K
Required VRAM
48.0 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
38.4 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Offload BF16 Offload Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

NousResearch · 71B

Hermes 3 Llama 3.1 70B

Quantization
Q4
Context
8K
Required VRAM
48.0 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
38.4 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Offload BF16 Offload Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Meta · 71B

Llama 3.1 70B Instruct

Quantization
Q4
Context
8K
Required VRAM
48.0 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
38.4 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Offload BF16 Offload Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Meta · 71B

Llama 3.3 70B Instruct

Quantization
Q4
Context
8K
Required VRAM
48.0 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
38.4 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Offload BF16 Offload Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Mistral · 47B

Mixtral 8x7B Instruct

Quantization
Q4
Context
8K
Required VRAM
31.3 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
55.1 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Offload BF16 Offload Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Meta · 34B

Code Llama 34B Instruct

Quantization
Q4
Context
8K
Required VRAM
23.6 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
62.8 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 4K Fits 8K Fits

Full calculationFind a GPU

DeepSeek · 33B

DeepSeek Coder 33B Instruct

Quantization
Q4
Context
8K
Required VRAM
23.8 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
62.6 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 33B

Qwen3 32B

Quantization
Q4
Context
8K
Required VRAM
23.5 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
62.9 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits

Full calculationFind a GPU

DeepSeek · 33B

DeepSeek R1 Distill Qwen 32B

Quantization
Q4
Context
8K
Required VRAM
23.3 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
63.1 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 33B

Qwen2.5 32B Instruct

Quantization
Q4
Context
8K
Required VRAM
23.3 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
63.1 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 33B

Qwen2.5-Coder 32B Instruct

Quantization
Q4
Context
8K
Required VRAM
23.3 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
63.1 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 33B

QwQ 32B

Quantization
Q4
Context
8K
Required VRAM
23.3 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
63.1 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits

Full calculationFind a GPU

Google · 31B

Gemma 4 31B Instruct

Quantization
Q4
Context
8K
Required VRAM
28.1 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
58.3 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Limited context 16K Fits 2K Fits 256K Limited context 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 31B

Qwen3 30B-A3B

Quantization
Q4
Context
8K
Required VRAM
20.8 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
65.6 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 31B

Qwen3-Coder 30B-A3B Instruct

Quantization
Q4
Context
8K
Required VRAM
20.8 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
65.6 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 256K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Google · 27B

Gemma 3 27B Instruct

Quantization
Q4
Context
8K
Required VRAM
22.0 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
64.4 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Google · 27B

Gemma 2 27B Instruct

Quantization
Q4
Context
8K
Required VRAM
20.9 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
65.5 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 2K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Google · 27B

Gemma 4 26B-A4B Instruct

Quantization
Q4
Context
8K
Required VRAM
19.4 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
67.0 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 256K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Mistral · 24B

Mistral Small 3.2 24B Instruct

Quantization
Q4
Context
8K
Required VRAM
17.2 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
69.2 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Mistral · 24B

Mistral Small 24B Instruct

Quantization
Q4
Context
8K
Required VRAM
17.0 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
69.4 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Mistral · 24B

Magistral Small

Quantization
Q4
Context
8K
Required VRAM
16.9 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
69.5 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits

Full calculationFind a GPU

OpenAI · 22B

GPT-OSS 20B

Quantization
Q4
Context
8K
Required VRAM
14.7 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
71.7 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

BigCode · 16B

StarCoder2 15B

Quantization
Q4
Context
8K
Required VRAM
11.5 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
74.9 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 15B

Qwen3 14B

Quantization
Q4
Context
8K
Required VRAM
11.4 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
75.0 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits

Full calculationFind a GPU

DeepSeek · 15B

DeepSeek R1 Distill Qwen 14B

Quantization
Q4
Context
8K
Required VRAM
11.6 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
74.8 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Microsoft · 15B

Phi-4

Quantization
Q4
Context
8K
Required VRAM
11.6 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
74.8 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 15B

Qwen2.5 14B Instruct

Quantization
Q4
Context
8K
Required VRAM
11.6 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
74.8 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Microsoft · 15B

Phi-4 Reasoning

Quantization
Q4
Context
8K
Required VRAM
11.6 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
74.8 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Microsoft · 14B

Phi-3 Medium 128K Instruct

Quantization
Q4
Context
8K
Required VRAM
11.2 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
75.2 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Meta · 13B

Code Llama 13B Instruct

Quantization
Q4
Context
8K
Required VRAM
15.2 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
71.2 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Mistral · 12B

Mistral Nemo 12B Instruct

Quantization
Q4
Context
8K
Required VRAM
9.7 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
76.7 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Google · 12B

Gemma 3 12B Instruct

Quantization
Q4
Context
8K
Required VRAM
11.5 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
74.9 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Google · 12B

Gemma 4 12B Instruct

Quantization
Q4
Context
8K
Required VRAM
11.3 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
75.1 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 256K Limited context 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

TII · 10B

Falcon 3 10B Instruct

Quantization
Q4
Context
8K
Required VRAM
8.5 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
77.9 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Zhipu AI · 9.4B

GLM-4-9B-0414

Quantization
Q4
Context
8K
Required VRAM
7.0 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
79.4 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Zhipu AI · 9.4B

GLM-4 9B Chat

Quantization
Q4
Context
8K
Required VRAM
11.7 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
74.7 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Limited context 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Google · 9.2B

Gemma 2 9B Instruct

Quantization
Q4
Context
8K
Required VRAM
9.2 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
77.2 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 2K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 8.2B

Qwen3 8B

Quantization
Q4
Context
8K
Required VRAM
7.1 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
79.3 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits

Full calculationFind a GPU

DeepSeek · 8.0B

DeepSeek R1 Distill Llama 8B

Quantization
Q4
Context
8K
Required VRAM
6.8 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
79.6 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

NousResearch · 8.0B

Hermes 3 Llama 3.1 8B

Quantization
Q4
Context
8K
Required VRAM
6.8 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
79.6 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Meta · 8.0B

Llama 3.1 8B Instruct

Quantization
Q4
Context
8K
Required VRAM
6.8 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
79.6 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Meta · 8.0B

Llama 3 8B Instruct

Quantization
Q4
Context
8K
Required VRAM
6.8 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
79.6 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 2K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Mistral · 8.0B

Ministral 8B Instruct

Quantization
Q4
Context
8K
Required VRAM
7.0 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
79.4 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Google · 8.0B

Gemma 4 E4B Instruct

Quantization
Q4
Context
8K
Required VRAM
6.5 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
79.9 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

DeepSeek · 7.6B

DeepSeek R1 Distill Qwen 7B

Quantization
Q4
Context
8K
Required VRAM
6.0 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
80.4 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 7.6B

Qwen2.5 7B Instruct

Quantization
Q4
Context
8K
Required VRAM
6.0 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
80.4 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 7.6B

Qwen2.5-Coder 7B Instruct

Quantization
Q4
Context
8K
Required VRAM
6.0 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
80.4 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Qwen · 7.6B

Qwen2 7B Instruct

Quantization
Q4
Context
8K
Required VRAM
6.0 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
80.4 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

TII · 7.5B

Falcon 3 7B Instruct

Quantization
Q4
Context
8K
Required VRAM
6.4 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
80.0 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Mistral · 7.3B

Mistral 7B Instruct v0.3

Quantization
Q4
Context
8K
Required VRAM
6.3 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
80.1 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

OpenChat · 7.2B

OpenChat 3.5 7B

Quantization
Q4
Context
8K
Required VRAM
6.3 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
80.1 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 2K Fits 4K Fits 8K Fits

Full calculationFind a GPU

TII · 7.2B

Falcon 7B Instruct

Quantization
Q4
Context
8K
Required VRAM
9.8 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
76.6 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

BigCode · 7.2B

StarCoder2 7B

Quantization
Q4
Context
8K
Required VRAM
5.8 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
80.6 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Meta · 6.7B

Code Llama 7B Instruct

Quantization
Q4
Context
8K
Required VRAM
9.0 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
77.4 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Google · 5.1B

Gemma 4 E2B Instruct

Quantization
Q4
Context
8K
Required VRAM
4.3 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
82.1 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Google · 4.3B

Gemma 3 4B Instruct

Quantization
Q4
Context
8K
Required VRAM
4.5 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
81.9 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 4.0B

Qwen3 4B

Quantization
Q4
Context
8K
Required VRAM
4.4 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
82.0 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits

Full calculationFind a GPU

Microsoft · 3.8B

Phi-4 Mini Instruct

Quantization
Q4
Context
8K
Required VRAM
4.2 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
82.2 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Microsoft · 3.8B

Phi-4 Mini Reasoning

Quantization
Q4
Context
8K
Required VRAM
4.2 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
82.2 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Microsoft · 3.8B

Phi-3.5 Mini Instruct

Quantization
Q4
Context
8K
Required VRAM
6.2 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
80.2 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Microsoft · 3.8B

Phi-3 Mini 128K Instruct

Quantization
Q4
Context
8K
Required VRAM
6.2 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
80.2 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

TII · 3.2B

Falcon 3 3B Instruct

Quantization
Q4
Context
8K
Required VRAM
3.5 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
82.9 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Meta · 3.2B

Llama 3.2 3B Instruct

Quantization
Q4
Context
8K
Required VRAM
3.7 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
82.7 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 3.1B

Qwen2.5 3B Instruct

Quantization
Q4
Context
8K
Required VRAM
3.0 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
83.4 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

HuggingFaceTB · 3.1B

SmolLM3 3B

Quantization
Q4
Context
8K
Required VRAM
3.3 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
83.1 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Google · 2.6B

Gemma 2 2B Instruct

Quantization
Q4
Context
8K
Required VRAM
3.2 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
83.2 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 2K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 2.0B

Qwen3 1.7B

Quantization
Q4
Context
8K
Required VRAM
2.9 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
83.5 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits

Full calculationFind a GPU

DeepSeek · 1.8B

DeepSeek R1 Distill Qwen 1.5B

Quantization
Q4
Context
8K
Required VRAM
2.1 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
84.3 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

HuggingFaceTB · 1.7B

SmolLM2 1.7B Instruct

Quantization
Q4
Context
8K
Required VRAM
3.3 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
83.1 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 2K Fits 4K Fits 8K Fits

Full calculationFind a GPU

TII · 1.7B

Falcon 3 1B Instruct

Quantization
Q4
Context
8K
Required VRAM
2.4 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
84.0 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 2K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Qwen · 1.5B

Qwen2.5 1.5B Instruct

Quantization
Q4
Context
8K
Required VRAM
1.9 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
84.5 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Qwen · 1.5B

Qwen2.5-Coder 1.5B Instruct

Quantization
Q4
Context
8K
Required VRAM
1.9 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
84.5 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Meta · 1.2B

Llama 3.2 1B Instruct

Quantization
Q4
Context
8K
Required VRAM
1.8 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
84.6 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits

Full calculationFind a GPU

Google · 1000M

Gemma 3 1B Instruct

Quantization
Q4
Context
8K
Required VRAM
1.6 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
84.8 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Alibaba · 752M

Qwen3 0.6B

Quantization
Q4
Context
8K
Required VRAM
2.1 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
84.3 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits

Full calculationFind a GPU

Qwen · 494M

Qwen2.5 0.5B Instruct

Quantization
Q4
Context
8K
Required VRAM
1.2 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
85.2 GB leftover after the calculated requirement
Fit
Fits Runs fully in GPU VRAM

At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits

Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits

Full calculationFind a GPU

Requires CPU/RAM offload

3 models at the selected quantization and context. Memory fit only — not a speed ranking.

Zhipu AI · 358B

GLM-4.5

Quantization
Q4
Context
8K
Required VRAM
230.7 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
144.3 GB over usable VRAM
Fit
Offload Requires CPU/RAM offload

At 8K: FP16 No BF16 No Q8 No Q6 No Q5 No Q4 Offload Q3 Offload

Q4: 128K Offload 16K Offload 2K Offload 32K Offload 4K Offload 64K Offload 8K Offload

Full calculationFind a GPU

Alibaba · 235B

Qwen3 235B-A22B

Quantization
Q4
Context
8K
Required VRAM
151.2 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
64.8 GB over usable VRAM
Fit
Offload Requires CPU/RAM offload

At 8K: FP16 No BF16 No Q8 No Q6 Offload Q5 Offload Q4 Offload Q3 Offload

Q4: 16K Offload 2K Offload 32K Offload 4K Offload 40K Offload 8K Offload

Full calculationFind a GPU

Mistral · 141B

Mixtral 8x22B Instruct

Quantization
Q4
Context
8K
Required VRAM
91.8 GB
Usable GPU VRAM
86.4 GB
VRAM headroom
5.4 GB over usable VRAM
Fit
Offload Requires CPU/RAM offload

At 8K: FP16 No BF16 No Q8 Offload Q6 Offload Q5 Offload Q4 Offload Q3 Fits

Q4: 16K Offload 2K Offload 32K Offload 4K Offload 64K Offload 8K Offload

Full calculationFind a GPU