Home / What can I run
What AI models can this GPU run?
Choose a canonical GPU chip (for example RTX 3090 24 GB). The calculator returns actual catalog models that fit in usable VRAM at the selected quantization and context. This does not rank speed.
Assumptions used
- GPU
- NVIDIA Tesla V100S 32 GB · 32 GB advertised
- Usable VRAM
- 28.8 GB (advertised × 0.90)
- Quantization
- Q4
- Context
- 8K
71 of 82 catalog models calculate as a full VRAM fit under these assumptions.
Runs fully in GPU VRAM
71 models at the selected quantization and context. Memory fit only — not a speed ranking.
Code Llama 34B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 23.6 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 5.2 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Offload Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 4K Fits 8K Fits
DeepSeek Coder 33B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 23.8 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 5.0 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Offload Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 4K Fits 8K Fits
Qwen3 32B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 23.5 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 5.3 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Offload Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Limited context 4K Fits 40K Limited context 8K Fits
DeepSeek R1 Distill Qwen 32B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 23.3 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 5.5 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Offload Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 32K Limited context 4K Fits 64K Limited context 8K Fits
Qwen2.5 32B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 23.3 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 5.5 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Offload Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Limited context 4K Fits 8K Fits
Qwen2.5-Coder 32B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 23.3 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 5.5 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Offload Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Limited context 4K Fits 8K Fits
QwQ 32B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 23.3 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 5.5 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Offload Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Limited context 4K Fits 40K Limited context 8K Fits
Gemma 4 31B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 28.1 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 0.7 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Offload Q5 Limited context Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Limited context 2K Fits 256K Limited context 32K Limited context 4K Fits 64K Limited context 8K Fits
Qwen3 30B-A3B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 20.8 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 8.0 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Limited context Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits
Qwen3-Coder 30B-A3B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 20.8 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 8.0 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Limited context Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 256K Limited context 32K Fits 4K Fits 64K Fits 8K Fits
Gemma 3 27B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 22.0 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 6.8 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Limited context Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 32K Limited context 4K Fits 64K Limited context 8K Fits
Gemma 2 27B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 20.9 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 7.9 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 2K Fits 4K Fits 8K Fits
Gemma 4 26B-A4B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 19.4 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 9.4 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Offload Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 256K Limited context 32K Fits 4K Fits 64K Limited context 8K Fits
Mistral Small 3.2 24B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 17.2 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 11.6 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Limited context Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Mistral Small 24B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 17.0 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 11.8 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Limited context Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
Magistral Small
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 16.9 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 11.9 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Limited context Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits
GPT-OSS 20B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 14.7 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 14.1 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
StarCoder2 15B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 11.5 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 17.3 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 4K Fits 8K Fits
Qwen3 14B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 11.4 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 17.4 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits
DeepSeek R1 Distill Qwen 14B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 11.6 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 17.2 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Phi-4
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 11.6 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 17.2 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 4K Fits 8K Fits
Qwen2.5 14B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 11.6 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 17.2 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
Phi-4 Reasoning
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 11.6 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 17.2 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
Phi-3 Medium 128K Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 11.2 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 17.6 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Code Llama 13B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 15.2 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 13.6 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Offload BF16 Offload Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 4K Fits 8K Fits
Mistral Nemo 12B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 9.7 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 19.1 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Gemma 3 12B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 11.5 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 17.3 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 32K Fits 4K Fits 64K Limited context 8K Fits
Gemma 4 12B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 11.3 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 17.5 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 256K Limited context 32K Fits 4K Fits 64K Limited context 8K Fits
Falcon 3 10B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 8.5 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 20.3 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
GLM-4-9B-0414
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 7.0 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 21.8 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
GLM-4 9B Chat
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 11.7 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 17.1 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 32K Fits 4K Fits 64K Limited context 8K Fits
Gemma 2 9B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 9.2 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 19.6 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 2K Fits 4K Fits 8K Fits
Qwen3 8B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 7.1 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 21.7 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits
DeepSeek R1 Distill Llama 8B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 6.8 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 22.0 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Hermes 3 Llama 3.1 8B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 6.8 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 22.0 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Llama 3.1 8B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 6.8 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 22.0 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Llama 3 8B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 6.8 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 22.0 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 2K Fits 4K Fits 8K Fits
Ministral 8B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 7.0 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 21.8 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
Gemma 4 E4B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 6.5 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 22.3 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
DeepSeek R1 Distill Qwen 7B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 6.0 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 22.8 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Qwen2.5 7B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 6.0 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 22.8 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
Qwen2.5-Coder 7B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 6.0 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 22.8 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
Qwen2 7B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 6.0 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 22.8 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
Falcon 3 7B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 6.4 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 22.4 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
Mistral 7B Instruct v0.3
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 6.3 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 22.5 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
OpenChat 3.5 7B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 6.3 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 22.5 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 2K Fits 4K Fits 8K Fits
Falcon 7B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 9.8 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 19.0 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 32K Fits 4K Fits 64K Limited context 8K Fits
StarCoder2 7B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 5.8 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 23.0 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 4K Fits 8K Fits
Code Llama 7B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 9.0 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 19.8 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 4K Fits 8K Fits
Gemma 4 E2B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 4.3 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 24.5 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Gemma 3 4B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 4.5 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 24.3 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Qwen3 4B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 4.4 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 24.4 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits
Phi-4 Mini Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 4.2 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 24.6 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Phi-4 Mini Reasoning
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 4.2 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 24.6 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Phi-3.5 Mini Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 6.2 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 22.6 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Phi-3 Mini 128K Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 6.2 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 22.6 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Limited context 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Falcon 3 3B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 3.5 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 25.3 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
Llama 3.2 3B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 3.7 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 25.1 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Qwen2.5 3B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 3.0 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 25.8 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
SmolLM3 3B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 3.3 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 25.5 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Gemma 2 2B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 3.2 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 25.6 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 2K Fits 4K Fits 8K Fits
Qwen3 1.7B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 2.9 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 25.9 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits
DeepSeek R1 Distill Qwen 1.5B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 2.1 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 26.7 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
SmolLM2 1.7B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 3.3 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 25.5 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 2K Fits 4K Fits 8K Fits
Falcon 3 1B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 2.4 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 26.4 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 2K Fits 4K Fits 8K Fits
Qwen2.5 1.5B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 1.9 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 26.9 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
Qwen2.5-Coder 1.5B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 1.9 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 26.9 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
Llama 3.2 1B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 1.8 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 27.0 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 128K Fits 16K Fits 2K Fits 32K Fits 4K Fits 64K Fits 8K Fits
Gemma 3 1B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 1.6 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 27.2 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
Qwen3 0.6B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 2.1 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 26.7 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 40K Fits 8K Fits
Qwen2.5 0.5B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 1.2 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 27.6 GB leftover after the calculated requirement
- Fit
- Fits Runs fully in GPU VRAM
At 8K: FP16 Fits BF16 Fits Q8 Fits Q6 Fits Q5 Fits Q4 Fits Q3 Fits
Q4: 16K Fits 2K Fits 32K Fits 4K Fits 8K Fits
Requires CPU/RAM offload
8 models at the selected quantization and context. Memory fit only — not a speed ranking.
GPT-OSS 120B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 75.5 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 46.7 GB over usable VRAM
- Fit
- Offload Requires CPU/RAM offload
At 8K: FP16 No BF16 No Q8 No Q6 No Q5 No Q4 Offload Q3 Offload
Q4: 128K Offload 16K Offload 2K Offload 32K Offload 4K Offload 64K Offload 8K Offload
GLM-4.5 Air
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 72.2 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 43.4 GB over usable VRAM
- Fit
- Offload Requires CPU/RAM offload
At 8K: FP16 No BF16 No Q8 No Q6 No Q5 No Q4 Offload Q3 Offload
Q4: 128K Offload 16K Offload 2K Offload 32K Offload 4K Offload 64K Offload 8K Offload
Qwen2.5 72B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 49.3 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 20.5 GB over usable VRAM
- Fit
- Offload Requires CPU/RAM offload
At 8K: FP16 No BF16 No Q8 No Q6 Offload Q5 Offload Q4 Offload Q3 Offload
Q4: 16K Offload 2K Offload 32K Offload 4K Offload 8K Offload
DeepSeek R1 Distill Llama 70B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 48.0 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 19.2 GB over usable VRAM
- Fit
- Offload Requires CPU/RAM offload
At 8K: FP16 No BF16 No Q8 No Q6 Offload Q5 Offload Q4 Offload Q3 Offload
Q4: 128K Offload 16K Offload 2K Offload 32K Offload 4K Offload 64K Offload 8K Offload
Hermes 3 Llama 3.1 70B
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 48.0 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 19.2 GB over usable VRAM
- Fit
- Offload Requires CPU/RAM offload
At 8K: FP16 No BF16 No Q8 No Q6 Offload Q5 Offload Q4 Offload Q3 Offload
Q4: 128K Offload 16K Offload 2K Offload 32K Offload 4K Offload 64K Offload 8K Offload
Llama 3.1 70B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 48.0 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 19.2 GB over usable VRAM
- Fit
- Offload Requires CPU/RAM offload
At 8K: FP16 No BF16 No Q8 No Q6 Offload Q5 Offload Q4 Offload Q3 Offload
Q4: 128K Offload 16K Offload 2K Offload 32K Offload 4K Offload 64K Offload 8K Offload
Llama 3.3 70B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 48.0 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 19.2 GB over usable VRAM
- Fit
- Offload Requires CPU/RAM offload
At 8K: FP16 No BF16 No Q8 No Q6 Offload Q5 Offload Q4 Offload Q3 Offload
Q4: 128K Offload 16K Offload 2K Offload 32K Offload 4K Offload 64K Offload 8K Offload
Mixtral 8x7B Instruct
- Quantization
- Q4
- Context
- 8K
- Required VRAM
- 31.3 GB
- Usable GPU VRAM
- 28.8 GB
- VRAM headroom
- 2.5 GB over usable VRAM
- Fit
- Offload Requires CPU/RAM offload
At 8K: FP16 No BF16 No Q8 Offload Q6 Offload Q5 Offload Q4 Offload Q3 Fits
Q4: 16K Offload 2K Offload 32K Offload 4K Offload 8K Offload
Does not fit
| Model | Required VRAM | Usable VRAM | Status |
|---|---|---|---|
| GLM-4.5 | 230.7 GB | 28.8 GB | No |
| Qwen3 235B-A22B | 151.2 GB | 28.8 GB | No |
| Mixtral 8x22B Instruct | 91.8 GB | 28.8 GB | No |