What AI Models Can You Run on a 24GB GPU?
A 24GB GPU can run most 7B–14B models comfortably and many 30B–34B models in 4-bit quantization, but context length and runtime overhead determine the real limit. This guide explains what fits, when you need offload, and how to check a specific model before downloading it.
A 24GB GPU is a strong local-AI configuration. It can run most 7B–14B language models with substantial headroom, and many 30B–34B models in 4-bit quantization. Some smaller models can run in 8-bit or FP16, while 70B-class models generally require CPU/RAM offload or multiple GPUs.
The important qualification is that model weights are only one part of the VRAM budget. You also need memory for the KV cache, temporary inference buffers, the model runtime, CUDA overhead, and sometimes the operating system or display.
As a practical starting point:
| Model size | FP16/BF16 | 8-bit | 4-bit | Typical 24GB-GPU result |
|---|---|---|---|---|
| 7B–8B | Usually comfortable | Comfortable | Comfortable | Good fit, including longer contexts |
| 12B–14B | Often too large for full precision | Usually comfortable | Comfortable | Good fit, with room depending on context |
| 20B–24B | Usually too large | May fit tightly | Usually comfortable | 4-bit is the practical choice |
| 30B–34B | Too large | Generally too large | Often fits, but headroom matters | Good candidate for 4-bit |
| 40B–ย? | Too large | Too large | Often tight or unsuitable | Check exact file and context requirements |
| 70B | Too large | Too large | Usually does not fit entirely | Offload or multiple GPUs normally required |
The table is a planning guide, not a compatibility guarantee. Quantization format, architecture, context length, batch size, and inference software can change the result significantly.
How much of a 24GB GPU is actually available?
A card marketed as having 24GB does not normally provide 24GB for model weights.
Some of the capacity may be consumed by:
- The desktop display and graphics driver
- CUDA or ROCm runtime allocations
- The inference framework
- Temporary computation buffers
- The KV cache for the conversation
- Multiple concurrent sequences or batches
For conservative planning, treat the usable budget as less than 24GB. The exact amount depends on the operating system, whether the GPU drives a monitor, the backend, and the model loader. A model that needs 23.8GB on paper may fail to load even though the card is labeled 24GB.
For a reliable setup, leave explicit headroom rather than targeting the theoretical maximum.
Estimating model-weight memory
Parameter count is useful, but quantization determines how much storage each parameter needs.
Approximate weight-memory formulas are:
- FP32:
parameters × 4 bytes - FP16 or BF16:
parameters × 2 bytes - 8-bit:
parameters × about 1 byte - 4-bit:
parameters × about 0.5 bytes
These are estimates. Real model files are larger because quantization includes scales and metadata, and some tensors may remain at higher precision.
A more practical formula is:
Total VRAM ≈ model-file memory + KV-cache memory + runtime overhead + safety headroom
For example, a 32B model at idealized 4-bit storage would require:
32 billion × 0.5 bytes ≈ 16GB
The actual loaded model can require more than that. Quantization metadata, unquantized tensors, temporary buffers, and the KV cache may push the total substantially higher. It may still be a reasonable 24GB candidate at moderate context lengths, but it should not be treated as a guaranteed fit from the parameter count alone.
Why the model file size is a better first check
For a specific quantized download, the file size is often more useful than multiplying parameter count by a theoretical bit width. It includes much of the quantization overhead already.
However, file size is still not the complete VRAM requirement. A loader may allocate additional memory during model conversion or loading, and runtime memory grows as the context grows.
A useful screening process is:
- Check the model's total parameter count.
- Identify the exact quantization and file format.
- Check the quantized file size.
- Reserve memory for the KV cache and runtime.
- Compare the result with the GPU's usable, rather than advertised, VRAM.
What fits at each quantization level?
7B–8B models
Models in the 7B–8B range are the easiest fit for a 24GB GPU.
They commonly run in:
- FP16 or BF16
- 8-bit quantization
- 4-bit quantization
Full-precision inference leaves more room for a large context and may preserve maximum model fidelity, but it consumes considerably more VRAM than 4-bit inference. For many users, 4-bit or 8-bit is chosen not because FP16 is impossible, but because the saved memory can be used for longer context, larger batches, or multiple concurrent requests.
12B–14B models
This is a particularly comfortable range for 24GB cards when using 4-bit or 8-bit quantization.
FP16 memory alone is approximately:
14 billion × 2 bytes ≈ 28GB
That already exceeds the card's capacity before runtime overhead, so a 14B model generally needs quantization or offload for a 24GB GPU.
At 4-bit, the idealized weight estimate is:
14 billion × 0.5 bytes ≈ 7GB
The loaded model will be larger than this estimate, but usually leaves useful room for the KV cache and inference runtime.
20B–24B models
In this range, 4-bit quantization is usually the practical starting point.
An idealized 24B 4-bit estimate is:
24 billion × 0.5 bytes ≈ 12GB
That does not mean every 24B model will use only 12GB in practice. The exact architecture, quantization scheme, runtime, and context length still matter. 8-bit versions may fit in some configurations but can leave little room for context or may require a model loader with efficient memory handling.
30B–34B models
Many 30B–34B models are realistic targets for a 24GB GPU in a well-chosen 4-bit quantization.
An idealized 32B estimate is:
32 billion × 0.5 bytes ≈ 16GB
The remaining capacity may be enough for inference overhead and a moderate context, but this is where headroom becomes important. A model with a large quantized file, a high context setting, or a memory-hungry backend may not fit fully in VRAM.
8-bit and FP16 versions in this range are generally too large for a single 24GB card without offload.
40B-class models
40B-class models are borderline even in 4-bit quantization. The idealized weight estimate is already about 20GB:
40 billion × 0.5 bytes ≈ 20GB
Real memory use can exceed the remaining capacity once the runtime and KV cache are included. Some carefully selected files may be possible with a short context or partial offload, but this is not the range to assume will work smoothly on every 24GB setup.
70B-class models
A 70B model in idealized 4-bit storage requires:
70 billion × 0.5 bytes ≈ 35GB
That is above the capacity of a 24GB GPU before any runtime memory is considered. A 70B model can sometimes be made to run with CPU/RAM offload, but the GPU will not hold all weights, and generation speed may be limited by transfers between system memory and VRAM.
If your goal is a fully GPU-resident 70B model, plan for more VRAM, such as multiple GPUs or a higher-memory accelerator.
Context length can determine whether a model fits
Loading the weights is only the first VRAM test. During generation, the model stores attention keys and values for the current prompt and conversation. This storage is called the KV cache.
KV-cache usage grows with:
- Context length
- Number of layers
- Number of KV heads
- Head dimension
- Bytes per cache element
- Batch size or number of simultaneous sequences
A simplified estimate for one sequence is:
KV cache bytes ≈ 2 × layers × context tokens × KV heads × head dimension × bytes per element
The factor of 2 represents keys and values.
This formula is useful for understanding scaling, but it is not a universal calculator. Some runtimes use different cache types, cache quantization, paging, or memory layouts. Architectures with grouped-query attention or multi-query attention can use much less KV memory than an otherwise similar model.
The practical consequence is straightforward:
- A model that fits at 4,096 tokens may not fit at 32,000 or 128,000 tokens.
- Increasing context length can reduce generation speed even if the model still loads.
- A long context setting may reserve memory before you use the full context.
- Batch size multiplies KV-cache requirements across simultaneous sequences.
If a 30B–34B 4-bit model is close to the VRAM limit, start with a moderate context setting and increase it gradually. Do not assume that the model's advertised maximum context is practical on a 24GB card.
Headroom matters more than the exact parameter count
Two models with the same parameter count can have different VRAM requirements because of:
- Different quantization formats
- Different numbers of layers and attention heads
- Different KV-head configurations
- Different vocabulary or embedding sizes
- Unquantized output or embedding layers
- Runtime-specific temporary buffers
- Context and batch settings
For this reason, avoid choosing a model solely because:
parameter count × 0.5 bytes < 24GB
That calculation is a first-pass filter, not a loading guarantee.
A better rule is to divide candidates into three groups:
Comfortable fit: The model loads with meaningful room for context and normal runtime overhead.
Tight fit: The model may load, but context length, batch size, or backend settings need careful control.
Offload candidate: The model exceeds the practical GPU budget or leaves too little room for useful inference.
When CPU or system-RAM offload becomes necessary
Offload places some model weights or runtime data in system RAM instead of VRAM. It can make a model run when it would otherwise fail to load, but it changes the performance characteristics.
Offload is commonly needed when:
- The quantized model weights alone approach or exceed usable VRAM.
- You want a larger model than the GPU can hold.
- You need a context length that exhausts VRAM.
- You are using a higher-precision version for quality or compatibility.
- You are running multiple models or simultaneous requests.
The trade-off is data movement. GPU memory has much higher bandwidth and lower access latency than ordinary system memory, while CPU-to-GPU transfers are constrained by the platform's interconnect and workload pattern. A partially offloaded model may work, but tokens can be generated much more slowly than with a fully GPU-resident model.
Offload is most attractive when:
- You value model capability more than maximum tokens per second.
- The amount offloaded is small.
- You have sufficient system RAM.
- You can tolerate slower prompt processing or generation.
- The application supports configurable layer placement.
It is less attractive when you expect interactive speed, long contexts, or multiple concurrent users. In those cases, a smaller fully GPU-resident model may provide a better experience.
System RAM also needs room for the operating system, the model file, the loader, and any other applications. Do not plan to use all installed RAM for offload.
Mixture-of-experts models need special care
A mixture-of-experts, or MoE, model may advertise a relatively small number of active parameters per token while having a much larger total parameter count.
The active-parameter figure describes how much computation is used for each token. It does not automatically describe how much memory is required to store the model. Depending on the architecture and inference engine, many or all expert weights may still need to be available in memory.
When evaluating an MoE model, check:
- Total parameters
- Active parameters
- Quantized file size
- Whether experts can be dynamically loaded or offloaded
- The runtime's documented memory behavior
Do not assume that an MoE labeled “12B active” has the same memory requirements as a dense 12B model.
What about image and other AI models?
“24GB GPU” compatibility is not limited to language models, but the memory calculation differs by workload.
For image generation, VRAM use depends on:
- The base model or checkpoint
- Precision and quantization
- Image resolution
- Batch size
- Control modules, adapters, and additional encoders
- The user interface and attention implementation
A setup that handles a language model comfortably may still need memory optimization for high-resolution image generation or multiple additional components. The same principle applies to video, speech, embedding, and multimodal models: model files are only part of the working-set memory.
For a specific non-LLM workflow, check the application's documented requirements and test the intended resolution, batch size, and add-ons rather than relying on parameter count alone.
A practical selection workflow
Use this process before downloading a large model:
- Define the workload. Decide whether you need chat, coding, reasoning, multimodal input, image generation, or another task.
- Choose a parameter range. On 24GB, start with 7B–14B for flexibility, or consider 30B–34B when 4-bit quality and model capability are more important.
- Select the quantization. Use FP16/BF16 when the model is small enough and you want maximum precision; use 8-bit or 4-bit to increase the feasible model size.
- Inspect the exact file size. Do not estimate from the model name alone.
- Budget the context. Decide whether you need a short chat context or a long document context.
- Leave safety margin. Keep room for the runtime, KV cache, and normal system allocations.
- Check backend compatibility. The model format must be supported by the inference application and GPU backend you plan to use.
- Plan for offload only intentionally. If the model exceeds the GPU budget, verify that you have enough system RAM and accept the likely speed trade-off.
For a model-specific answer, use RigForAI's What can I run tool. It is more useful than a generic parameter table when you want to compare a particular model, quantization, context length, and hardware configuration.
Bottom line
A 24GB GPU is best viewed as a 7B–14B high-flexibility machine and a 30B–34B 4-bit machine, with the exact boundary determined by context length and runtime overhead.
- 7B–8B models: usually comfortable across common precisions.
- 12B–14B models: excellent in 4-bit or 8-bit; usually not full FP16.
- 20B–24B models: generally practical in 4-bit.
- 30B–34B models: often practical in 4-bit, but check headroom carefully.
- 40B-class models: borderline and configuration-dependent.
- 70B models: normally require offload or additional GPUs.
The safest choice is not the largest model whose estimated weights barely fit. It is the largest model that leaves enough VRAM for the context, runtime, and workload you actually intend to use.