Home / Guides / What AI Models Can You Run on a 24GB GPU?

guide

What AI Models Can You Run on a 24GB GPU?

Updated 2026-08-21

A 24GB GPU can run most 7B–14B models comfortably and many 30B–34B models in 4-bit quantization, but context length and runtime overhead determine the real limit. This guide explains what fits, when you need offload, and how to check a specific model before downloading it.

A 24GB GPU is a strong local-AI configuration. It can run most 7B–14B language models with substantial headroom, and many 30B–34B models in 4-bit quantization. Some smaller models can run in 8-bit or FP16, while 70B-class models generally require CPU/RAM offload or multiple GPUs.

The important qualification is that model weights are only one part of the VRAM budget. You also need memory for the KV cache, temporary inference buffers, the model runtime, CUDA overhead, and sometimes the operating system or display.

As a practical starting point:

Model sizeFP16/BF168-bit4-bitTypical 24GB-GPU result
7B–8BUsually comfortableComfortableComfortableGood fit, including longer contexts
12B–14BOften too large for full precisionUsually comfortableComfortableGood fit, with room depending on context
20B–24BUsually too largeMay fit tightlyUsually comfortable4-bit is the practical choice
30B–34BToo largeGenerally too largeOften fits, but headroom mattersGood candidate for 4-bit
40B–ย?Too largeToo largeOften tight or unsuitableCheck exact file and context requirements
70BToo largeToo largeUsually does not fit entirelyOffload or multiple GPUs normally required

The table is a planning guide, not a compatibility guarantee. Quantization format, architecture, context length, batch size, and inference software can change the result significantly.

How much of a 24GB GPU is actually available?

A card marketed as having 24GB does not normally provide 24GB for model weights.

Some of the capacity may be consumed by:

  • The desktop display and graphics driver
  • CUDA or ROCm runtime allocations
  • The inference framework
  • Temporary computation buffers
  • The KV cache for the conversation
  • Multiple concurrent sequences or batches

For conservative planning, treat the usable budget as less than 24GB. The exact amount depends on the operating system, whether the GPU drives a monitor, the backend, and the model loader. A model that needs 23.8GB on paper may fail to load even though the card is labeled 24GB.

For a reliable setup, leave explicit headroom rather than targeting the theoretical maximum.

Estimating model-weight memory

Parameter count is useful, but quantization determines how much storage each parameter needs.

Approximate weight-memory formulas are:

  • FP32: parameters × 4 bytes
  • FP16 or BF16: parameters × 2 bytes
  • 8-bit: parameters × about 1 byte
  • 4-bit: parameters × about 0.5 bytes

These are estimates. Real model files are larger because quantization includes scales and metadata, and some tensors may remain at higher precision.

A more practical formula is:

Total VRAM ≈ model-file memory + KV-cache memory + runtime overhead + safety headroom

For example, a 32B model at idealized 4-bit storage would require:

32 billion × 0.5 bytes ≈ 16GB

The actual loaded model can require more than that. Quantization metadata, unquantized tensors, temporary buffers, and the KV cache may push the total substantially higher. It may still be a reasonable 24GB candidate at moderate context lengths, but it should not be treated as a guaranteed fit from the parameter count alone.

Why the model file size is a better first check

For a specific quantized download, the file size is often more useful than multiplying parameter count by a theoretical bit width. It includes much of the quantization overhead already.

However, file size is still not the complete VRAM requirement. A loader may allocate additional memory during model conversion or loading, and runtime memory grows as the context grows.

A useful screening process is:

  1. Check the model's total parameter count.
  2. Identify the exact quantization and file format.
  3. Check the quantized file size.
  4. Reserve memory for the KV cache and runtime.
  5. Compare the result with the GPU's usable, rather than advertised, VRAM.

What fits at each quantization level?

7B–8B models

Models in the 7B–8B range are the easiest fit for a 24GB GPU.

They commonly run in:

  • FP16 or BF16
  • 8-bit quantization
  • 4-bit quantization

Full-precision inference leaves more room for a large context and may preserve maximum model fidelity, but it consumes considerably more VRAM than 4-bit inference. For many users, 4-bit or 8-bit is chosen not because FP16 is impossible, but because the saved memory can be used for longer context, larger batches, or multiple concurrent requests.

12B–14B models

This is a particularly comfortable range for 24GB cards when using 4-bit or 8-bit quantization.

FP16 memory alone is approximately:

14 billion × 2 bytes ≈ 28GB

That already exceeds the card's capacity before runtime overhead, so a 14B model generally needs quantization or offload for a 24GB GPU.

At 4-bit, the idealized weight estimate is:

14 billion × 0.5 bytes ≈ 7GB

The loaded model will be larger than this estimate, but usually leaves useful room for the KV cache and inference runtime.

20B–24B models

In this range, 4-bit quantization is usually the practical starting point.

An idealized 24B 4-bit estimate is:

24 billion × 0.5 bytes ≈ 12GB

That does not mean every 24B model will use only 12GB in practice. The exact architecture, quantization scheme, runtime, and context length still matter. 8-bit versions may fit in some configurations but can leave little room for context or may require a model loader with efficient memory handling.

30B–34B models

Many 30B–34B models are realistic targets for a 24GB GPU in a well-chosen 4-bit quantization.

An idealized 32B estimate is:

32 billion × 0.5 bytes ≈ 16GB

The remaining capacity may be enough for inference overhead and a moderate context, but this is where headroom becomes important. A model with a large quantized file, a high context setting, or a memory-hungry backend may not fit fully in VRAM.

8-bit and FP16 versions in this range are generally too large for a single 24GB card without offload.

40B-class models

40B-class models are borderline even in 4-bit quantization. The idealized weight estimate is already about 20GB:

40 billion × 0.5 bytes ≈ 20GB

Real memory use can exceed the remaining capacity once the runtime and KV cache are included. Some carefully selected files may be possible with a short context or partial offload, but this is not the range to assume will work smoothly on every 24GB setup.

70B-class models

A 70B model in idealized 4-bit storage requires:

70 billion × 0.5 bytes ≈ 35GB

That is above the capacity of a 24GB GPU before any runtime memory is considered. A 70B model can sometimes be made to run with CPU/RAM offload, but the GPU will not hold all weights, and generation speed may be limited by transfers between system memory and VRAM.

If your goal is a fully GPU-resident 70B model, plan for more VRAM, such as multiple GPUs or a higher-memory accelerator.

Context length can determine whether a model fits

Loading the weights is only the first VRAM test. During generation, the model stores attention keys and values for the current prompt and conversation. This storage is called the KV cache.

KV-cache usage grows with:

  • Context length
  • Number of layers
  • Number of KV heads
  • Head dimension
  • Bytes per cache element
  • Batch size or number of simultaneous sequences

A simplified estimate for one sequence is:

KV cache bytes ≈ 2 × layers × context tokens × KV heads × head dimension × bytes per element

The factor of 2 represents keys and values.

This formula is useful for understanding scaling, but it is not a universal calculator. Some runtimes use different cache types, cache quantization, paging, or memory layouts. Architectures with grouped-query attention or multi-query attention can use much less KV memory than an otherwise similar model.

The practical consequence is straightforward:

  • A model that fits at 4,096 tokens may not fit at 32,000 or 128,000 tokens.
  • Increasing context length can reduce generation speed even if the model still loads.
  • A long context setting may reserve memory before you use the full context.
  • Batch size multiplies KV-cache requirements across simultaneous sequences.

If a 30B–34B 4-bit model is close to the VRAM limit, start with a moderate context setting and increase it gradually. Do not assume that the model's advertised maximum context is practical on a 24GB card.

Headroom matters more than the exact parameter count

Two models with the same parameter count can have different VRAM requirements because of:

  • Different quantization formats
  • Different numbers of layers and attention heads
  • Different KV-head configurations
  • Different vocabulary or embedding sizes
  • Unquantized output or embedding layers
  • Runtime-specific temporary buffers
  • Context and batch settings

For this reason, avoid choosing a model solely because:

parameter count × 0.5 bytes < 24GB

That calculation is a first-pass filter, not a loading guarantee.

A better rule is to divide candidates into three groups:

Comfortable fit: The model loads with meaningful room for context and normal runtime overhead.

Tight fit: The model may load, but context length, batch size, or backend settings need careful control.

Offload candidate: The model exceeds the practical GPU budget or leaves too little room for useful inference.

When CPU or system-RAM offload becomes necessary

Offload places some model weights or runtime data in system RAM instead of VRAM. It can make a model run when it would otherwise fail to load, but it changes the performance characteristics.

Offload is commonly needed when:

  • The quantized model weights alone approach or exceed usable VRAM.
  • You want a larger model than the GPU can hold.
  • You need a context length that exhausts VRAM.
  • You are using a higher-precision version for quality or compatibility.
  • You are running multiple models or simultaneous requests.

The trade-off is data movement. GPU memory has much higher bandwidth and lower access latency than ordinary system memory, while CPU-to-GPU transfers are constrained by the platform's interconnect and workload pattern. A partially offloaded model may work, but tokens can be generated much more slowly than with a fully GPU-resident model.

Offload is most attractive when:

  • You value model capability more than maximum tokens per second.
  • The amount offloaded is small.
  • You have sufficient system RAM.
  • You can tolerate slower prompt processing or generation.
  • The application supports configurable layer placement.

It is less attractive when you expect interactive speed, long contexts, or multiple concurrent users. In those cases, a smaller fully GPU-resident model may provide a better experience.

System RAM also needs room for the operating system, the model file, the loader, and any other applications. Do not plan to use all installed RAM for offload.

Mixture-of-experts models need special care

A mixture-of-experts, or MoE, model may advertise a relatively small number of active parameters per token while having a much larger total parameter count.

The active-parameter figure describes how much computation is used for each token. It does not automatically describe how much memory is required to store the model. Depending on the architecture and inference engine, many or all expert weights may still need to be available in memory.

When evaluating an MoE model, check:

  • Total parameters
  • Active parameters
  • Quantized file size
  • Whether experts can be dynamically loaded or offloaded
  • The runtime's documented memory behavior

Do not assume that an MoE labeled “12B active” has the same memory requirements as a dense 12B model.

What about image and other AI models?

“24GB GPU” compatibility is not limited to language models, but the memory calculation differs by workload.

For image generation, VRAM use depends on:

  • The base model or checkpoint
  • Precision and quantization
  • Image resolution
  • Batch size
  • Control modules, adapters, and additional encoders
  • The user interface and attention implementation

A setup that handles a language model comfortably may still need memory optimization for high-resolution image generation or multiple additional components. The same principle applies to video, speech, embedding, and multimodal models: model files are only part of the working-set memory.

For a specific non-LLM workflow, check the application's documented requirements and test the intended resolution, batch size, and add-ons rather than relying on parameter count alone.

A practical selection workflow

Use this process before downloading a large model:

  1. Define the workload. Decide whether you need chat, coding, reasoning, multimodal input, image generation, or another task.
  2. Choose a parameter range. On 24GB, start with 7B–14B for flexibility, or consider 30B–34B when 4-bit quality and model capability are more important.
  3. Select the quantization. Use FP16/BF16 when the model is small enough and you want maximum precision; use 8-bit or 4-bit to increase the feasible model size.
  4. Inspect the exact file size. Do not estimate from the model name alone.
  5. Budget the context. Decide whether you need a short chat context or a long document context.
  6. Leave safety margin. Keep room for the runtime, KV cache, and normal system allocations.
  7. Check backend compatibility. The model format must be supported by the inference application and GPU backend you plan to use.
  8. Plan for offload only intentionally. If the model exceeds the GPU budget, verify that you have enough system RAM and accept the likely speed trade-off.

For a model-specific answer, use RigForAI's What can I run tool. It is more useful than a generic parameter table when you want to compare a particular model, quantization, context length, and hardware configuration.

Bottom line

A 24GB GPU is best viewed as a 7B–14B high-flexibility machine and a 30B–34B 4-bit machine, with the exact boundary determined by context length and runtime overhead.

  • 7B–8B models: usually comfortable across common precisions.
  • 12B–14B models: excellent in 4-bit or 8-bit; usually not full FP16.
  • 20B–24B models: generally practical in 4-bit.
  • 30B–34B models: often practical in 4-bit, but check headroom carefully.
  • 40B-class models: borderline and configuration-dependent.
  • 70B models: normally require offload or additional GPUs.

The safest choice is not the largest model whose estimated weights barely fit. It is the largest model that leaves enough VRAM for the context, runtime, and workload you actually intend to use.

Related guides