Home / Multi-GPU

llama.cpp · homogeneous 2–4 GPUs · not a PC builder

Can this model run across multiple GPUs?

Phase 8 answers memory feasibility for identical cards under llama.cpp. Default split is layer (pipeline parallel: each GPU owns a slice of layers; KV stays with those layers). This is not 2 × VRAM = pool, and it is not a complete PC build.

2× NVIDIA RTX 5070 Ti 16 GB

Fits only at a shorter context

Model
Gemma 4 31B Instruct · Q4 · 8K
Runtime / split
llama.cpp · layer · calculator 1.0.0
Single-GPU status
CPU_RAM_OFFLOAD_REQUIRED · required 28.1 GB vs usable 14.4 GB
Per card advertised / usable
16 GB / 14.4 GB
Aggregate advertised
32.0 GB — capacity, not effective model capacity
Nominal aggregate usable
28.8 GB (N × advertised × 0.90)
Effective per-GPU peak
14.4 GB vs usable 14.4 GB
Per-GPU split (approx.)
weights 9.2 GB · KV 3.8 GB · runtime 1.0 GB · safety 0.5 GB
Why
Requested context does not fit. A shorter context (up to 4096 tokens) does under the same split.
Topology
Requires space for 2 discrete GPUs. Exact cooler slot width is not in the RigForAI graph. layer split is pipeline parallel and can run over PCIe. KV stays with the layers on each GPU. NVLink is not treated as a 1× aggregate VRAM pool.
Performance
Phase 5 predictions are single-GPU. Multi-GPU tok/s is not published.
GPU-only current cost
$2099.98 = 2 × $1049.99 at the current per-card matched price. This does not mean 2 units are in stock, and it is not a complete build cost.
MSI GAMING GeForce RTX 5070 Ti 16G TRIO OC NVIDIA 16 GB GDDR7 · ASIN B0F11KQQF4
Amazon
Formula
per_gpu_peak = ceil(weight_bytes/N) + ceil(kv_bytes/N) + runtime_base + safety_base + weight_fractions×ceil(weight_bytes/N). Compare to usable = advertised×0.90. Not (single_gpu_required × N) and not advertised_vram × N as effective capacity.

NVIDIA RTX 5070 Ti 16 GB · Single-GPU can-run · Find a single GPU · Build a machine for this GPU setup

Query combinations are not indexed. Source: llama.cpp multi-GPU documentation.