Home / Multi-GPU
llama.cpp · homogeneous 2–4 GPUs · not a PC builder
Can this model run across multiple GPUs?
Phase 8 answers memory feasibility for identical cards under llama.cpp. Default split is layer (pipeline parallel: each GPU owns a slice of layers; KV stays with those layers). This is not 2 × VRAM = pool, and it is not a complete PC build.
2× NVIDIA RTX 3060 12 GB
Fits fully across the GPUs
- Model
- Qwen3 30B-A3B · Q4 · 8K
- Runtime / split
- llama.cpp ·
layer· calculator 1.0.0 - Single-GPU status
- CPU_RAM_OFFLOAD_REQUIRED · required 20.8 GB vs usable 10.8 GB
- Per card advertised / usable
- 12 GB / 10.8 GB
- Aggregate advertised
- 24.0 GB — capacity, not effective model capacity
- Nominal aggregate usable
- 21.6 GB (N × advertised × 0.90)
- Effective per-GPU peak
- 10.8 GB vs usable 10.8 GB
- Per-GPU split (approx.)
- weights 9.0 GB · KV 0.4 GB · runtime 0.9 GB · safety 0.5 GB
- Why
- Each homogeneous GPU holds about 1/2 of weights and 1/2 of KV under llama.cpp layer split, plus a full per-device runtime/safety reserve. Peak per GPU is within usable VRAM.
- Topology
- Requires space for 2 discrete GPUs. Exact cooler slot width is not in the RigForAI graph. layer split is pipeline parallel and can run over PCIe. KV stays with the layers on each GPU. NVLink is not treated as a 1× aggregate VRAM pool.
- Performance
- Phase 5 predictions are single-GPU. Multi-GPU tok/s is not published.
- GPU-only current cost
- $639.98 = 2 × $319.99 at the current per-card matched price. This does not mean 2 units are in stock, and it is not a complete build cost.GIGABYTE GeForce RTX 3060 VISION OC 12G NVIDIA 12 GB GDDR6 · ASIN B0971BRCM4
- Formula
- per_gpu_peak = ceil(weight_bytes/N) + ceil(kv_bytes/N) + runtime_base + safety_base + weight_fractions×ceil(weight_bytes/N). Compare to usable = advertised×0.90. Not (single_gpu_required × N) and not advertised_vram × N as effective capacity.
NVIDIA RTX 3060 12 GB · Single-GPU can-run · Find a single GPU · Build a machine for this GPU setup
Query combinations are not indexed. Source: llama.cpp multi-GPU documentation.