Home / Multi-GPU
llama.cpp · homogeneous 2–4 GPUs · not a PC builder
Can this model run across multiple GPUs?
Phase 8 answers memory feasibility for identical cards under llama.cpp. Default split is layer (pipeline parallel: each GPU owns a slice of layers; KV stays with those layers). This is not 2 × VRAM = pool, and it is not a complete PC build.
2× AMD Radeon RX 570 8 GB
Does not fit on this configuration
- Model
- Qwen3 235B-A22B · Q4 · 8K
- Runtime / split
- llama.cpp ·
layer· calculator 1.0.0 - Single-GPU status
- DOES_NOT_FIT · required 151.2 GB vs usable 7.2 GB
- Per card advertised / usable
- 8 GB / 7.2 GB
- Aggregate advertised
- 16.0 GB — capacity, not effective model capacity
- Nominal aggregate usable
- 14.4 GB (N × advertised × 0.90)
- Effective per-GPU peak
- 76.0 GB vs usable 7.2 GB
- Per-GPU split (approx.)
- weights 69.0 GB · KV 0.7 GB · runtime 3.9 GB · safety 2.3 GB
- Why
- Even a 2K context exceeds per-GPU usable VRAM after partition and per-device reserves.
- Topology
- Requires space for 2 discrete GPUs. Exact cooler slot width is not in the RigForAI graph. layer split is pipeline parallel and can run over PCIe. KV stays with the layers on each GPU. NVLink is not treated as a 1× aggregate VRAM pool.
- Performance
- Phase 5 predictions are single-GPU. Multi-GPU tok/s is not published.
- GPU-only current cost
- Cost unavailable — no fresh matched Amazon offer for this canonical GPU.
- Formula
- per_gpu_peak = ceil(weight_bytes/N) + ceil(kv_bytes/N) + runtime_base + safety_base + weight_fractions×ceil(weight_bytes/N). Compare to usable = advertised×0.90. Not (single_gpu_required × N) and not advertised_vram × N as effective capacity.
AMD Radeon RX 570 8 GB · Single-GPU can-run · Find a single GPU · Build a machine for this GPU setup
Query combinations are not indexed. Source: llama.cpp multi-GPU documentation.