Home / Multi-GPU

llama.cpp · homogeneous 2–4 GPUs · not a PC builder

Can this model run across multiple GPUs?

Phase 8 answers memory feasibility for identical cards under llama.cpp. Default split is layer (pipeline parallel: each GPU owns a slice of layers; KV stays with those layers). This is not 2 × VRAM = pool, and it is not a complete PC build.

Cheapest current GPU-only multi-GPU configuration in RigForAI

Homogeneous llama.cpp layer configs that fully fit, among cards with a fresh matched Amazon offer. Cost = GPU count × current per-card price. This is not a complete build and does not prove that N units are in stock.

#ConfigurationPer-GPU peakGPU-only costWinning SKU
14× NVIDIA RTX 5000 32 GB23.5 GB$17919.84
4 × $4479.96
Amazon
Lenovo NVIDIA RTX 5000 Ada 32 GB GDDR6
ASIN B0FGQG418C
24× NVIDIA RTX 5090 32 GB23.5 GB$19199.96
4 × $4799.99
Amazon
GIGABYTE GeForce RTX 5090 WINDFORCE OC 32G Graphics Card
ASIN B0DT7GMXHB
33× NVIDIA RTX 6000 48 GB31.1 GB$62664.00
3 × $20888.00
Amazon
HP NVIDIA RTX 6000 Ada 48 GB 4DP Graphics
ASIN B0CTP2HHBN
44× NVIDIA RTX 6000 48 GB23.5 GB$83552.00
4 × $20888.00
Amazon
HP NVIDIA RTX 6000 Ada 48 GB 4DP Graphics
ASIN B0CTP2HHBN

Feasible — current price unavailable

Query combinations are not indexed. Source: llama.cpp multi-GPU documentation.