Home / Multi-GPU
llama.cpp · homogeneous 2–4 GPUs · not a PC builder
Can this model run across multiple GPUs?
Phase 8 answers memory feasibility for identical cards under llama.cpp. Default split is layer (pipeline parallel: each GPU owns a slice of layers; KV stays with those layers). This is not 2 × VRAM = pool, and it is not a complete PC build.
Cheapest current GPU-only multi-GPU configuration in RigForAI
Homogeneous llama.cpp layer configs that fully fit, among cards with a fresh matched Amazon offer. Cost = GPU count × current per-card price. This is not a complete build and does not prove that N units are in stock.
| # | Configuration | Per-GPU peak | GPU-only cost | Winning SKU |
|---|---|---|---|---|
| 1 | 4× NVIDIA RTX 5000 32 GB | 23.5 GB | $17919.84 4 × $4479.96 Amazon | Lenovo NVIDIA RTX 5000 Ada 32 GB GDDR6 ASIN B0FGQG418C |
| 2 | 4× NVIDIA RTX 5090 32 GB | 23.5 GB | $19199.96 4 × $4799.99 Amazon | GIGABYTE GeForce RTX 5090 WINDFORCE OC 32G Graphics Card ASIN B0DT7GMXHB |
| 3 | 3× NVIDIA RTX 6000 48 GB | 31.1 GB | $62664.00 3 × $20888.00 Amazon | HP NVIDIA RTX 6000 Ada 48 GB 4DP Graphics ASIN B0CTP2HHBN |
| 4 | 4× NVIDIA RTX 6000 48 GB | 23.5 GB | $83552.00 4 × $20888.00 Amazon | HP NVIDIA RTX 6000 Ada 48 GB 4DP Graphics ASIN B0CTP2HHBN |
Feasible — current price unavailable
- 4× NVIDIA Tesla M10 32 GB
- 4× NVIDIA Quadro GV100 32 GB
- 3× NVIDIA RTX 8000 48 GB
- 4× NVIDIA RTX 8000 48 GB
- 3× NVIDIA RTX A6000 48 GB
- 4× NVIDIA RTX A6000 48 GB
Query combinations are not indexed. Source: llama.cpp multi-GPU documentation.