Home / Servers / 2× NVIDIA RTX 5080 16 GB
Canonical hardware · not a commercial offer
2× NVIDIA RTX 5080 16 GB
126 GB system RAM · AMD EPYC 7532 32-Core Processor
Hardware
- GPU
- NVIDIA RTX 5080 16 GB
- GPU count
- 2
- VRAM per GPU
- 16.0 GB advertised
- Aggregate VRAM
- 32.0 GB advertised
- System RAM
- 126 GB
- CPU
- AMD EPYC 7532 32-Core Processor
- Storage
- 1555 GB SSD
- Network
- 734.9 down · 892.9 up
AI memory capability
Default catalog below is Q4 at 8K. Statuses reuse Phase 2 (1 GPU) or Phase 8 (2–4 identical GPUs). GPU counts above 4 are not modeled. Public wording is memory-capable, not serving throughput.
Memory-capable at Q4/8K: 47 published models
- Qwen2.5 0.5B Instruct
- Llama 3.2 1B Instruct
- Qwen2.5 1.5B Instruct
- Qwen2.5-Coder 1.5B Instruct
- SmolLM2 1.7B Instruct
- Gemma 2 2B Instruct
- Qwen2.5 3B Instruct
- Llama 3.2 3B Instruct
- Phi-3.5 Mini Instruct
- Phi-3 Mini 128K Instruct
- Phi-4 Mini Instruct
- Code Llama 7B Instruct
- StarCoder2 7B
- Falcon 7B Instruct
- OpenChat 3.5 7B
- Mistral 7B Instruct v0.3
- Falcon 3 7B Instruct
- DeepSeek R1 Distill Qwen 7B
- Qwen2.5 7B Instruct
- Qwen2.5-Coder 7B Instruct
- Qwen2 7B Instruct
- DeepSeek R1 Distill Llama 8B
- Hermes 3 Llama 3.1 8B
- Llama 3.1 8B Instruct
Serving performance
Measured vLLM serving results for this GPU identity (2× NVIDIA RTX 5080 16 GB), not a specific rental host. Tensor parallel size is stored explicitly and is not assumed equal to GPU count. Request rate is separate from concurrency; this corpus is closed-loop / max-concurrency.
No measured vLLM serving rows map to this GPU count yet.
Current rental offers
| Provider | Location | Price | Stock | |
|---|---|---|---|---|
| Vast.ai | US | $0.430/hour · ~$314/month equivalent (hourly × 730) | Listed as available | Provider |
Want to own the hardware? Build a PC · Prefer a prebuilt? Workstations