Home / Servers / 4× NVIDIA RTX 3090 24 GB
Canonical hardware · not a commercial offer
4× NVIDIA RTX 3090 24 GB
252 GB system RAM · AMD EPYC 7502 32-Core Processor
Hardware
- GPU
- NVIDIA RTX 3090 24 GB
- GPU count
- 4
- VRAM per GPU
- 24.0 GB advertised
- Aggregate VRAM
- 96.0 GB advertised
- System RAM
- 252 GB
- CPU
- AMD EPYC 7502 32-Core Processor
- Storage
- 1280 GB SSD
- Network
- 353 down · 299.1 up
AI memory capability
Default catalog below is Q4 at 8K. Statuses reuse Phase 2 (1 GPU) or Phase 8 (2–4 identical GPUs). GPU counts above 4 are not modeled. Public wording is memory-capable, not serving throughput.
Memory-capable at Q4/8K: 54 published models
- Qwen2.5 0.5B Instruct
- Llama 3.2 1B Instruct
- Qwen2.5 1.5B Instruct
- Qwen2.5-Coder 1.5B Instruct
- SmolLM2 1.7B Instruct
- Gemma 2 2B Instruct
- Qwen2.5 3B Instruct
- Llama 3.2 3B Instruct
- Phi-3.5 Mini Instruct
- Phi-3 Mini 128K Instruct
- Phi-4 Mini Instruct
- Code Llama 7B Instruct
- StarCoder2 7B
- Falcon 7B Instruct
- OpenChat 3.5 7B
- Mistral 7B Instruct v0.3
- Falcon 3 7B Instruct
- DeepSeek R1 Distill Qwen 7B
- Qwen2.5 7B Instruct
- Qwen2.5-Coder 7B Instruct
- Qwen2 7B Instruct
- DeepSeek R1 Distill Llama 8B
- Hermes 3 Llama 3.1 8B
- Llama 3.1 8B Instruct
Serving performance
Measured vLLM serving results for this GPU identity (4× NVIDIA RTX 3090 24 GB), not a specific rental host. Tensor parallel size is stored explicitly and is not assumed equal to GPU count. Request rate is separate from concurrency; this corpus is closed-loop / max-concurrency.
No measured vLLM serving rows map to this GPU count yet.
Current rental offers
| Provider | Location | Price | Stock | |
|---|---|---|---|---|
| Vast.ai | JP | $0.406/hour · ~$296/month equivalent (hourly × 730) | Listed as available | Provider |
Want to own the hardware? Build a PC · Prefer a prebuilt? Workstations