Home / Servers
Find a server for your AI model
Memory-capable means the selected model, quantization and context fit the server's GPU memory (and system RAM) using the same calculators as local hardware. Serving numbers, when shown, are official vLLM nightly measurements for matching GPU identity — not Phase 5 single-user tok/s and not this exact rental host.
Compare buy vs rent using these rental prices against a current local build or workstation.
Serving evaluation uses measured vLLM TTFT and TPOT at an explicit input/output shape and concurrency. Missing evidence is reported as insufficient, not as a fail.
Workload shapes are only those present in official vLLM nightly evidence. Request rate is not interchangeable with concurrency. P95 is not published by this source. Targets are yours, not an industry SLA.
Memory filter: Vicuna 13B v1.5 at Q4 / 8K. Offers that are memory-capable remain listed even when serving evidence is absent.
No matching GPU server offers in the current fresh inventory.
Providers
GPU / dedicated AI compute only. Not a generic hosting directory. All providers with inventory
Want to own the hardware? Build a PC · Prefer a prebuilt? Workstations