Why Tokens per Second Is Not Enough to Compare AI Servers
Tokens per second is useful, but it does not describe responsiveness, concurrency, context capacity, or user experience. This guide shows how to compare AI servers using throughput, latency, memory, and workload fit together.
Tokens per second (tok/s) is one of the most visible numbers in local AI testing. It is also one of the easiest numbers to misuse.
A server that generates 80 tokens per second in a single-request benchmark may be a worse choice than one generating 50 tokens per second if the second system:
- Responds faster before generation starts
- Handles several users without a large slowdown
- Supports longer prompts and larger context windows
- Maintains enough VRAM for the model, KV cache, and runtime overhead
- Meets your actual latency and throughput targets
The better question is not “Which server has the highest tokens-per-second number?” It is:
Does this system meet the latency, concurrency, context, and reliability requirements of the workload?
Start with the serving objective
Before comparing hardware, define what the server is supposed to do. Different workloads optimize for different outcomes.
| Workload | Primary concern | Important supporting metrics |
|---|---|---|
| Interactive chat | Fast time to first token and smooth streaming | Decode speed, prompt processing time, p95 latency |
| Coding assistant | Low latency on repeated short requests | TTFT, inter-token latency, concurrency |
| Document analysis | Prompt processing and context capacity | Prefill throughput, KV-cache memory, total request time |
| API serving | Aggregate work completed per unit of time | Requests per second, tokens per second, p95/p99 latency |
| Batch generation | Maximum total throughput | Aggregate tok/s, batch size, completion time |
| Multi-user local server | Predictable performance under contention | Per-user latency, queueing, VRAM headroom |
A useful serving objective should specify at least:
- Model and quantization
- Prompt length distribution
- Expected output length
- Number of simultaneous requests
- Target time to first token
- Target inter-token latency
- Required context length
- Acceptable p95 or p99 latency
- Whether throughput or responsiveness is more important
Without these details, a tok/s result is only a partial measurement.
What tokens per second actually measures
Generation normally has two materially different phases.
Prefill: processing the prompt
During prefill, the runtime processes the input prompt and builds the intermediate state needed for generation. Long prompts make this phase more important.
Prefill affects:
- Time to first token (TTFT)
- How quickly a request enters the decode phase
- Throughput for document and retrieval-heavy workloads
- The cost of processing large context windows
A benchmark that reports only output generation speed can hide a slow prefill phase.
Decode: generating output tokens
During decode, the model generates output one token at a time. Each new token depends on the previous tokens, so this phase has a sequential component.
Decode speed affects:
- Streaming smoothness
- Time required to produce the completion
- Per-user interactive experience
- Aggregate output throughput
A common approximation for one request is:
Total latency ≈ prefill time + (number of output tokens × time per output token)
If a benchmark reports decode speed as 50 tok/s, the approximate decode time for 200 output tokens is:
200 tokens / 50 tokens per second = 4 seconds
That does not include prompt processing, scheduling, queueing, network time, or final response handling.
Clarify what the benchmark counts
“Tokens per second” can refer to several different measurements:
- Per-request decode speed: output speed for one active request
- Aggregate throughput: total output tok/s across all active requests
- Prefill throughput: prompt tokens processed per second
- End-to-end throughput: completed work including waiting and overhead
- Average speed: arithmetic average that may hide slow tail requests
These numbers are not interchangeable. A server may produce fewer tokens per individual request while producing more total tokens per second with batching.
The latency metrics that tok/s misses
Time to first token
TTFT is the time from request submission until the first output token arrives.
A simplified model is:
TTFT ≈ queue time + prefill time + scheduling overhead
For an interactive assistant, TTFT may matter more than peak decode speed. A response that streams at 100 tok/s but takes several seconds to begin can feel worse than one that starts quickly and streams at 40 tok/s.
Inter-token latency
Inter-token latency is the time between generated tokens. It is the inverse of decode speed:
Inter-token latency (seconds) ≈ 1 / decode speed (tokens per second)
For example:
1 / 40 tok/s = 0.025 seconds per token
That is approximately 25 milliseconds per token under the measured conditions. Actual user-visible timing can differ because of batching, output buffering, network delivery, and interface behavior.
End-to-end request latency
For a request with a prompt and completion, a more complete approximation is:
End-to-end latency = queue time + prefill time + decode time + transfer and application overhead
Where:
Decode time ≈ output tokens / sustained decode tok/s
This is the number to use when estimating how long a complete request takes.
Tail latency
Average latency is not enough for a multi-user service. Measure at least p95 latency, and preferably p99 when the application is sensitive to stalls.
- p50: typical request
- p95: 95% of requests finish at or below this latency
- p99: exposes rarer but consequential slowdowns
Queueing, memory pressure, batch changes, and context-length variation can make tail latency much worse than the average.
Concurrency changes the meaning of performance
A one-request benchmark answers:
How fast can this system run this workload when it has the machine to itself?
It does not answer:
How will each user experience the server when several requests are active?
With concurrency, the runtime may use continuous batching or another scheduling strategy. It can combine work from multiple requests to improve hardware utilization, but that creates trade-offs.
Per-request speed versus aggregate throughput
Suppose a hypothetical test produces these results:
| Concurrent requests | Per-request decode speed | Aggregate output throughput |
|---|---|---|
| 1 | 60 tok/s | 60 tok/s |
| 2 | 42 tok/s | 84 tok/s |
| 4 | 25 tok/s | 100 tok/s |
| 8 | 12 tok/s | 96 tok/s |
The four-request result has lower per-user speed than the one-request result, but higher aggregate throughput. The eight-request result may indicate that the system has reached a saturation point: adding requests increases contention without increasing useful throughput.
For interactive use, the four-request setting might already feel too slow even though it maximizes aggregate output. For offline batch work, it may be the preferred operating point.
Queueing can dominate the result
Once active work exceeds what the runtime can serve promptly, new requests wait. A fast GPU does not eliminate queueing if:
- The service accepts more simultaneous work than the configured batch can handle
- VRAM limits the number of resident sequences
- Long-context requests occupy cache space
- CPU-side preprocessing or tokenization becomes a bottleneck
- Requests are unevenly distributed in length
Measure latency at the concurrency level you actually expect, not only at concurrency one.
Context length and KV cache affect capacity
Longer context consumes memory for the key-value cache, commonly called the KV cache. This cache stores attention-related state for each active sequence so the runtime does not need to recompute the entire history at every decode step.
KV-cache use grows with factors such as:
- Number of active sequences
- Context length per sequence
- Model architecture
- Data type or cache quantization
- Runtime implementation
- Whether prefix caching or other optimizations are used
A useful planning relationship is:
Total KV-cache memory ≈ KV memory per token × cached tokens across active sequences
And:
Cached tokens ≈ sum of context tokens for all active requests
This explains why a model that fits comfortably for one short conversation may fail with several long conversations.
A practical VRAM budget
VRAM is not just model-file size. A rough planning model is:
Required VRAM ≈ model weights + KV cache + runtime/workspace overhead + safety headroom
The exact values depend on the model, quantization, backend, context, batch behavior, and offload configuration. Treat this as a budgeting framework, not a universal sizing formula.
If a deployment barely fits at startup, it may not be viable in production. Leave headroom for:
- Longer-than-average prompts
- Multiple active users
- Temporary workspace allocations
- Runtime fragmentation
- Different batching decisions
- Monitoring or serving components that share memory
Out-of-memory failures can appear only under concurrency, which is why a single-request load test is insufficient.
Hardware factors behind the measurements
Tokens per second is an outcome of several interacting limits rather than a direct specification of a server.
Memory bandwidth
Decode often repeatedly reads model weights and cache data. Memory bandwidth can therefore be important, especially for workloads that do not keep enough parallel work in flight to fully exploit compute resources.
A rough bandwidth-limited intuition is:
Approximate throughput ∝ usable memory bandwidth / bytes moved per generated token
This is not a benchmark formula. The bytes moved per token depend on architecture, quantization, cache behavior, batching, and kernel implementation.
Compute throughput
Prefill and larger batches can expose more parallel computation. Matrix-multiplication capability, supported numerical formats, and optimized kernels may matter more during prompt processing or high-concurrency serving than during a lightly loaded decode test.
VRAM capacity
Capacity determines whether the model, KV cache, and runtime can remain on the intended device. A larger memory pool can enable:
- Larger models or less aggressive quantization
- Longer contexts
- More simultaneous sequences
- More room for batching
Capacity alone does not guarantee higher tok/s. It may simply allow the workload to run without offloading or memory failures.
Interconnect and multi-GPU placement
When a model or workload is distributed across devices, data movement between GPUs can affect latency and throughput. The result depends on the runtime’s sharding strategy and communication pattern.
Do not assume that adding a second GPU doubles performance. It may improve capacity, provide more parallel resources, or introduce communication overhead. Validate the exact model and serving stack.
CPU, system memory, and storage
The GPU may not be the only bottleneck. CPU resources can affect tokenization, request scheduling, sampling, and data preparation. System memory and storage influence model loading and application behavior, although they do not substitute for adequate VRAM during active inference.
A small deployment example
Imagine a local team wants to serve a model for four simultaneous users. Its stated requirements are:
- Four active requests
- Prompts commonly around 4,000 tokens
- Completions around 300 tokens
- First token within 1.5 seconds at p95
- At least 20 output tok/s per active user at p95
- No out-of-memory failures during longer sessions
These requirements cannot be evaluated with a single single-user tok/s result.
Step 1: Estimate completion time
If the target sustained per-request decode speed is 20 tok/s:
300 output tokens / 20 tok/s = 15 seconds of decode time
The total request time is higher because it also includes queueing, prefill, and overhead.
Step 2: Account for concurrency
At four users, the server needs to sustain approximately:
4 users × 20 output tok/s = 80 aggregate output tok/s
That does not mean a single-user benchmark must show 80 tok/s per request. It means the system needs to be tested at concurrency four, with per-request and aggregate results reported separately.
Step 3: Check the prompt phase
A 4,000-token prompt may make prefill and TTFT significant. The test should report:
- Prompt-processing time or prefill throughput
- TTFT at concurrency four
- Decode speed after the first token
- p95 latency, not only the average
Step 4: Check memory under the real cache load
The deployment must fit:
model weights + KV cache for four active conversations + runtime overhead + headroom
If users can send longer prompts or retain longer conversation histories, size for the expected upper range rather than the average alone.
Step 5: Decide based on the objective
A configuration that delivers 90 aggregate tok/s but only 12 tok/s per user may satisfy a batch-processing objective but fail the interactive requirement. Conversely, a configuration delivering 22 tok/s per user with lower aggregate utilization may be the better choice for a chat service.
How to benchmark an AI server properly
Use a workload matrix rather than one headline number.
Vary the important dimensions
At minimum, test:
- Prompt lengths: short, typical, and long
- Output lengths: short and representative
- Concurrency: 1, expected load, and overload
- Context-window usage: current history and projected maximum
- Batch or scheduler settings
- Model and quantization used in deployment
- Cold start versus warmed-up operation
Record the right metrics
For each test, record:
- Prefill throughput
- Time to first token
- Per-request decode tok/s
- Aggregate output tok/s
- Total request latency
- p50, p95, and p99 latency
- Requests completed per minute or hour
- Peak VRAM usage
- System RAM usage
- CPU utilization
- Errors, retries, and out-of-memory events
- Power draw and thermal behavior when relevant
Use sustained results after warm-up. Short bursts can benefit from transient conditions and may not reflect a long-running service.
Keep comparisons controlled
When comparing hardware, hold these constant where possible:
- Exact model revision
- Quantization format
- Prompt and completion data
- Runtime and version
- Backend and kernel settings
- Context length
- Sampling configuration
- Concurrency and scheduler policy
A faster result obtained with a different model, quantization, context, or runtime is not a clean hardware comparison.
Common failure modes and checks
Failure: comparing single-user tok/s for a multi-user service
Symptom: The benchmark looks excellent, but users experience delays.
Check: Test at expected concurrency and report both per-request latency and aggregate throughput.
Failure: ignoring prefill
Symptom: Short completions feel responsive, but long prompts take too long to start.
Check: Measure TTFT and prompt-processing throughput using realistic context lengths.
Failure: sizing VRAM from model weights alone
Symptom: The server starts successfully, then fails when several users connect.
Check: Include KV cache, runtime overhead, batch workspace, and headroom in the memory budget.
Failure: maximizing batch size without a latency target
Symptom: Aggregate tok/s improves, but interactive responses become sluggish.
Check: Find the highest batch or concurrency level that still meets p95 TTFT and inter-token latency targets.
Failure: trusting averages
Symptom: The average looks acceptable, but occasional requests stall badly.
Check: Track p95 and p99 latency, along with queue time and memory pressure.
Failure: assuming more GPUs scale linearly
Symptom: Additional hardware increases cost and complexity without proportional throughput.
Check: Benchmark the exact model placement, interconnect, runtime, and concurrency pattern.
A practical decision framework
When comparing candidate AI servers, use this sequence:
- Define the workload: model, quantization, context, output length, and concurrency.
- Set service targets: TTFT, per-user decode speed, aggregate throughput, and tail latency.
- Estimate memory: weights, KV cache, runtime overhead, and headroom.
- Identify the likely bottleneck: memory capacity, memory bandwidth, compute, CPU, or interconnect.
- Benchmark realistic load: include representative prompts and concurrent requests.
- Check sustained behavior: temperature, power limits, throttling, and long-run stability.
- Choose the least complex system that meets the target: peak benchmark performance is not the only consideration.
For planning a complete build, use RigForAI’s Find a GPU tool to narrow hardware choices by memory and system requirements. If you are comparing complete multi-GPU or rack-style configurations, review the GPU servers guide as well.
Tokens per second remains a useful metric. It becomes useful in context: alongside TTFT, inter-token latency, concurrency, aggregate throughput, context capacity, VRAM usage, and tail latency. The best AI server is the one that meets the workload’s service objective consistently—not necessarily the one with the largest number in a single benchmark.