Home / Guides / Why Tokens per Second Is Not Enough to Compare AI Servers

guide

Why Tokens per Second Is Not Enough to Compare AI Servers

Updated 2026-09-01

Tokens per second is useful, but it does not describe responsiveness, concurrency, context capacity, or user experience. This guide shows how to compare AI servers using throughput, latency, memory, and workload fit together.

Tokens per second (tok/s) is one of the most visible numbers in local AI testing. It is also one of the easiest numbers to misuse.

A server that generates 80 tokens per second in a single-request benchmark may be a worse choice than one generating 50 tokens per second if the second system:

  • Responds faster before generation starts
  • Handles several users without a large slowdown
  • Supports longer prompts and larger context windows
  • Maintains enough VRAM for the model, KV cache, and runtime overhead
  • Meets your actual latency and throughput targets

The better question is not “Which server has the highest tokens-per-second number?” It is:

Does this system meet the latency, concurrency, context, and reliability requirements of the workload?

Start with the serving objective

Before comparing hardware, define what the server is supposed to do. Different workloads optimize for different outcomes.

WorkloadPrimary concernImportant supporting metrics
Interactive chatFast time to first token and smooth streamingDecode speed, prompt processing time, p95 latency
Coding assistantLow latency on repeated short requestsTTFT, inter-token latency, concurrency
Document analysisPrompt processing and context capacityPrefill throughput, KV-cache memory, total request time
API servingAggregate work completed per unit of timeRequests per second, tokens per second, p95/p99 latency
Batch generationMaximum total throughputAggregate tok/s, batch size, completion time
Multi-user local serverPredictable performance under contentionPer-user latency, queueing, VRAM headroom

A useful serving objective should specify at least:

  1. Model and quantization
  2. Prompt length distribution
  3. Expected output length
  4. Number of simultaneous requests
  5. Target time to first token
  6. Target inter-token latency
  7. Required context length
  8. Acceptable p95 or p99 latency
  9. Whether throughput or responsiveness is more important

Without these details, a tok/s result is only a partial measurement.

What tokens per second actually measures

Generation normally has two materially different phases.

Prefill: processing the prompt

During prefill, the runtime processes the input prompt and builds the intermediate state needed for generation. Long prompts make this phase more important.

Prefill affects:

  • Time to first token (TTFT)
  • How quickly a request enters the decode phase
  • Throughput for document and retrieval-heavy workloads
  • The cost of processing large context windows

A benchmark that reports only output generation speed can hide a slow prefill phase.

Decode: generating output tokens

During decode, the model generates output one token at a time. Each new token depends on the previous tokens, so this phase has a sequential component.

Decode speed affects:

  • Streaming smoothness
  • Time required to produce the completion
  • Per-user interactive experience
  • Aggregate output throughput

A common approximation for one request is:

Total latency ≈ prefill time + (number of output tokens × time per output token)

If a benchmark reports decode speed as 50 tok/s, the approximate decode time for 200 output tokens is:

200 tokens / 50 tokens per second = 4 seconds

That does not include prompt processing, scheduling, queueing, network time, or final response handling.

Clarify what the benchmark counts

“Tokens per second” can refer to several different measurements:

  • Per-request decode speed: output speed for one active request
  • Aggregate throughput: total output tok/s across all active requests
  • Prefill throughput: prompt tokens processed per second
  • End-to-end throughput: completed work including waiting and overhead
  • Average speed: arithmetic average that may hide slow tail requests

These numbers are not interchangeable. A server may produce fewer tokens per individual request while producing more total tokens per second with batching.

The latency metrics that tok/s misses

Time to first token

TTFT is the time from request submission until the first output token arrives.

A simplified model is:

TTFT ≈ queue time + prefill time + scheduling overhead

For an interactive assistant, TTFT may matter more than peak decode speed. A response that streams at 100 tok/s but takes several seconds to begin can feel worse than one that starts quickly and streams at 40 tok/s.

Inter-token latency

Inter-token latency is the time between generated tokens. It is the inverse of decode speed:

Inter-token latency (seconds) ≈ 1 / decode speed (tokens per second)

For example:

1 / 40 tok/s = 0.025 seconds per token

That is approximately 25 milliseconds per token under the measured conditions. Actual user-visible timing can differ because of batching, output buffering, network delivery, and interface behavior.

End-to-end request latency

For a request with a prompt and completion, a more complete approximation is:

End-to-end latency = queue time + prefill time + decode time + transfer and application overhead

Where:

Decode time ≈ output tokens / sustained decode tok/s

This is the number to use when estimating how long a complete request takes.

Tail latency

Average latency is not enough for a multi-user service. Measure at least p95 latency, and preferably p99 when the application is sensitive to stalls.

  • p50: typical request
  • p95: 95% of requests finish at or below this latency
  • p99: exposes rarer but consequential slowdowns

Queueing, memory pressure, batch changes, and context-length variation can make tail latency much worse than the average.

Concurrency changes the meaning of performance

A one-request benchmark answers:

How fast can this system run this workload when it has the machine to itself?

It does not answer:

How will each user experience the server when several requests are active?

With concurrency, the runtime may use continuous batching or another scheduling strategy. It can combine work from multiple requests to improve hardware utilization, but that creates trade-offs.

Per-request speed versus aggregate throughput

Suppose a hypothetical test produces these results:

Concurrent requestsPer-request decode speedAggregate output throughput
160 tok/s60 tok/s
242 tok/s84 tok/s
425 tok/s100 tok/s
812 tok/s96 tok/s

The four-request result has lower per-user speed than the one-request result, but higher aggregate throughput. The eight-request result may indicate that the system has reached a saturation point: adding requests increases contention without increasing useful throughput.

For interactive use, the four-request setting might already feel too slow even though it maximizes aggregate output. For offline batch work, it may be the preferred operating point.

Queueing can dominate the result

Once active work exceeds what the runtime can serve promptly, new requests wait. A fast GPU does not eliminate queueing if:

  • The service accepts more simultaneous work than the configured batch can handle
  • VRAM limits the number of resident sequences
  • Long-context requests occupy cache space
  • CPU-side preprocessing or tokenization becomes a bottleneck
  • Requests are unevenly distributed in length

Measure latency at the concurrency level you actually expect, not only at concurrency one.

Context length and KV cache affect capacity

Longer context consumes memory for the key-value cache, commonly called the KV cache. This cache stores attention-related state for each active sequence so the runtime does not need to recompute the entire history at every decode step.

KV-cache use grows with factors such as:

  • Number of active sequences
  • Context length per sequence
  • Model architecture
  • Data type or cache quantization
  • Runtime implementation
  • Whether prefix caching or other optimizations are used

A useful planning relationship is:

Total KV-cache memory ≈ KV memory per token × cached tokens across active sequences

And:

Cached tokens ≈ sum of context tokens for all active requests

This explains why a model that fits comfortably for one short conversation may fail with several long conversations.

A practical VRAM budget

VRAM is not just model-file size. A rough planning model is:

Required VRAM ≈ model weights + KV cache + runtime/workspace overhead + safety headroom

The exact values depend on the model, quantization, backend, context, batch behavior, and offload configuration. Treat this as a budgeting framework, not a universal sizing formula.

If a deployment barely fits at startup, it may not be viable in production. Leave headroom for:

  • Longer-than-average prompts
  • Multiple active users
  • Temporary workspace allocations
  • Runtime fragmentation
  • Different batching decisions
  • Monitoring or serving components that share memory

Out-of-memory failures can appear only under concurrency, which is why a single-request load test is insufficient.

Hardware factors behind the measurements

Tokens per second is an outcome of several interacting limits rather than a direct specification of a server.

Memory bandwidth

Decode often repeatedly reads model weights and cache data. Memory bandwidth can therefore be important, especially for workloads that do not keep enough parallel work in flight to fully exploit compute resources.

A rough bandwidth-limited intuition is:

Approximate throughput ∝ usable memory bandwidth / bytes moved per generated token

This is not a benchmark formula. The bytes moved per token depend on architecture, quantization, cache behavior, batching, and kernel implementation.

Compute throughput

Prefill and larger batches can expose more parallel computation. Matrix-multiplication capability, supported numerical formats, and optimized kernels may matter more during prompt processing or high-concurrency serving than during a lightly loaded decode test.

VRAM capacity

Capacity determines whether the model, KV cache, and runtime can remain on the intended device. A larger memory pool can enable:

  • Larger models or less aggressive quantization
  • Longer contexts
  • More simultaneous sequences
  • More room for batching

Capacity alone does not guarantee higher tok/s. It may simply allow the workload to run without offloading or memory failures.

Interconnect and multi-GPU placement

When a model or workload is distributed across devices, data movement between GPUs can affect latency and throughput. The result depends on the runtime’s sharding strategy and communication pattern.

Do not assume that adding a second GPU doubles performance. It may improve capacity, provide more parallel resources, or introduce communication overhead. Validate the exact model and serving stack.

CPU, system memory, and storage

The GPU may not be the only bottleneck. CPU resources can affect tokenization, request scheduling, sampling, and data preparation. System memory and storage influence model loading and application behavior, although they do not substitute for adequate VRAM during active inference.

A small deployment example

Imagine a local team wants to serve a model for four simultaneous users. Its stated requirements are:

  • Four active requests
  • Prompts commonly around 4,000 tokens
  • Completions around 300 tokens
  • First token within 1.5 seconds at p95
  • At least 20 output tok/s per active user at p95
  • No out-of-memory failures during longer sessions

These requirements cannot be evaluated with a single single-user tok/s result.

Step 1: Estimate completion time

If the target sustained per-request decode speed is 20 tok/s:

300 output tokens / 20 tok/s = 15 seconds of decode time

The total request time is higher because it also includes queueing, prefill, and overhead.

Step 2: Account for concurrency

At four users, the server needs to sustain approximately:

4 users × 20 output tok/s = 80 aggregate output tok/s

That does not mean a single-user benchmark must show 80 tok/s per request. It means the system needs to be tested at concurrency four, with per-request and aggregate results reported separately.

Step 3: Check the prompt phase

A 4,000-token prompt may make prefill and TTFT significant. The test should report:

  • Prompt-processing time or prefill throughput
  • TTFT at concurrency four
  • Decode speed after the first token
  • p95 latency, not only the average

Step 4: Check memory under the real cache load

The deployment must fit:

model weights + KV cache for four active conversations + runtime overhead + headroom

If users can send longer prompts or retain longer conversation histories, size for the expected upper range rather than the average alone.

Step 5: Decide based on the objective

A configuration that delivers 90 aggregate tok/s but only 12 tok/s per user may satisfy a batch-processing objective but fail the interactive requirement. Conversely, a configuration delivering 22 tok/s per user with lower aggregate utilization may be the better choice for a chat service.

How to benchmark an AI server properly

Use a workload matrix rather than one headline number.

Vary the important dimensions

At minimum, test:

  • Prompt lengths: short, typical, and long
  • Output lengths: short and representative
  • Concurrency: 1, expected load, and overload
  • Context-window usage: current history and projected maximum
  • Batch or scheduler settings
  • Model and quantization used in deployment
  • Cold start versus warmed-up operation

Record the right metrics

For each test, record:

  • Prefill throughput
  • Time to first token
  • Per-request decode tok/s
  • Aggregate output tok/s
  • Total request latency
  • p50, p95, and p99 latency
  • Requests completed per minute or hour
  • Peak VRAM usage
  • System RAM usage
  • CPU utilization
  • Errors, retries, and out-of-memory events
  • Power draw and thermal behavior when relevant

Use sustained results after warm-up. Short bursts can benefit from transient conditions and may not reflect a long-running service.

Keep comparisons controlled

When comparing hardware, hold these constant where possible:

  • Exact model revision
  • Quantization format
  • Prompt and completion data
  • Runtime and version
  • Backend and kernel settings
  • Context length
  • Sampling configuration
  • Concurrency and scheduler policy

A faster result obtained with a different model, quantization, context, or runtime is not a clean hardware comparison.

Common failure modes and checks

Failure: comparing single-user tok/s for a multi-user service

Symptom: The benchmark looks excellent, but users experience delays.

Check: Test at expected concurrency and report both per-request latency and aggregate throughput.

Failure: ignoring prefill

Symptom: Short completions feel responsive, but long prompts take too long to start.

Check: Measure TTFT and prompt-processing throughput using realistic context lengths.

Failure: sizing VRAM from model weights alone

Symptom: The server starts successfully, then fails when several users connect.

Check: Include KV cache, runtime overhead, batch workspace, and headroom in the memory budget.

Failure: maximizing batch size without a latency target

Symptom: Aggregate tok/s improves, but interactive responses become sluggish.

Check: Find the highest batch or concurrency level that still meets p95 TTFT and inter-token latency targets.

Failure: trusting averages

Symptom: The average looks acceptable, but occasional requests stall badly.

Check: Track p95 and p99 latency, along with queue time and memory pressure.

Failure: assuming more GPUs scale linearly

Symptom: Additional hardware increases cost and complexity without proportional throughput.

Check: Benchmark the exact model placement, interconnect, runtime, and concurrency pattern.

A practical decision framework

When comparing candidate AI servers, use this sequence:

  1. Define the workload: model, quantization, context, output length, and concurrency.
  2. Set service targets: TTFT, per-user decode speed, aggregate throughput, and tail latency.
  3. Estimate memory: weights, KV cache, runtime overhead, and headroom.
  4. Identify the likely bottleneck: memory capacity, memory bandwidth, compute, CPU, or interconnect.
  5. Benchmark realistic load: include representative prompts and concurrent requests.
  6. Check sustained behavior: temperature, power limits, throttling, and long-run stability.
  7. Choose the least complex system that meets the target: peak benchmark performance is not the only consideration.

For planning a complete build, use RigForAI’s Find a GPU tool to narrow hardware choices by memory and system requirements. If you are comparing complete multi-GPU or rack-style configurations, review the GPU servers guide as well.

Tokens per second remains a useful metric. It becomes useful in context: alongside TTFT, inter-token latency, concurrency, aggregate throughput, context capacity, VRAM usage, and tail latency. The best AI server is the one that meets the workload’s service objective consistently—not necessarily the one with the largest number in a single benchmark.

Related guides