LLM latency targets: time to first token and tokens per second for chat, RAG and agents
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- NVIDIA’s NIM benchmarking guide defines time to first token (TTFT) as the time from query submission to the first received token, including queuing, prefill and network latency, and inter-token latency (ITL) as the average gap between tokens, also known as time per output token (TPOT)
- vllm bench serve reports TPOT as each request’s average gap after the first token and ITL as every single gap, so a figure called ITL means a different thing in vLLM than in AIPerf; note the tool with every number
- TTFT is driven by the queue and by prefill, which is compute-bound and grows with prompt length; per-user tokens per second is driven by decode, which is bound by memory bandwidth and slows as the batch grows
- Our example targets at the 95th percentile: chat 1 s to first token and a TPOT of 100 ms (10 tokens per second), RAG 2.5 s and 100 ms, agent steps 1 s and 50 ms, batch jobs a throughput target instead
- The number of concurrent users one GPU serves depends on the latency target: NVIDIA’s guide says total throughput rises with concurrency while throughput per user falls, so measure goodput, the requests per second that meet every target
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
LLM latency metrics: TTFT, inter-token latency and throughput
LLM latency is described by four numbers: time to first token (TTFT), the gap between output tokens, called inter-token latency (ITL) or time per output token (TPOT), end-to-end latency per request, and throughput. TTFT is what a user waits before text appears; the gap between tokens sets how fast the text then streams. Both have to be stated with a percentile and with the tool that measured them, because tools compute them differently. For a chat assistant our example targets, derived below, are 1 s to the first token and 100 ms per token (10 tokens per second per user) at the 95th percentile.
NVIDIA’s NIM benchmarking guide, as read on 10 October 2026, defines TTFT as “the time from query submission to the first received token, if the response is not empty”, and says it “generally includes request queuing time, prefill time, and network latency”. It defines ITL as “the average time between consecutive tokens” and adds that it “is also known as time per output token (TPOT)”. End-to-end latency is TTFT plus generation time, the time from the first to the last token received.
| METRIC | NVIDIA NIM GUIDE | VLLM BENCH SERVE | DRIVEN BY |
|---|---|---|---|
| Time to first token | submission to first received token; queuing, prefill and network included | ttft, per request, at the client | queue wait, prompt length, prefill compute |
| Inter-token latency | average gap between consecutive tokens, the same as TPOT | itl, every gap between outputs | memory bandwidth, batch size, prefills sharing a step |
| Time per output token | the same metric as ITL | tpot, each request’s average gap after the first token | as ITL |
| End-to-end latency | submission to the full response: TTFT plus generation time | e2el, per request | TTFT, gap between tokens, answer length |
| Throughput per user | output length divided by end-to-end latency, approaching 1/ITL | not reported separately | as ITL |
| System throughput | total tokens per second and requests per second | request and output token throughput | batch size, until compute saturates |
NVIDIA NIM LLM benchmarking guide, metrics page; vLLM bench serve CLI reference and benchmark source; all read on 10 October 2026. The “driven by” column is our summary of NVIDIA’s inference optimisation blog of 17 November 2023 and vLLM’s optimisation guide.
Throughput per user and system throughput answer different questions. NVIDIA’s guide states that “total system TPS increases while TPS per user decreases as latency increases”. For sizing, a figure in tokens per second has to say which of the two it is.
TTFT vs TPOT and ITL: where vLLM and AIPerf differ
The definitions agree on TTFT and end-to-end latency and diverge on the gap between tokens. NVIDIA’s guide notes that “tools differ on whether TTFT is included in the average” and says to “compare results only when definitions align”.
AIPerf, the benchmarking tool NVIDIA now points to, computes ITL per request as end-to-end latency less TTFT, divided by the output tokens less one, and reports the gaps between all streamed chunks separately as inter chunk latency. Its output token throughput per user is 1 divided by that ITL. vllm bench serve computes the same per-request average and calls it TPOT, while its ITL is the list of every individual gap, pooled across requests. An ITL percentile from vLLM therefore shows stalls that an AIPerf ITL percentile, built from averages, smooths away.
Where the measurement starts matters too. Both tools time from the client, so their TTFT includes the network and, when the test runs through it, the gateway in front of the server. vLLM’s own server histogram starts inside the server and leaves both out, as our guide to LLM serving metrics in Prometheus explains; a target written for one cannot be checked against the other. For reasoning models AIPerf adds time to first output token, which it calculates as the time “from request start to the first non-reasoning output token”, the point at which the user sees the answer rather than the model’s thinking.
What drives time to first token and tokens per second
TTFT has two parts: waiting in the queue and prefill, the processing of the whole prompt before the first token. NVIDIA’s inference optimisation blog describes prefill as “a matrix-matrix operation that’s highly parallelized”, which “effectively saturates GPU utilization”, so its time grows with prompt length and falls with compute. NVIDIA’s benchmarking guide adds that “time to first token can increase due to queueing delay” and that longer input sequences increase TTFT. A RAG prompt with several retrieved passages costs more prefill than a chat question, and an agent that re-sends its growing context pays it again at every step.
The gap between tokens comes from decode, where each step reads the weights and the KV cache. The same NVIDIA blog calls this “a memory-bound operation” whose latency is dominated by the speed at which data is transferred from memory. For one user, tokens per second are capped by memory bandwidth divided by the bytes read per token. By that arithmetic, Llama 3.3 70B in FP8 tops out at about 23 tokens per second on an RTX PRO 6000 Server Edition and about 68 on an H200 NVL, as our RTX PRO 6000 vs H200 NVL comparison shows; these are upper bounds that no system reaches.
Prefill and decode share the GPU, so the scheduler trades one against the other. In vLLM, chunked prefill is on by default where possible, and the scheduler “prioritizes decode requests”. Its optimisation guide says that smaller values of max_num_batched_tokens “achieve better ITL because there are fewer prefills slowing down decodes”, while “higher values achieve better time to first token”. A server tuned for RAG prompts and a server tuned for fast streaming are configured differently.
Latency and concurrent users on one GPU
Concurrency raises total throughput and lowers per-user speed until the GPU saturates. NVIDIA’s parameters page states that “throughput generally saturates near the max batch size while latency steadily increases”. Past that point more concurrent requests add queue time to TTFT and give no extra throughput.
The number of users a card serves is therefore set by the latency target as well as by memory. Memory decides how many sessions fit in the KV cache; our guide to concurrent users per RTX PRO 6000 gives that arithmetic and the batch sweep from NVIDIA’s TensorRT-LLM tuning guide. The latency target decides how many of those sessions can run at once and still meet it. When the per-user target is above the single-stream ceiling, no batch size helps: a 40 ms target is 25 tokens per second, above the ceiling of about 23 for a 70B FP8 model on one RTX PRO 6000 Server Edition. NVFP4 weights, a smaller model, a card with more bandwidth, tensor parallelism across two cards or speculative decoding change that figure; more concurrency does not.
Example latency targets for chat, RAG, agents and batch jobs
Targets come from the use case, and the values below are our examples, not vendor recommendations. We placed them on edges of vLLM’s histogram buckets, so production monitoring can count exactly how many requests meet them. Our guide to users per RTX PRO 6000 puts mean silent reading at about 5.3 tokens per second, from a 2019 meta-analysis, so 10 tokens per second streams faster than the average reader. We state the per-token targets as TPOT, the per-request average that vLLM’s goodput checks and AIPerf reports as ITL.
| USE CASE | EXAMPLE TARGETS | REQUEST SHAPE | WHAT TO SIZE |
|---|---|---|---|
| Chat assistant | TTFT 1 s, TPOT 100 ms, at p95 | short prompts, answers of a few hundred tokens | KV cache for peak sessions, bandwidth per user |
| RAG question answering | TTFT 2.5 s, TPOT 100 ms, at p95 | prompts of several thousand tokens | prefill compute, KV cache per long prompt |
| Agent step | TTFT 1 s, TPOT 50 ms, plus a time per task | many calls, growing context, short outputs | prefill per step, per-user speed |
| Reasoning model | time to first output token, TPOT 50 ms | hundreds to thousands of thinking tokens | per-user speed, answer length |
| Batch documents | documents per hour, end-to-end time per document | long input, short output | system throughput |
Example targets and request shapes are ours, at vLLM’s TTFT bucket edges of 1 and 2.5 s and its inter-token and TPOT bucket edges of 50 and 100 ms; test them with your users. NVIDIA’s NIM benchmarking guide lists retrieval under summarisation, with ISL near 1,000 and OSL near 100 tokens.
Turn the targets into end-to-end times before you agree them. A chat answer of 400 tokens with 1 s to the first token and 100 ms per token takes 1 + 399 × 0.1, about 41 s, of which the user reads along from the first second. An agent task of 8 steps, each with 1 s to the first token and 150 output tokens at 50 ms, takes 8 × (1 + 149 × 0.05), about 68 s before tool calls; at 100 ms per token it takes about 127 s. A reasoning model that thinks for 1,500 tokens at 50 ms shows its first answer token after about 75 s plus TTFT. For agents and reasoning models the gap between tokens weighs more than for chat, because the intermediate tokens go unread and each one adds to the wait. Our guide to GPU sizing for AI agents counts the calls and tokens per task.
We size inference servers by model size and concurrent users, with configuration and quote within one business day. Send us your example targets, the model and the peak requests in flight through the form below.
How to measure LLM latency with vllm bench serve and AIPerf
A load test shows whether one GPU meets the targets at a given concurrency. vllm bench serve ships with vLLM; --percentile-metrics defaults to ttft, tpot and itl, and --metric-percentiles defaults to the 99th only. --goodput takes targets as metric and value pairs in milliseconds and counts the requests per second that meet all of them. AIPerf installs with pip install aiperf, release 0.13.0 of 24 September 2026, and runs as aiperf profile with --streaming for TTFT and ITL and --concurrency for a fixed number of requests in flight. GenAI-Perf’s documentation states that it “is being phased out” and tells users “For new performance benchmarking needs, please use AIPerf instead”. Our guide to what a private AI pilot should measure covers the datasets and flags for replaying your own prompts.
- Write each target as metric, percentile and value, and say whether it is measured at the client or in the server.
- Replay prompts with the input and output lengths of your use case; random 1,024-token prompts with 128-token answers, vLLM’s default, fit no use case above.
- Fix the number of requests in flight with
--max-concurrencyand step it up, 1, 2, 4, 8 and so on; without it vLLM’s default request rate of inf sends all requests at time 0. - Add
--goodput ttft:1000 tpot:100and--percentile-metrics ttft,tpot,itl,e2elwith--metric-percentiles 50,95,99. - Take the highest concurrency at which the 95th percentiles meet the targets and goodput stays close to request throughput, and record the GPU, engine version and settings with it.
vLLM’s goodput checks TPOT, the per-request average, so a request with one long stall can still count as good; read the ITL percentiles beside it. NVIDIA’s guide says to “prefer concurrency for most benchmarks”, because with a fixed request rate outstanding requests “can grow without bound” once arrivals exceed what the system serves. Results carry over only to the same GPU type and engine settings.
In our Private AI/ML service, the technical assessment produces a pilot plan with metrics, at a price fixed before work begins. Describe in the form below the use case you would measure first and the targets you have in mind.
What we supply
We build AI servers to order for inference and RAG, with 2 to 8 GPUs per node sized by model size and concurrent users, and we supply the cards on their own for servers you already run. For latency targets the choice is between bandwidth and memory: the H200 NVL with 141 GB of HBM3e at 4.8 TB/s, the RTX PRO 6000 Blackwell in all three editions with 96 GB, and the L40S and L4 for smaller models. Cards and servers come with manufacturer warranty, and NVIDIA AI Enterprise licences come on the same invoice as the hardware, on one EU contract. If you want the targets and a pilot plan with metrics worked out for you, that is our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What is time to first token in LLM inference?
What is the difference between TTFT and TPOT?
How many tokens per second per user does an LLM chat need?
What are LLM inference latency requirements for chat, RAG and agents?
Which metrics describe LLM inference performance?
How do I measure LLM latency on my own server?
Send us the use case, the model and its precision, the prompt and answer lengths, the peak requests in flight and the latency targets you have in mind. We reply within one business day with a configuration and quote for the GPU server those figures point to.
Talk to an expertWe reply within one business day