BLOG · GUIDE ·

LLM latency targets: time to first token and tokens per second for chat, RAG and agents

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • NVIDIA’s NIM benchmarking guide defines time to first token (TTFT) as the time from query submission to the first received token, including queuing, prefill and network latency, and inter-token latency (ITL) as the average gap between tokens, also known as time per output token (TPOT)
  • vllm bench serve reports TPOT as each request’s average gap after the first token and ITL as every single gap, so a figure called ITL means a different thing in vLLM than in AIPerf; note the tool with every number
  • TTFT is driven by the queue and by prefill, which is compute-bound and grows with prompt length; per-user tokens per second is driven by decode, which is bound by memory bandwidth and slows as the batch grows
  • Our example targets at the 95th percentile: chat 1 s to first token and a TPOT of 100 ms (10 tokens per second), RAG 2.5 s and 100 ms, agent steps 1 s and 50 ms, batch jobs a throughput target instead
  • The number of concurrent users one GPU serves depends on the latency target: NVIDIA’s guide says total throughput rises with concurrency while throughput per user falls, so measure goodput, the requests per second that meet every target

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

LLM latency metrics: TTFT, inter-token latency and throughput

LLM latency is described by four numbers: time to first token (TTFT), the gap between output tokens, called inter-token latency (ITL) or time per output token (TPOT), end-to-end latency per request, and throughput. TTFT is what a user waits before text appears; the gap between tokens sets how fast the text then streams. Both have to be stated with a percentile and with the tool that measured them, because tools compute them differently. For a chat assistant our example targets, derived below, are 1 s to the first token and 100 ms per token (10 tokens per second per user) at the 95th percentile.

NVIDIA’s NIM benchmarking guide, as read on 10 October 2026, defines TTFT as “the time from query submission to the first received token, if the response is not empty”, and says it “generally includes request queuing time, prefill time, and network latency”. It defines ITL as “the average time between consecutive tokens” and adds that it “is also known as time per output token (TPOT)”. End-to-end latency is TTFT plus generation time, the time from the first to the last token received.

METRICNVIDIA NIM GUIDEVLLM BENCH SERVEDRIVEN BY
Time to first tokensubmission to first received token; queuing, prefill and network includedttft, per request, at the clientqueue wait, prompt length, prefill compute
Inter-token latencyaverage gap between consecutive tokens, the same as TPOTitl, every gap between outputsmemory bandwidth, batch size, prefills sharing a step
Time per output tokenthe same metric as ITLtpot, each request’s average gap after the first tokenas ITL
End-to-end latencysubmission to the full response: TTFT plus generation timee2el, per requestTTFT, gap between tokens, answer length
Throughput per useroutput length divided by end-to-end latency, approaching 1/ITLnot reported separatelyas ITL
System throughputtotal tokens per second and requests per secondrequest and output token throughputbatch size, until compute saturates

NVIDIA NIM LLM benchmarking guide, metrics page; vLLM bench serve CLI reference and benchmark source; all read on 10 October 2026. The “driven by” column is our summary of NVIDIA’s inference optimisation blog of 17 November 2023 and vLLM’s optimisation guide.

Throughput per user and system throughput answer different questions. NVIDIA’s guide states that “total system TPS increases while TPS per user decreases as latency increases”. For sizing, a figure in tokens per second has to say which of the two it is.

TTFT vs TPOT and ITL: where vLLM and AIPerf differ

The definitions agree on TTFT and end-to-end latency and diverge on the gap between tokens. NVIDIA’s guide notes that “tools differ on whether TTFT is included in the average” and says to “compare results only when definitions align”.

AIPerf, the benchmarking tool NVIDIA now points to, computes ITL per request as end-to-end latency less TTFT, divided by the output tokens less one, and reports the gaps between all streamed chunks separately as inter chunk latency. Its output token throughput per user is 1 divided by that ITL. vllm bench serve computes the same per-request average and calls it TPOT, while its ITL is the list of every individual gap, pooled across requests. An ITL percentile from vLLM therefore shows stalls that an AIPerf ITL percentile, built from averages, smooths away.

Where the measurement starts matters too. Both tools time from the client, so their TTFT includes the network and, when the test runs through it, the gateway in front of the server. vLLM’s own server histogram starts inside the server and leaves both out, as our guide to LLM serving metrics in Prometheus explains; a target written for one cannot be checked against the other. For reasoning models AIPerf adds time to first output token, which it calculates as the time “from request start to the first non-reasoning output token”, the point at which the user sees the answer rather than the model’s thinking.

What drives time to first token and tokens per second

TTFT has two parts: waiting in the queue and prefill, the processing of the whole prompt before the first token. NVIDIA’s inference optimisation blog describes prefill as “a matrix-matrix operation that’s highly parallelized”, which “effectively saturates GPU utilization”, so its time grows with prompt length and falls with compute. NVIDIA’s benchmarking guide adds that “time to first token can increase due to queueing delay” and that longer input sequences increase TTFT. A RAG prompt with several retrieved passages costs more prefill than a chat question, and an agent that re-sends its growing context pays it again at every step.

The gap between tokens comes from decode, where each step reads the weights and the KV cache. The same NVIDIA blog calls this “a memory-bound operation” whose latency is dominated by the speed at which data is transferred from memory. For one user, tokens per second are capped by memory bandwidth divided by the bytes read per token. By that arithmetic, Llama 3.3 70B in FP8 tops out at about 23 tokens per second on an RTX PRO 6000 Server Edition and about 68 on an H200 NVL, as our RTX PRO 6000 vs H200 NVL comparison shows; these are upper bounds that no system reaches.

Prefill and decode share the GPU, so the scheduler trades one against the other. In vLLM, chunked prefill is on by default where possible, and the scheduler “prioritizes decode requests”. Its optimisation guide says that smaller values of max_num_batched_tokens “achieve better ITL because there are fewer prefills slowing down decodes”, while “higher values achieve better time to first token”. A server tuned for RAG prompts and a server tuned for fast streaming are configured differently.

Latency and concurrent users on one GPU

Concurrency raises total throughput and lowers per-user speed until the GPU saturates. NVIDIA’s parameters page states that “throughput generally saturates near the max batch size while latency steadily increases”. Past that point more concurrent requests add queue time to TTFT and give no extra throughput.

The number of users a card serves is therefore set by the latency target as well as by memory. Memory decides how many sessions fit in the KV cache; our guide to concurrent users per RTX PRO 6000 gives that arithmetic and the batch sweep from NVIDIA’s TensorRT-LLM tuning guide. The latency target decides how many of those sessions can run at once and still meet it. When the per-user target is above the single-stream ceiling, no batch size helps: a 40 ms target is 25 tokens per second, above the ceiling of about 23 for a 70B FP8 model on one RTX PRO 6000 Server Edition. NVFP4 weights, a smaller model, a card with more bandwidth, tensor parallelism across two cards or speculative decoding change that figure; more concurrency does not.

Example latency targets for chat, RAG, agents and batch jobs

Targets come from the use case, and the values below are our examples, not vendor recommendations. We placed them on edges of vLLM’s histogram buckets, so production monitoring can count exactly how many requests meet them. Our guide to users per RTX PRO 6000 puts mean silent reading at about 5.3 tokens per second, from a 2019 meta-analysis, so 10 tokens per second streams faster than the average reader. We state the per-token targets as TPOT, the per-request average that vLLM’s goodput checks and AIPerf reports as ITL.

USE CASEEXAMPLE TARGETSREQUEST SHAPEWHAT TO SIZE
Chat assistantTTFT 1 s, TPOT 100 ms, at p95short prompts, answers of a few hundred tokensKV cache for peak sessions, bandwidth per user
RAG question answeringTTFT 2.5 s, TPOT 100 ms, at p95prompts of several thousand tokensprefill compute, KV cache per long prompt
Agent stepTTFT 1 s, TPOT 50 ms, plus a time per taskmany calls, growing context, short outputsprefill per step, per-user speed
Reasoning modeltime to first output token, TPOT 50 mshundreds to thousands of thinking tokensper-user speed, answer length
Batch documentsdocuments per hour, end-to-end time per documentlong input, short outputsystem throughput

Example targets and request shapes are ours, at vLLM’s TTFT bucket edges of 1 and 2.5 s and its inter-token and TPOT bucket edges of 50 and 100 ms; test them with your users. NVIDIA’s NIM benchmarking guide lists retrieval under summarisation, with ISL near 1,000 and OSL near 100 tokens.

Turn the targets into end-to-end times before you agree them. A chat answer of 400 tokens with 1 s to the first token and 100 ms per token takes 1 + 399 × 0.1, about 41 s, of which the user reads along from the first second. An agent task of 8 steps, each with 1 s to the first token and 150 output tokens at 50 ms, takes 8 × (1 + 149 × 0.05), about 68 s before tool calls; at 100 ms per token it takes about 127 s. A reasoning model that thinks for 1,500 tokens at 50 ms shows its first answer token after about 75 s plus TTFT. For agents and reasoning models the gap between tokens weighs more than for chat, because the intermediate tokens go unread and each one adds to the wait. Our guide to GPU sizing for AI agents counts the calls and tokens per task.

We size inference servers by model size and concurrent users, with configuration and quote within one business day. Send us your example targets, the model and the peak requests in flight through the form below.

How to measure LLM latency with vllm bench serve and AIPerf

A load test shows whether one GPU meets the targets at a given concurrency. vllm bench serve ships with vLLM; --percentile-metrics defaults to ttft, tpot and itl, and --metric-percentiles defaults to the 99th only. --goodput takes targets as metric and value pairs in milliseconds and counts the requests per second that meet all of them. AIPerf installs with pip install aiperf, release 0.13.0 of 24 September 2026, and runs as aiperf profile with --streaming for TTFT and ITL and --concurrency for a fixed number of requests in flight. GenAI-Perf’s documentation states that it “is being phased out” and tells users “For new performance benchmarking needs, please use AIPerf instead”. Our guide to what a private AI pilot should measure covers the datasets and flags for replaying your own prompts.

  1. Write each target as metric, percentile and value, and say whether it is measured at the client or in the server.
  2. Replay prompts with the input and output lengths of your use case; random 1,024-token prompts with 128-token answers, vLLM’s default, fit no use case above.
  3. Fix the number of requests in flight with --max-concurrency and step it up, 1, 2, 4, 8 and so on; without it vLLM’s default request rate of inf sends all requests at time 0.
  4. Add --goodput ttft:1000 tpot:100 and --percentile-metrics ttft,tpot,itl,e2el with --metric-percentiles 50,95,99.
  5. Take the highest concurrency at which the 95th percentiles meet the targets and goodput stays close to request throughput, and record the GPU, engine version and settings with it.

vLLM’s goodput checks TPOT, the per-request average, so a request with one long stall can still count as good; read the ITL percentiles beside it. NVIDIA’s guide says to “prefer concurrency for most benchmarks”, because with a fixed request rate outstanding requests “can grow without bound” once arrivals exceed what the system serves. Results carry over only to the same GPU type and engine settings.

In our Private AI/ML service, the technical assessment produces a pilot plan with metrics, at a price fixed before work begins. Describe in the form below the use case you would measure first and the targets you have in mind.

What we supply

We build AI servers to order for inference and RAG, with 2 to 8 GPUs per node sized by model size and concurrent users, and we supply the cards on their own for servers you already run. For latency targets the choice is between bandwidth and memory: the H200 NVL with 141 GB of HBM3e at 4.8 TB/s, the RTX PRO 6000 Blackwell in all three editions with 96 GB, and the L40S and L4 for smaller models. Cards and servers come with manufacturer warranty, and NVIDIA AI Enterprise licences come on the same invoice as the hardware, on one EU contract. If you want the targets and a pilot plan with metrics worked out for you, that is our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What is time to first token in LLM inference?
Time to first token (TTFT) is the time from sending a request to receiving the first token of the answer. NVIDIA’s NIM benchmarking guide says it generally includes request queuing time, prefill time and network latency, so it rises with longer prompts and with a queue. vLLM’s server-side metric starts inside the server and leaves the network out.
What is the difference between TTFT and TPOT?
TTFT measures the wait until the first token appears, and time per output token (TPOT) measures the average gap between the tokens that follow. NVIDIA’s benchmarking guide treats TPOT and inter-token latency as one metric, while vllm bench serve reports TPOT as each request’s average gap and ITL as every single gap. TTFT is driven mainly by prefill and queueing, TPOT by memory bandwidth and batch size.
How many tokens per second per user does an LLM chat need?
Mean silent reading is about 5.3 tokens per second by the meta-analysis our users-per-GPU guide cites, so a stream at 10 tokens per second, 100 ms per token, stays ahead of readers. That is our example target for chat at the 95th percentile, not a vendor figure. Agents and reasoning models need faster generation, because every intermediate token adds to the wait.
What are LLM inference latency requirements for chat, RAG and agents?
Our examples at the 95th percentile are 1 s to the first token and a TPOT of 100 ms for chat, 2.5 s and 100 ms for RAG with long prompts, and 1 s and 50 ms per agent step plus a time budget per task. They sit on vLLM’s histogram bucket edges, so monitoring can count exactly how many requests meet them. Batch document jobs need a throughput target, such as documents per hour, instead of TTFT.
Which metrics describe LLM inference performance?
Time to first token, inter-token latency or time per output token, end-to-end latency, throughput per user and system throughput in tokens or requests per second. Goodput, the requests per second that meet every latency target, links them to a service level. Each figure needs its percentile, concurrency, prompt and answer lengths and the tool that produced it.
How do I measure LLM latency on my own server?
Run vllm bench serve or NVIDIA’s AIPerf against the server with prompts of your use case’s lengths, at fixed concurrency levels stepped up from one. Set goodput targets for TTFT and TPOT and read the 95th and 99th percentiles of TTFT and inter-token latency. NVIDIA’s documentation says GenAI-Perf is being phased out and points new benchmarking to AIPerf.

Send us the use case, the model and its precision, the prompt and answer lengths, the peak requests in flight and the latency targets you have in mind. We reply within one business day with a configuration and quote for the GPU server those figures point to.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna