BLOG · GUIDE ·

LLM monitoring in production: vLLM metrics in Prometheus, alerts on queue, KV cache and TTFT

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • LLM monitoring on vLLM starts with its Prometheus metrics at /metrics on the API port, on by default: requests waiting and running, KV cache usage, preemptions, time to first token, inter-token and end-to-end latency, token counters and finish reasons
  • In vLLM 0.31.0 of 5 October 2026, vllm:kv_cache_usage_perc has replaced vllm:gpu_cache_usage_perc, deprecated in 0.10.0, and vllm:inter_token_latency_seconds has replaced vllm:time_per_output_token_seconds, deprecated in 0.10.2; dashboards built on the old names show no data
  • NVIDIA NIM for LLMs passes the same vLLM metrics through unchanged at /v1/metrics, and vLLM’s --api-key does not protect /metrics, so only the gateway and the components that read its metrics should reach the port
  • Our example alerts fire on a queue that lasts 10 minutes, KV cache above 0.9 for 15 minutes, any preemption, and fewer than 95 per cent of requests within a TTFT target placed on one of vLLM’s bucket edges, such as 1 or 2.5 s
  • Prompts and answers belong in a query log at the gateway, with restricted access and a set retention, never in metric labels; OpenTelemetry’s GenAI conventions make content attributes in traces Opt-In

Eurokommerz × Vixen.UNO: Private AI/ML  Talk to an expert →

What to monitor on a self-hosted LLM service

LLM monitoring for a self-hosted model watches what the serving engine reports about its queue, KV cache and latency, and raises an alert before users notice a slowdown. In vLLM these figures are Prometheus metrics at /metrics on the API port, 8000 by default, and vLLM’s Prometheus and Grafana example states that “Prometheus metric logging is enabled by default in the OpenAI-compatible server”. Requests waiting, KV cache usage, preemptions, time to first token (TTFT) and inter-token latency drive the main alerts; running requests, token counters and finish reasons add the trend data for capacity planning.

The gateway in front of vLLM adds users, HTTP errors and a time to first token close to what a user sees. GPU health comes from DCGM; our guide to GPU server monitoring with DCGM covers it and explains why GPU utilisation says little about a serving GPU’s load. TTFT, inter-token latency and time per output token (TPOT) are defined in our guide to what a private AI pilot should measure, which also sizes hardware from them.

vLLM metrics for Prometheus in vLLM 0.31.0

The names below come from the source code of vLLM 0.31.0, published on 5 October 2026 and current as of 6 October. The metrics in the table have the labels model_name and engine, the second numbering the engine cores under data parallelism. Counters end in _total, and histograms add _bucket, _sum and _count.

SIGNALVLLM 0.31.0 METRICWHAT IT TELLS YOUEXAMPLE ALERT (OURS)
Waiting requestsvllm:num_requests_waitingrequests in the waiting queue, preempted ones includedabove 0 for 10 minutes
Running requestsvllm:num_requests_runningrequests in the running batchnone; read it beside the queue
KV cache usagevllm:kv_cache_usage_percshare of the KV cache in use, 1 meaning fullabove 0.9 for 15 minutes
Preemptionsvllm:num_preemptions_totaltimes a request went back to the queue for lack of KV cache, to restart its prefillany increase within 15 minutes
Time to first tokenvllm:time_to_first_token_secondsqueue wait plus prompt processing inside vLLMunder 95 per cent of requests within target for 10 minutes
Inter-token latencyvllm:inter_token_latency_secondsgaps between output steps; stalls show heremore than 1 per cent of gaps above 0.5 s for 10 minutes
End-to-end latencyvllm:e2e_request_latency_secondsthe whole request inside vLLM99th percentile near the gateway’s timeout
Finish reasonsvllm:request_success_totalfinished requests by finished_reason, where length means max_tokens or max_model_len was reachedlength above 5 per cent of requests in an hour

vLLM 0.31.0 source code (metric definitions and finish reasons), released 5 October 2026. The alert values are our examples, not vLLM’s; test them against your own traffic.

Give each rule a for duration, which makes Prometheus wait “for a certain duration” before an alert fires, so that short bursts do not page anyone. Gauges such as the queue are read only at each scrape, every 5 s in vLLM’s own example and every minute by Prometheus’s default. Add a rule on Prometheus’s up series, 0 when a scrape fails, because a crashed server exports nothing and the other rules then have no data.

NVIDIA’s NIM for LLMs exposes the metrics of its vLLM backend at /v1/metrics, and its documentation says that “NIM passes through the inference backend’s native Prometheus metrics without modification”. Dashboards and rules carry over, within the vLLM version a NIM release contains. Other engines use names of their own; our comparison of vLLM, SGLang, TensorRT-LLM and Ollama shows how each exposes metrics.

vLLM’s security guide says that --api-key protects only endpoints under /v1, /v2, /inference and /cohere, so /metrics, outside those paths on the same port, answers without a key. Let only the gateway, Prometheus and, on Kubernetes, the endpoint picker reach that port, as our guide to a private LLM platform on Kubernetes sets out; it also covers autoscaling on these metrics.

Renamed metrics and the vLLM Grafana dashboards

vLLM deprecated its metrics with a gpu_ prefix in release 0.10.0 and hid them from 0.11.0. The KV cache gauge vllm:gpu_cache_usage_perc became vllm:kv_cache_usage_perc, and the prefix cache counters became vllm:prefix_cache_queries and vllm:prefix_cache_hits, which appear with the suffix _total. Release 0.10.2 deprecated vllm:time_per_output_token_seconds in favour of vllm:inter_token_latency_seconds, since what it measured was, in the pull request’s words, “the time between iterations”. We found none of these old names in the metric definitions of 0.31.0, so panels built on them show no data. Under vLLM’s deprecation policy, --show-hidden-metrics-for-version restores a deprecated metric only in the release that hides it, before the next one removes it.

vLLM’s Prometheus and Grafana example includes grafana.json, with panels for latency percentiles from the 50th to the 99th, token throughput, running and waiting requests, KV cache usage, prompt and answer lengths, finish reasons and queue, prefill and decode time. Its observability examples add Performance Statistics and Query Statistics dashboards, as JSON for Grafana and YAML for Perses. The set-up takes three steps.

  1. Add each vLLM server to Prometheus as a scrape target on its API port, path /metrics, or /v1/metrics for NIM, with a scrape interval of a few seconds.
  2. Import grafana.json, select the Prometheus data source and pick the model in the dashboard’s model_name variable.
  3. Add the alert rules from the table, each with a for duration, and before each vLLM upgrade check its release notes for deprecated metrics.

Monitoring time to first token in production

vLLM’s TTFT histogram starts when its frontend receives a request. When it rises, the queue time, vllm:request_queue_time_seconds, and the prefill time, vllm:request_prefill_time_seconds, show which part grew; if neither grew, look at the frontend. In vLLM’s metrics design document, the first runs from the engine core queueing the request to its most recent scheduling, the second from that scheduling to the first new token. A longer queue time means missing capacity; a longer prefill time with a flat queue points to more prefill work: longer prompts, which the vllm:request_prompt_tokens histogram shows, or fewer prefix cache hits. Both time histograms start their buckets at 0.3 s, so for shorter waits use the mean, the rate of _sum divided by the rate of _count.

For a TTFT target, alert on the share of requests within it rather than on a percentile. Prometheus’s documentation notes that with an SLO known in advance, “you could use the fixed bucket boundaries of a classic histogram to allow an accurate calculation”. vLLM’s TTFT buckets include 0.25, 0.5, 0.75, 1, 2.5 and 5 s, so put the target on one of these edges. The rate of vllm:time_to_first_token_seconds_bucket with le="1.0", divided by the rate of vllm:time_to_first_token_seconds_count, each summed by model_name, is the share of requests whose first token came within 1 s. Without the sums the result is empty, as PromQL pairs only series with identical labels and only the bucket has le; the share of answers that ended on length needs the same sums. A 2 s target falls between two edges, and any figure for it is interpolated.

Per-token speed works the same way with the per-request histogram vllm:request_time_per_output_token_seconds, whose bucket at 0.1 s counts the requests that averaged 10 tokens per second or faster after the first. The gateway should also record its own time to first token, network and proxies included, which OpenTelemetry’s GenAI conventions call gen_ai.client.operation.time_to_first_chunk; that figure is closer to what users notice.

Capacity signals for adding a GPU

Before buying, read the same metrics at the busiest hour of each working day, not as daily averages, and match the pattern: only two of the five below call for more GPU memory or compute.

PATTERNLIKELY LIMITFIRST RESPONSE
Queue, cache near fullKV cache memory; preemptions rise tooan FP8 KV cache if answer quality holds, then more GPU memory: a second card or another replica
Queue, cache has roomthe scheduler’s limits, max_num_seqs or max_num_batched_tokensraise the limit and watch inter-token latency, which grows with the batch
TTFT up, queue time flatlonger prompts, more prefill workprompt lengths and prefix cache hits; fewer or shorter retrieved passages
Slow tokens, no queuea batch too large for the per-user speed targetcap max_num_seqs per replica so each stream meets the target, then add a GPU
Waiting, reason deferredLoRA budget, KV transfer or blocked status in vllm:num_requests_waiting_by_reasonadapter limits and connector health, not hardware

Our reading of the vLLM 0.31.0 metric definitions and of vLLM’s optimisation guide; the responses are suggestions to test, not vLLM’s instructions.

For frequent preemptions, vLLM’s optimisation guide lists remedies from a higher gpu_memory_utilization (0.92 by default) to more tensor or pipeline parallelism, and warns that preemption and recomputation “can adversely affect end-to-end latency”. An FP8 KV cache (--kv-cache-dtype fp8) stores one byte per value where a 16-bit cache stores two, so the same memory holds twice the tokens; run your evaluation set first. As our suggestion, plan the next GPU when the first or fourth pattern appears at the peak hour on most working days of a month, and size it from peak token throughput and peak running plus waiting requests.

When the signals point to more hardware, we select and supply the GPU servers and calculate the TCO against cloud GPUs before the purchase. Send us the queue, KV cache and latency figures from your busiest hours through the form below.

Logging prompts and answers without leaking them

Serving metrics carry no request content, and custom labels should not add any. vLLM’s labels name the model, the engine and reasons such as the finish reason, not users, and Prometheus warns that “every unique combination of key-value label pairs represents a new time series”, advising against labels such as user IDs. Usage per user or department goes into the query log.

Keep the query log in the gateway or chat front end, the only component that knows who asked. Record the user, time, model, token counts, finish reason and a request ID, and the text of prompts and answers where your policy requires it, in a store that only the people your policy names can read, with a retention period set before the first entry. A model server set to log requests for debugging can write prompts into its application log and every log platform that collects it; keep that to test systems. How long the text may be kept is a legal assessment for your legal department. Our guide to a private ChatGPT alternative shows where a chat front end’s own audit log falls short.

vLLM sends OpenTelemetry traces when --otlp-traces-endpoint is set; in its example, vLLM’s span holds “metadata about the request” and the prompt text sits in the client’s span. OpenTelemetry’s GenAI conventions make the attributes for prompts, answers and system instructions Opt-In, and note that prompts and answers are “likely to contain sensitive information including user/PII data”. Traces sent to an external service take any recorded content out of the company. NIM accepts an X-Request-Id header as a correlation identifier and forwards the W3C traceparent header, from which its backend creates spans when OTEL_EXPORTER_OTLP_TRACES_ENDPOINT is set, so a query log entry and its trace can share one ID.

In our private AI/ML service, queries and answers are logged so that security and legal see who accesses what, and how. Describe in the form below who should read that log and how long you plan to keep it.

OpenTelemetry GenAI conventions for LLM observability

Observability adds the traces and logs that explain a single slow or wrong answer, and OpenTelemetry’s semantic conventions for generative AI give them common names. Spans carry gen_ai.operation.name and gen_ai.provider.name as required attributes and token counts in gen_ai.usage.input_tokens and gen_ai.usage.output_tokens. The metrics include gen_ai.client.operation.duration and gen_ai.client.token.usage for clients and gen_ai.server.time_to_first_token for servers. As of October 2026 their status is Development, and opentelemetry.io, at semantic conventions 1.44.0, says they “have moved” to a repository of their own.

vLLM keeps its own Prometheus names and uses OpenTelemetry for tracing. Its request span, llm_request, counts tokens in gen_ai.usage.prompt_tokens and gen_ai.usage.completion_tokens instead of the conventions’ input and output names. The option --collect-detailed-traces collects detailed traces for the modules model, worker or all, which, its help text warns, “might have a performance impact”. Record the client metrics in the gateway; their recommended buckets for time to first chunk double from 10 ms and do not match vLLM’s, so set explicit bucket boundaries there that put your TTFT target on an edge.

What we do

Our Private AI/ML service deploys models on-premise with vLLM, Ollama or NVIDIA AI Enterprise and builds the platform around them. You get a query log, data and permissions management and a team trained to run the platform, with ongoing support under an agreed SLA if you want it, and we supply the GPU servers when the metrics call for more capacity. Eurokommerz holds the contract, with engineering by our partner Vixen.UNO; the first call is free of charge, and the price of the technical assessment is fixed before work begins. Data handling during a project is described on our security and compliance page.

FAQ

What should LLM monitoring cover for a self-hosted model?
It should cover the serving engine’s queue, KV cache and latency: requests waiting and running, KV cache usage, preemptions, time to first token and inter-token latency, plus token counters and finish reasons for trends. The gateway adds users, HTTP errors and a time to first token close to what users see, and DCGM adds GPU health.
Which vLLM metrics should Prometheus scrape?
Scrape all of /metrics on the API port; in vLLM 0.31.0 our alert examples use vllm:num_requests_waiting, vllm:kv_cache_usage_perc, vllm:num_preemptions_total and the histograms vllm:time_to_first_token_seconds and vllm:inter_token_latency_seconds. vLLM’s own example scrapes every 5 seconds, while Prometheus defaults to 1 minute, which can miss a short queue.
Is there a Grafana dashboard for vLLM?
vLLM’s repository ships grafana.json with its Prometheus and Grafana example, with panels for latency percentiles, token throughput, running and waiting requests, KV cache usage, finish reasons and queue time, and a second set of Performance Statistics and Query Statistics dashboards for Grafana and Perses. Dashboards built for older releases may still query vllm:gpu_cache_usage_perc, which current releases no longer export.
How do you monitor time to first token in production?
Alert on the share of requests whose first token arrives within the target: the rate of the vllm:time_to_first_token_seconds bucket at the target divided by the rate of its count, both summed by model_name, with the target on a bucket edge such as 1 or 2.5 seconds. When the share drops, compare vllm:request_queue_time_seconds with vllm:request_prefill_time_seconds to see whether waiting or more prefill work caused it. Measure the time to first token at the gateway as well, because vLLM’s figure leaves out the network.
What replaced vllm:gpu_cache_usage_perc in vLLM?
vLLM replaced it with vllm:kv_cache_usage_perc, where 1 means a full cache, after deprecating its metrics with a gpu_ prefix in 0.10.0 and hiding them from 0.11.0; vllm:prefix_cache_queries and vllm:prefix_cache_hits replaced the gpu_ prefix cache counters. The old names are not in vLLM 0.31.0, so the --show-hidden-metrics-for-version option cannot bring them back.
Should LLM observability record prompts and answers?
Record them only in a query log with restricted access and a retention period that your legal department has assessed, kept at the gateway that knows the user. OpenTelemetry’s GenAI conventions make the attributes for prompts and answers Opt-In and note that they are likely to contain sensitive information, and metric labels should carry neither user IDs nor text.

Send us the model server and version you run, how you monitor it today and the latency targets your users expect. We reply within one business day with next steps, starting with a first call that leaves you with two or three possible solution scenarios. The first call is free of charge.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry – see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna