LLM monitoring in production: vLLM metrics in Prometheus, alerts on queue, KV cache and TTFT
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- LLM monitoring on vLLM starts with its Prometheus metrics at /metrics on the API port, on by default: requests waiting and running, KV cache usage, preemptions, time to first token, inter-token and end-to-end latency, token counters and finish reasons
- In vLLM 0.31.0 of 5 October 2026, vllm:kv_cache_usage_perc has replaced vllm:gpu_cache_usage_perc, deprecated in 0.10.0, and vllm:
inter_ token_ latency_ seconds has replaced vllm: time_ per_ output_ token_ seconds, deprecated in 0.10.2; dashboards built on the old names show no data - NVIDIA NIM for LLMs passes the same vLLM metrics through unchanged at /v1/metrics, and vLLM’s --api-key does not protect /metrics, so only the gateway and the components that read its metrics should reach the port
- Our example alerts fire on a queue that lasts 10 minutes, KV cache above 0.9 for 15 minutes, any preemption, and fewer than 95 per cent of requests within a TTFT target placed on one of vLLM’s bucket edges, such as 1 or 2.5 s
- Prompts and answers belong in a query log at the gateway, with restricted access and a set retention, never in metric labels; OpenTelemetry’s GenAI conventions make content attributes in traces Opt-In
Eurokommerz × Vixen.UNO: Private AI/ML Talk to an expert →
What to monitor on a self-hosted LLM service
LLM monitoring for a self-hosted model watches what the serving engine reports about its queue, KV cache and latency, and raises an alert before users notice a slowdown. In vLLM these figures are Prometheus metrics at /metrics on the API port, 8000 by default, and vLLM’s Prometheus and Grafana example states that “Prometheus metric logging is enabled by default in the OpenAI-compatible server”. Requests waiting, KV cache usage, preemptions, time to first token (TTFT) and inter-token latency drive the main alerts; running requests, token counters and finish reasons add the trend data for capacity planning.
The gateway in front of vLLM adds users, HTTP errors and a time to first token close to what a user sees. GPU health comes from DCGM; our guide to GPU server monitoring with DCGM covers it and explains why GPU utilisation says little about a serving GPU’s load. TTFT, inter-token latency and time per output token (TPOT) are defined in our guide to what a private AI pilot should measure, which also sizes hardware from them.
vLLM metrics for Prometheus in vLLM 0.31.0
The names below come from the source code of vLLM 0.31.0, published on 5 October 2026 and current as of 6 October. The metrics in the table have the labels model_name and engine, the second numbering the engine cores under data parallelism. Counters end in _total, and histograms add _bucket, _sum and _count.
| SIGNAL | VLLM 0.31.0 METRIC | WHAT IT TELLS YOU | EXAMPLE ALERT (OURS) |
|---|---|---|---|
| Waiting requests | vllm:num_ | requests in the waiting queue, preempted ones included | above 0 for 10 minutes |
| Running requests | vllm:num_ | requests in the running batch | none; read it beside the queue |
| KV cache usage | vllm:kv_ | share of the KV cache in use, 1 meaning full | above 0.9 for 15 minutes |
| Preemptions | vllm:num_ | times a request went back to the queue for lack of KV cache, to restart its prefill | any increase within 15 minutes |
| Time to first token | vllm:time_ | queue wait plus prompt processing inside vLLM | under 95 per cent of requests within target for 10 minutes |
| Inter-token latency | vllm:inter_ | gaps between output steps; stalls show here | more than 1 per cent of gaps above 0.5 s for 10 minutes |
| End-to-end latency | vllm:e2e_ | the whole request inside vLLM | 99th percentile near the gateway’s timeout |
| Finish reasons | vllm:request_ | finished requests by finished_reason, where length means max_tokens or max_model_len was reached | length above 5 per cent of requests in an hour |
vLLM 0.31.0 source code (metric definitions and finish reasons), released 5 October 2026. The alert values are our examples, not vLLM’s; test them against your own traffic.
Give each rule a for duration, which makes Prometheus wait “for a certain duration” before an alert fires, so that short bursts do not page anyone. Gauges such as the queue are read only at each scrape, every 5 s in vLLM’s own example and every minute by Prometheus’s default. Add a rule on Prometheus’s up series, 0 when a scrape fails, because a crashed server exports nothing and the other rules then have no data.
NVIDIA’s NIM for LLMs exposes the metrics of its vLLM backend at /v1/metrics, and its documentation says that “NIM passes through the inference backend’s native Prometheus metrics without modification”. Dashboards and rules carry over, within the vLLM version a NIM release contains. Other engines use names of their own; our comparison of vLLM, SGLang, TensorRT-LLM and Ollama shows how each exposes metrics.
vLLM’s security guide says that --api-key protects only endpoints under /v1, /v2, /inference and /cohere, so /metrics, outside those paths on the same port, answers without a key. Let only the gateway, Prometheus and, on Kubernetes, the endpoint picker reach that port, as our guide to a private LLM platform on Kubernetes sets out; it also covers autoscaling on these metrics.
Renamed metrics and the vLLM Grafana dashboards
vLLM deprecated its metrics with a gpu_ prefix in release 0.10.0 and hid them from 0.11.0. The KV cache gauge vllm:gpu_cache_usage_perc became vllm:kv_cache_usage_perc, and the prefix cache counters became vllm:prefix_cache_queries and vllm:prefix_cache_hits, which appear with the suffix _total. Release 0.10.2 deprecated vllm:--show-hidden-metrics-for-version restores a deprecated metric only in the release that hides it, before the next one removes it.
vLLM’s Prometheus and Grafana example includes grafana.json, with panels for latency percentiles from the 50th to the 99th, token throughput, running and waiting requests, KV cache usage, prompt and answer lengths, finish reasons and queue, prefill and decode time. Its observability examples add Performance Statistics and Query Statistics dashboards, as JSON for Grafana and YAML for Perses. The set-up takes three steps.
- Add each vLLM server to Prometheus as a scrape target on its API port, path /metrics, or /v1/metrics for NIM, with a scrape interval of a few seconds.
- Import grafana.json, select the Prometheus data source and pick the model in the dashboard’s model_name variable.
- Add the alert rules from the table, each with a
forduration, and before each vLLM upgrade check its release notes for deprecated metrics.
Monitoring time to first token in production
vLLM’s TTFT histogram starts when its frontend receives a request. When it rises, the queue time, vllm:
For a TTFT target, alert on the share of requests within it rather than on a percentile. Prometheus’s documentation notes that with an SLO known in advance, “you could use the fixed bucket boundaries of a classic histogram to allow an accurate calculation”. vLLM’s TTFT buckets include 0.25, 0.5, 0.75, 1, 2.5 and 5 s, so put the target on one of these edges. The rate of vllm:le="1.0", divided by the rate of vllm:
Per-token speed works the same way with the per-request histogram vllm:
Capacity signals for adding a GPU
Before buying, read the same metrics at the busiest hour of each working day, not as daily averages, and match the pattern: only two of the five below call for more GPU memory or compute.
| PATTERN | LIKELY LIMIT | FIRST RESPONSE |
|---|---|---|
| Queue, cache near full | KV cache memory; preemptions rise too | an FP8 KV cache if answer quality holds, then more GPU memory: a second card or another replica |
| Queue, cache has room | the scheduler’s limits, max_ or max_ | raise the limit and watch inter-token latency, which grows with the batch |
| TTFT up, queue time flat | longer prompts, more prefill work | prompt lengths and prefix cache hits; fewer or shorter retrieved passages |
| Slow tokens, no queue | a batch too large for the per-user speed target | cap max_ per replica so each stream meets the target, then add a GPU |
| Waiting, reason deferred | LoRA budget, KV transfer or blocked status in vllm:num_ | adapter limits and connector health, not hardware |
Our reading of the vLLM 0.31.0 metric definitions and of vLLM’s optimisation guide; the responses are suggestions to test, not vLLM’s instructions.
For frequent preemptions, vLLM’s optimisation guide lists remedies from a higher gpu_memory_utilization (0.92 by default) to more tensor or pipeline parallelism, and warns that preemption and recomputation “can adversely affect end-to-end latency”. An FP8 KV cache (--kv-cache-dtype fp8) stores one byte per value where a 16-bit cache stores two, so the same memory holds twice the tokens; run your evaluation set first. As our suggestion, plan the next GPU when the first or fourth pattern appears at the peak hour on most working days of a month, and size it from peak token throughput and peak running plus waiting requests.
When the signals point to more hardware, we select and supply the GPU servers and calculate the TCO against cloud GPUs before the purchase. Send us the queue, KV cache and latency figures from your busiest hours through the form below.
Logging prompts and answers without leaking them
Serving metrics carry no request content, and custom labels should not add any. vLLM’s labels name the model, the engine and reasons such as the finish reason, not users, and Prometheus warns that “every unique combination of key-value label pairs represents a new time series”, advising against labels such as user IDs. Usage per user or department goes into the query log.
Keep the query log in the gateway or chat front end, the only component that knows who asked. Record the user, time, model, token counts, finish reason and a request ID, and the text of prompts and answers where your policy requires it, in a store that only the people your policy names can read, with a retention period set before the first entry. A model server set to log requests for debugging can write prompts into its application log and every log platform that collects it; keep that to test systems. How long the text may be kept is a legal assessment for your legal department. Our guide to a private ChatGPT alternative shows where a chat front end’s own audit log falls short.
vLLM sends OpenTelemetry traces when --otlp-traces-endpoint is set; in its example, vLLM’s span holds “metadata about the request” and the prompt text sits in the client’s span. OpenTelemetry’s GenAI conventions make the attributes for prompts, answers and system instructions Opt-In, and note that prompts and answers are “likely to contain sensitive information including user/PII data”. Traces sent to an external service take any recorded content out of the company. NIM accepts an X-Request-Id header as a correlation identifier and forwards the W3C traceparent header, from which its backend creates spans when OTEL_
In our private AI/ML service, queries and answers are logged so that security and legal see who accesses what, and how. Describe in the form below who should read that log and how long you plan to keep it.
OpenTelemetry GenAI conventions for LLM observability
Observability adds the traces and logs that explain a single slow or wrong answer, and OpenTelemetry’s semantic conventions for generative AI give them common names. Spans carry gen_ai.operation.name and gen_ai.provider.name as required attributes and token counts in gen_ai.usage.input_tokens and gen_ai.usage.output_tokens. The metrics include gen_
vLLM keeps its own Prometheus names and uses OpenTelemetry for tracing. Its request span, llm_request, counts tokens in gen_ai.usage.prompt_tokens and gen_
What we do
Our Private AI/ML service deploys models on-premise with vLLM, Ollama or NVIDIA AI Enterprise and builds the platform around them. You get a query log, data and permissions management and a team trained to run the platform, with ongoing support under an agreed SLA if you want it, and we supply the GPU servers when the metrics call for more capacity. Eurokommerz holds the contract, with engineering by our partner Vixen.UNO; the first call is free of charge, and the price of the technical assessment is fixed before work begins. Data handling during a project is described on our security and compliance page.
FAQ
What should LLM monitoring cover for a self-hosted model?
Which vLLM metrics should Prometheus scrape?
Is there a Grafana dashboard for vLLM?
How do you monitor time to first token in production?
What replaced vllm:gpu_cache_usage_perc in vLLM?
Should LLM observability record prompts and answers?
Send us the model server and version you run, how you monitor it today and the latency targets your users expect. We reply within one business day with next steps, starting with a first call that leaves you with two or three possible solution scenarios. The first call is free of charge.
Talk to an expertWe reply within one business day