Operating an LLM platform in production: LLMOps for updates, incidents, roles and on-call
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- LLMOps on a private platform keeps model checkpoints, the serving engine, GPU drivers, Kubernetes and the RAG index current without changing answers unnoticed, and runs capacity reviews, incidents, access reviews and log retention around them
- vLLM aims at a regular release every 2 weeks (v0.30.0 on 22 September, v0.31.0 on 5 October 2026) and guarantees backwards compatibility only for a limited number of minor releases, so every skipped version’s notes need reading
- NVIDIA’s driver table of 9 September 2026 lists R580 as a Long Term Support Branch to June 2028 and R595 as a Production Branch to March 2027; a production branch gets bug and security fixes for up to one year
- Each model or engine change passes the evaluation set, then serves a small weighted share of traffic, for example 10 per cent, while the old version keeps running, so a rollback is one routing change
- Four owners share the work: platform, model, data and security; Google’s SRE book sets a minimum of eight engineers on one site for on-call cover at all hours with a primary and a secondary, so a mid-size company decides which hours need cover and who provides it
Eurokommerz × Vixen.UNO: Private AI/ML Talk to an expert →
What LLMOps covers on a private LLM platform
LLMOps is the day-2 work of operating a private LLM platform in production: it keeps models, the serving engine, GPU drivers and the RAG index current without breaking the answers people rely on, and it covers capacity, incidents, access reviews and log retention. After the build, every component has its own release calendar, and each change has to be tested, rolled out in steps and reversible.
Take a platform for 500 to 2,000 employees on two to four GPU servers. It typically runs a chat model, an embedding model and a reranker on vLLM, a gateway with keys and a query log, a chat front end, a vector index fed from document sources, and Kubernetes with NVIDIA’s GPU Operator. That is about ten components, each updated at its own pace. The work is to change one at a time and to know, before users notice, whether answers stayed as good as they were.
LLMOps vs MLOps: serving models you did not train
MLOps grew around models a company trains itself: data pipelines, training runs, a model registry and retraining when the data drifts. On an LLM platform the model usually arrives as a checkpoint from its publisher. The levers you control are which checkpoint and precision you serve, the system prompt and chat template, the retrieval pipeline and the engine settings, and each of them can change answers as much as a new model can.
Quality is therefore measured on a fixed set of questions with expected answers or sources, run before every change. Where a company also fine-tunes, that part follows MLOps practice, and the result enters the platform like any other new checkpoint.
Release cadences of vLLM, NVIDIA drivers and models
vLLM’s release document says “We aim to have a regular release every 2 weeks”, and since v0.12.0 each regular release increments the minor version. v0.30.0 came out on 22 September 2026 and v0.31.0 on 5 October. The same document warns that “backwards compatibility is only guaranteed for a limited number of minor releases”. Under vLLM’s deprecation policy a feature is first deprecated, then switched off by default and then removed, across several minor releases, and the warning names the version that removes it. The breaking changes of v0.31.0 include the removal of tokenizer_mode="slow" and a renamed Mamba prefix cache option, so a start script that still passes either has to be changed before the upgrade.
NVIDIA’s driver lifecycle page, updated on 9 September 2026, says that two production branches are released per year, each with bug fixes and security updates “for up to 1 year”. A Long Term Support Branch receives quarterly, or as needed, bug and security releases for three years. NVIDIA’s table of supported drivers, of the same date, lists R580 as Long Term Support to June 2028 and R595 as Production to March 2027. Our guide to NVIDIA driver branches and CUDA versions explains how to pin a branch and which CUDA versions each one runs.
Model publishers release refreshed checkpoints when they are ready, often as a new repository. Qwen, for example, introduced Qwen3-235B-A22B-Instruct-2507 as “the updated version” of its Qwen3-235B-A22B non-thinking mode. vLLM’s --revision takes “a branch name, a tag name, or a commit id”, so pin the commit you evaluated. Give applications one stable model name and map it to the version each vLLM instance serves. vLLM answers to every name given in --served-model-name and labels its Prometheus metrics with the first. Start each version with its own name first and the stable name second, so it answers the requests applications send and stays apart on a dashboard.
Day-2 tasks, frequency and owners
Four roles share the platform work. The platform owner, usually in infrastructure, runs the GPU servers, drivers, Kubernetes and the serving engine. The model owner chooses checkpoints, keeps the evaluation set and signs off each change that can alter answers. A data owner for each RAG source decides which documents go in, which directory groups may read them and when they are removed. Security owns access reviews, the query log and vulnerability notices. One person can hold two roles.
| TASK | FREQUENCY, EXAMPLE | OWNER |
|---|---|---|
| GPU driver update | quarterly within the branch, node by node; a branch change before the branch ends | platform owner |
| Serving engine upgrade | every second or third vLLM minor release, sooner for a security fix | platform owner, model owner signs off evaluation |
| Model refresh or change | when a publisher releases a checkpoint worth testing | model owner |
| RAG index sync | daily incremental sync; full re-embedding when the embedding model changes | data owners, platform owner |
| Capacity review | monthly, from peak-hour queue and KV cache figures | platform owner |
| Access review | quarterly and when people change role | security, data owners |
| Query log retention | automated deletion, checked monthly | security |
| Restore test | twice a year: model store, vector index, configuration | platform owner |
| Incident review | after every significant incident | platform owner with all owners |
Frequencies and owners are our examples for a platform of 500 to 2,000 users, not vendor requirements. Release facts from vLLM’s RELEASE.md and NVIDIA’s driver lifecycle page of 9 September 2026.
An index built with one embedding model cannot be searched with query vectors from another. Qdrant’s migration guide says “Switching models requires re-embedding all vectors in your collection”, builds the new collection beside the old one and switches to it once it performs at least as well. The figures for the capacity review come from the serving metrics our guide to LLM monitoring with vLLM metrics defines.
Updating LLM models in production: evaluation, canary and rollback
A model change, an engine upgrade and a new chat template follow the same procedure. Our guide to LLM evaluation before model upgrades covers the evaluation step and its release gates.
- Write down what changes, the pinned revision or version, and the old state you will return to.
- Run the evaluation set against the new version on a test instance and compare it with the current version’s results.
- Deploy the new version beside the old one, with its own served model name first and the stable name second.
- Send a small share of traffic to it, or only a pilot group, and watch errors, latency and user feedback.
- Move the remaining traffic in steps, and keep the old version running until the change is closed.
- Roll back by setting the new version’s share to zero, then record the result in the change log.
The Kubernetes Gateway API splits traffic by weights on an HTTPRoute, and its guide states that “weight indicates a proportional split of traffic (rather than percentage)”, with 90 and 10 as its canary example. Running both versions at once needs GPU memory for both, so plan the canary on a server that has room, or move one replica of the old version for the duration. For a change inside one Deployment, kubectl rollout undo returns to the previous revision, and Kubernetes keeps 10 old ReplicaSets for rollback by default. A rollout that makes no progress within progressDeadlineSeconds, 600 s by default, is considered failed and reports ProgressDeadlineExceeded; the controller keeps processing it and does not roll back by itself. A replica that loads a large checkpoint over the network can take longer than that, so set the deadline above the load time you see in tests.
Our Private AI/ML service builds and supports the platform and trains your team to run it. Tell us which models and engines you run and how changes reach production today.
Runbooks for common LLM platform incidents
Write a one-page runbook for each recurring symptom.
| SYMPTOM | FIRST CHECK | NEXT STEP |
|---|---|---|
| Out of memory at start | other processes on the GPU in nvidia-smi; context length and memory share set for vLLM | shorter context or smaller batch, a quantised checkpoint, or the model split across cards |
| Out of memory after upgrade | release notes for changed defaults; memory taken by CUDA graphs | roll back, then retest with fewer captured graphs or eager mode |
| Queue grows, nothing runs | health endpoint, engine log, running requests at zero | take the replica out at the gateway, collect debug logs, restart it |
| Queue grows, cache full | KV cache usage and preemptions at the peak hour | the capacity patterns in the monitoring guide |
| XID in the kernel log | the XID number in NVIDIA’s catalogue, DCGM health of the GPU | drain the node where the catalogue says so |
| Answers changed overnight | model revision, chat template, index sync log, prompt changes | return to the pinned revision and rerun the evaluation set |
| Replica slow to start | model loaded over the network or downloaded at start | keep checkpoints on a local model store |
First checks from vLLM’s troubleshooting and memory guides (September 2026) and NVIDIA’s XID catalogue; the next steps are our suggestions.
vLLM’s memory guide lists tensor parallelism, quantised models, a shorter max_model_len, a smaller max_num_seqs and fewer CUDA graphs as ways to reduce memory, and the enforce_eager flag disables graph capture completely. For an instance that stops making progress, its troubleshooting guide suggests VLLM_LOGGING_LEVEL=DEBUG, CUDA_LAUNCH_BLOCKING=1 and, for multi-GPU communication, NCCL_DEBUG=TRACE. Use them on a replica that carries no user traffic. Our guide to GPU server monitoring with DCGM explains which XID messages call for a reset or a drained node. The SRE book asks for postmortems “after significant incidents” with a full timeline, and each one should end with a changed runbook or alert.
SLOs and on-call for an internal LLM service
The SRE book defines an SLO as “a target value or range of values for a service level that is measured by an SLI”, and adds that insisting on meeting SLOs 100 per cent of the time is “both unrealistic and undesirable”. It prefers percentiles to averages and treats the allowed shortfall as an error budget. For an internal assistant, three SLOs cover most of what users notice. Our example targets are 99.5 per cent of requests answered without error during business hours each month; 95 per cent of first tokens within 2.5 s, measured at the gateway; and no release whose evaluation results fall below those of the version it replaces. When the error budget for the month is spent, changes wait until the cause is fixed.
On-call follows from these targets. Google’s SRE book sizes on-call cover at all hours, with a primary and a secondary on call, at a minimum of eight engineers on one site, or six per site for two sites, and caps on-call at 25 per cent of an engineer’s time. For a platform run by two or three people, decide which hours the assistant must be available, whether a failure out of hours waits for the next working day, and who covers which part, with any support agreement stating its hours and response times. Our guide to LLM high availability on two GPU nodes covers the redundancy that makes calls rarer.
After the build, your own staff run the platform or we support it under an agreed SLA. Tell us who runs your platform today and which support you want from a supplier.
Access reviews and query log retention
Access to models and RAG sources should follow directory groups, so that a review means checking group membership against the data owners’ decisions. Implementing Regulation (EU) 2024/2690, which binds the digital infrastructure and digital service providers listed in its Article 1, requires in point 11.2.3 of its Annex that they “review access rights at planned intervals” and document the results. For other companies it is a usable reference. Review service accounts and gateway keys in the same cycle.
The query log records who asked what and which sources were used, so it contains personal data. Under Article 5(1)(e) of the GDPR, personal data are kept in a form that identifies people “for no longer than is necessary” for their purpose, and point 3.2.5 of the same Annex asks entities in its scope to “maintain and back up logs for a predefined period” and protect them from unauthorised access or changes. Set the retention period and the people who may read the log before the first entry, and automate deletion. How long the text may be kept is a legal assessment for your legal department.
What we do
Our Private AI/ML service deploys open and commercial models on-premise with vLLM, Ollama or NVIDIA AI Enterprise, on a Kubernetes-based platform where it fits, with logging of queries and answers and data and permissions management. We build and support the platform and train your team to run it and develop it further; afterwards it is your choice whether your own staff run it or we support it under an agreed SLA. Eurokommerz holds the contract and supplies the GPU servers, with engineering by our partner Vixen.UNO. The first call is free of charge, and the price of the technical assessment is fixed before work begins. How we handle data during a project is described on our security and compliance page.
FAQ
What is LLMOps?
What is the difference between MLOps and LLMOps?
How do you update an LLM model in production?
How often does vLLM release new versions?
Which roles does an LLM platform team need?
Does a private LLM platform need on-call?
Send us the models, serving engines and GPU servers you run, the number of users and who operates the platform today. We reply within one business day, and in the first call we work through your process and data with you, so you leave with 2 to 3 possible solution scenarios. The first call is free of charge.
Talk to an expertWe reply within one business day