BLOG · GUIDE ·

Operating an LLM platform in production: LLMOps for updates, incidents, roles and on-call

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • LLMOps on a private platform keeps model checkpoints, the serving engine, GPU drivers, Kubernetes and the RAG index current without changing answers unnoticed, and runs capacity reviews, incidents, access reviews and log retention around them
  • vLLM aims at a regular release every 2 weeks (v0.30.0 on 22 September, v0.31.0 on 5 October 2026) and guarantees backwards compatibility only for a limited number of minor releases, so every skipped version’s notes need reading
  • NVIDIA’s driver table of 9 September 2026 lists R580 as a Long Term Support Branch to June 2028 and R595 as a Production Branch to March 2027; a production branch gets bug and security fixes for up to one year
  • Each model or engine change passes the evaluation set, then serves a small weighted share of traffic, for example 10 per cent, while the old version keeps running, so a rollback is one routing change
  • Four owners share the work: platform, model, data and security; Google’s SRE book sets a minimum of eight engineers on one site for on-call cover at all hours with a primary and a secondary, so a mid-size company decides which hours need cover and who provides it

Eurokommerz × Vixen.UNO: Private AI/ML  Talk to an expert →

What LLMOps covers on a private LLM platform

LLMOps is the day-2 work of operating a private LLM platform in production: it keeps models, the serving engine, GPU drivers and the RAG index current without breaking the answers people rely on, and it covers capacity, incidents, access reviews and log retention. After the build, every component has its own release calendar, and each change has to be tested, rolled out in steps and reversible.

Take a platform for 500 to 2,000 employees on two to four GPU servers. It typically runs a chat model, an embedding model and a reranker on vLLM, a gateway with keys and a query log, a chat front end, a vector index fed from document sources, and Kubernetes with NVIDIA’s GPU Operator. That is about ten components, each updated at its own pace. The work is to change one at a time and to know, before users notice, whether answers stayed as good as they were.

LLMOps vs MLOps: serving models you did not train

MLOps grew around models a company trains itself: data pipelines, training runs, a model registry and retraining when the data drifts. On an LLM platform the model usually arrives as a checkpoint from its publisher. The levers you control are which checkpoint and precision you serve, the system prompt and chat template, the retrieval pipeline and the engine settings, and each of them can change answers as much as a new model can.

Quality is therefore measured on a fixed set of questions with expected answers or sources, run before every change. Where a company also fine-tunes, that part follows MLOps practice, and the result enters the platform like any other new checkpoint.

Release cadences of vLLM, NVIDIA drivers and models

vLLM’s release document says “We aim to have a regular release every 2 weeks”, and since v0.12.0 each regular release increments the minor version. v0.30.0 came out on 22 September 2026 and v0.31.0 on 5 October. The same document warns that “backwards compatibility is only guaranteed for a limited number of minor releases”. Under vLLM’s deprecation policy a feature is first deprecated, then switched off by default and then removed, across several minor releases, and the warning names the version that removes it. The breaking changes of v0.31.0 include the removal of tokenizer_mode="slow" and a renamed Mamba prefix cache option, so a start script that still passes either has to be changed before the upgrade.

NVIDIA’s driver lifecycle page, updated on 9 September 2026, says that two production branches are released per year, each with bug fixes and security updates “for up to 1 year”. A Long Term Support Branch receives quarterly, or as needed, bug and security releases for three years. NVIDIA’s table of supported drivers, of the same date, lists R580 as Long Term Support to June 2028 and R595 as Production to March 2027. Our guide to NVIDIA driver branches and CUDA versions explains how to pin a branch and which CUDA versions each one runs.

Model publishers release refreshed checkpoints when they are ready, often as a new repository. Qwen, for example, introduced Qwen3-235B-A22B-Instruct-2507 as “the updated version” of its Qwen3-235B-A22B non-thinking mode. vLLM’s --revision takes “a branch name, a tag name, or a commit id”, so pin the commit you evaluated. Give applications one stable model name and map it to the version each vLLM instance serves. vLLM answers to every name given in --served-model-name and labels its Prometheus metrics with the first. Start each version with its own name first and the stable name second, so it answers the requests applications send and stays apart on a dashboard.

Day-2 tasks, frequency and owners

Four roles share the platform work. The platform owner, usually in infrastructure, runs the GPU servers, drivers, Kubernetes and the serving engine. The model owner chooses checkpoints, keeps the evaluation set and signs off each change that can alter answers. A data owner for each RAG source decides which documents go in, which directory groups may read them and when they are removed. Security owns access reviews, the query log and vulnerability notices. One person can hold two roles.

TASKFREQUENCY, EXAMPLEOWNER
GPU driver updatequarterly within the branch, node by node; a branch change before the branch endsplatform owner
Serving engine upgradeevery second or third vLLM minor release, sooner for a security fixplatform owner, model owner signs off evaluation
Model refresh or changewhen a publisher releases a checkpoint worth testingmodel owner
RAG index syncdaily incremental sync; full re-embedding when the embedding model changesdata owners, platform owner
Capacity reviewmonthly, from peak-hour queue and KV cache figuresplatform owner
Access reviewquarterly and when people change rolesecurity, data owners
Query log retentionautomated deletion, checked monthlysecurity
Restore testtwice a year: model store, vector index, configurationplatform owner
Incident reviewafter every significant incidentplatform owner with all owners

Frequencies and owners are our examples for a platform of 500 to 2,000 users, not vendor requirements. Release facts from vLLM’s RELEASE.md and NVIDIA’s driver lifecycle page of 9 September 2026.

An index built with one embedding model cannot be searched with query vectors from another. Qdrant’s migration guide says “Switching models requires re-embedding all vectors in your collection”, builds the new collection beside the old one and switches to it once it performs at least as well. The figures for the capacity review come from the serving metrics our guide to LLM monitoring with vLLM metrics defines.

Updating LLM models in production: evaluation, canary and rollback

A model change, an engine upgrade and a new chat template follow the same procedure. Our guide to LLM evaluation before model upgrades covers the evaluation step and its release gates.

  1. Write down what changes, the pinned revision or version, and the old state you will return to.
  2. Run the evaluation set against the new version on a test instance and compare it with the current version’s results.
  3. Deploy the new version beside the old one, with its own served model name first and the stable name second.
  4. Send a small share of traffic to it, or only a pilot group, and watch errors, latency and user feedback.
  5. Move the remaining traffic in steps, and keep the old version running until the change is closed.
  6. Roll back by setting the new version’s share to zero, then record the result in the change log.

The Kubernetes Gateway API splits traffic by weights on an HTTPRoute, and its guide states that “weight indicates a proportional split of traffic (rather than percentage)”, with 90 and 10 as its canary example. Running both versions at once needs GPU memory for both, so plan the canary on a server that has room, or move one replica of the old version for the duration. For a change inside one Deployment, kubectl rollout undo returns to the previous revision, and Kubernetes keeps 10 old ReplicaSets for rollback by default. A rollout that makes no progress within progressDeadlineSeconds, 600 s by default, is considered failed and reports ProgressDeadlineExceeded; the controller keeps processing it and does not roll back by itself. A replica that loads a large checkpoint over the network can take longer than that, so set the deadline above the load time you see in tests.

Our Private AI/ML service builds and supports the platform and trains your team to run it. Tell us which models and engines you run and how changes reach production today.

Runbooks for common LLM platform incidents

Write a one-page runbook for each recurring symptom.

SYMPTOMFIRST CHECKNEXT STEP
Out of memory at startother processes on the GPU in nvidia-smi; context length and memory share set for vLLMshorter context or smaller batch, a quantised checkpoint, or the model split across cards
Out of memory after upgraderelease notes for changed defaults; memory taken by CUDA graphsroll back, then retest with fewer captured graphs or eager mode
Queue grows, nothing runshealth endpoint, engine log, running requests at zerotake the replica out at the gateway, collect debug logs, restart it
Queue grows, cache fullKV cache usage and preemptions at the peak hourthe capacity patterns in the monitoring guide
XID in the kernel logthe XID number in NVIDIA’s catalogue, DCGM health of the GPUdrain the node where the catalogue says so
Answers changed overnightmodel revision, chat template, index sync log, prompt changesreturn to the pinned revision and rerun the evaluation set
Replica slow to startmodel loaded over the network or downloaded at startkeep checkpoints on a local model store

First checks from vLLM’s troubleshooting and memory guides (September 2026) and NVIDIA’s XID catalogue; the next steps are our suggestions.

vLLM’s memory guide lists tensor parallelism, quantised models, a shorter max_model_len, a smaller max_num_seqs and fewer CUDA graphs as ways to reduce memory, and the enforce_eager flag disables graph capture completely. For an instance that stops making progress, its troubleshooting guide suggests VLLM_LOGGING_LEVEL=DEBUG, CUDA_LAUNCH_BLOCKING=1 and, for multi-GPU communication, NCCL_DEBUG=TRACE. Use them on a replica that carries no user traffic. Our guide to GPU server monitoring with DCGM explains which XID messages call for a reset or a drained node. The SRE book asks for postmortems “after significant incidents” with a full timeline, and each one should end with a changed runbook or alert.

SLOs and on-call for an internal LLM service

The SRE book defines an SLO as “a target value or range of values for a service level that is measured by an SLI”, and adds that insisting on meeting SLOs 100 per cent of the time is “both unrealistic and undesirable”. It prefers percentiles to averages and treats the allowed shortfall as an error budget. For an internal assistant, three SLOs cover most of what users notice. Our example targets are 99.5 per cent of requests answered without error during business hours each month; 95 per cent of first tokens within 2.5 s, measured at the gateway; and no release whose evaluation results fall below those of the version it replaces. When the error budget for the month is spent, changes wait until the cause is fixed.

On-call follows from these targets. Google’s SRE book sizes on-call cover at all hours, with a primary and a secondary on call, at a minimum of eight engineers on one site, or six per site for two sites, and caps on-call at 25 per cent of an engineer’s time. For a platform run by two or three people, decide which hours the assistant must be available, whether a failure out of hours waits for the next working day, and who covers which part, with any support agreement stating its hours and response times. Our guide to LLM high availability on two GPU nodes covers the redundancy that makes calls rarer.

After the build, your own staff run the platform or we support it under an agreed SLA. Tell us who runs your platform today and which support you want from a supplier.

Access reviews and query log retention

Access to models and RAG sources should follow directory groups, so that a review means checking group membership against the data owners’ decisions. Implementing Regulation (EU) 2024/2690, which binds the digital infrastructure and digital service providers listed in its Article 1, requires in point 11.2.3 of its Annex that they “review access rights at planned intervals” and document the results. For other companies it is a usable reference. Review service accounts and gateway keys in the same cycle.

The query log records who asked what and which sources were used, so it contains personal data. Under Article 5(1)(e) of the GDPR, personal data are kept in a form that identifies people “for no longer than is necessary” for their purpose, and point 3.2.5 of the same Annex asks entities in its scope to “maintain and back up logs for a predefined period” and protect them from unauthorised access or changes. Set the retention period and the people who may read the log before the first entry, and automate deletion. How long the text may be kept is a legal assessment for your legal department.

What we do

Our Private AI/ML service deploys open and commercial models on-premise with vLLM, Ollama or NVIDIA AI Enterprise, on a Kubernetes-based platform where it fits, with logging of queries and answers and data and permissions management. We build and support the platform and train your team to run it and develop it further; afterwards it is your choice whether your own staff run it or we support it under an agreed SLA. Eurokommerz holds the contract and supplies the GPU servers, with engineering by our partner Vixen.UNO. The first call is free of charge, and the price of the technical assessment is fixed before work begins. How we handle data during a project is described on our security and compliance page.

FAQ

What is LLMOps?
LLMOps is the operation of large language models in production: keeping checkpoints, serving engines, drivers and retrieval indexes current, testing each change against an evaluation set, and handling capacity, incidents, access and logs. On a private platform it is mostly day-2 work on models the company did not train.
What is the difference between MLOps and LLMOps?
MLOps centres on training a company’s own models from its data, with pipelines, a registry and retraining. LLMOps usually serves checkpoints from a publisher, so the controls are the choice of checkpoint and precision, prompts and chat templates, retrieval and engine settings, each tested on a fixed set of questions before release.
How do you update an LLM model in production?
Pin the new revision, run the evaluation set against it, and deploy it beside the current version under its own name. Send a small weighted share of traffic or a pilot group to it, move the rest in steps, and keep the old version running so that a rollback is one routing change.
How often does vLLM release new versions?
vLLM aims at a regular release every 2 weeks, and since v0.12.0 each one increments the minor version; v0.31.0 came out on 5 October 2026. Backwards compatibility is guaranteed only for a limited number of minor releases, and deprecated features are removed after several releases, so read the notes of every version you skip.
Which roles does an LLM platform team need?
A platform owner for servers, drivers, Kubernetes and the serving engine, a model owner for checkpoints and the evaluation set, a data owner for each document source, and security for access reviews and the query log. One person can hold two roles, but changes that alter answers need a sign-off by someone other than the person who made them.
Does a private LLM platform need on-call?
It needs the cover its users depend on, which the SLOs and service hours define. Google’s SRE book puts the minimum single-site on-call rotation at eight engineers, for cover at all hours with a primary and a secondary, so a mid-size company usually decides which hours need cover, adds redundant replicas and agrees who handles failures outside those hours.

Send us the models, serving engines and GPU servers you run, the number of users and who operates the platform today. We reply within one business day, and in the first call we work through your process and data with you, so you leave with 2 to 3 possible solution scenarios. The first call is free of charge.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna