Managed private AI: who runs the LLM platform after go-live, and what the support SLA covers
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- After go-live a private LLM platform needs owners for model and serving-engine updates, GPU driver and firmware updates, capacity reviews, access and query-log management, incidents and user support
- A company covers this with its own staff, with a supplier’s support under an agreed SLA, or with a mix in which it keeps the decisions on models, data and retention and the supplier carries out agreed technical tasks
- The update load is steady: vLLM aims at a regular release every 2 weeks, and NVIDIA releases two production driver branches a year, each with bug fixes and security updates for up to 1 year
- A managed platform can run on-premise on the company’s servers or on dedicated hardware in an EU data centre; in both cases the support provider should work under the company’s access controls
- The support SLA defines scope, service hours, priorities and response times, maintenance windows and change approval, access logging, reporting and the exit, with figures the parties agree for their own case
Eurokommerz × Vixen.UNO: Private AI/ML Talk to an expert →
Managed private AI: who runs the platform after go-live
After go-live, a private LLM platform needs an owner for model and serving-engine updates, GPU driver and firmware updates, capacity reviews, access and query-log management, incidents and user questions. A company can cover this with its own staff, with a supplier’s support under an agreed SLA, or with a mix in which it keeps the decisions and the supplier carries out agreed technical tasks. Managed private AI, or a private LLM managed service, usually means the second or third option, with the models still running on hardware dedicated to the company.
Take a platform for 1,200 employees on two GPU servers. It runs a chat model, an embedding model and a reranker on vLLM, a gateway with a query log, a chat front end and a vector index fed from six document sources. vLLM’s release document says “We aim to have a regular release every 2 weeks”. NVIDIA’s driver lifecycle page, updated on 9 September 2026, says that two production branches are released per year, each with bug fixes and security updates “for up to 1 year”. Each release has to be read, tested and applied in a maintenance window.
What has to be run after go-live
Our guide to operating a private LLM platform describes the day-2 work in depth: release cadences, canary rollouts, incident runbooks and access reviews. For the question of who does it, sort the work into three kinds.
Technical platform work covers the GPU servers, firmware and drivers, Kubernetes, the serving engine, the gateway, the front end and the index. It needs skills a mid-size IT team rarely has twice over, and it is the part a support provider can take on. Decisions about content and risk cover which model is served, which documents enter the index for which directory groups, how long the query log is kept and which tasks may use a public API. They stay with the company, because they rest on its data and its policies. User-facing work covers questions, access requests, training and feedback, and it fits the service desk the company already runs.
According to PeopleCert, ITIL 4 launched in 2019 with guidance for 34 management practices. Where IT already works with them, the AI platform joins the existing practices instead of getting its own. Faults go to incident management and problem management, access requests to service request management, and model and engine updates pass through change enablement and deployment management. Our AI rollout plan for 1,000 employees sets up the service desk category with three routes: access, platform and content.
Three operating models: own staff, support under an SLA, or a mix
With its own staff, the company holds every role. Our day-2 guide names four: a platform owner, a model owner, a data owner for each document source and security. One person can hold two roles, but each technical skill needs a second person for holidays and sick leave, so plan at least two people with GPU, Kubernetes and LLM serving skills for 1,200 users, beside their other work.
With support under an SLA, the supplier carries out the technical platform work within the hours and response times the agreement sets. The company still decides what goes in and who may read it. NIST’s AI Risk Management Framework (NIST AI 100-1, January 2023) asks in GOVERN 2.1 that roles, responsibilities and lines of communication for mapping, measuring and managing AI risks are documented and clear to individuals and teams throughout the organisation, and a support agreement does not replace that documentation.
The mix suits a company that has an IT team but no second person for each platform skill. The company keeps the model owner, the data owners, security and the first line at the service desk; the supplier is the second line for platform faults and carries out upgrades in agreed windows. The split can change once the company has trained its own platform owner.
After the build, our Private AI/ML service leaves the choice with you: your own staff, or our support under an agreed SLA. Tell us which tasks you want to keep in-house and who runs your systems today.
RACI for a private LLM platform: company and support provider
A RACI matrix gives each task one accountable party (A), the parties that carry it out (R), those consulted before (C) and those informed after (I). For the mix model the split can look like this.
| TASK | COMPANY | SUPPORT PROVIDER |
|---|---|---|
| Driver and firmware updates | A: approves the window and the change | R: tests on one node, applies node by node, reports |
| Serving engine upgrade | A: model owner signs off the evaluation results | R: upgrades, runs the evaluation set, keeps the rollback ready |
| Model change or refresh | A and R: model owner chooses the checkpoint | C: memory, throughput and serving settings |
| Document sources and access | A and R: data owners and security decide groups | C: maps groups to sources in the index |
| Query log retention | A: security and legal set the period | R: configures automated deletion and access to the log |
| Capacity review | A: decides on more GPUs or servers | R: prepares busy-hour figures and options |
| Platform incident | A: service owner, I: users via service desk | R: diagnosis and restore within the agreed hours |
| User questions and access | R: service desk, first line | C: second line for platform faults |
| Public API use (hybrid) | A: decides per task and data class | R: enables the route and logs what goes out |
Example split for a platform of 500 to 2,000 users, not a description of a specific contract; roles follow our day-2 guide and NIST AI RMF GOVERN 2.1.
The capacity review rests on the serving metrics described in our guide to LLM monitoring with vLLM metrics, and the decision on the next server is the subject of our GPU capacity planning guide.
Managed AI server on-premise or in an EU data centre
The location changes who provides the room, power and physical access, but not who decides about data. On-premise, the servers stand in the company’s server room, and the support provider reaches them over a connection the company controls. In an EU data centre on dedicated hardware, the data-centre operator provides power, cooling, physical security and connectivity, and the platform is reached over a VPN or a private line. Our article on private LLM hosting in an EU data centre compares the hosting options in detail.
| DEPLOYMENT OPTION | COMPANY KEEPS | MANAGED BY OTHERS |
|---|---|---|
| On-premise, own staff | room, power, network, servers and the whole software stack | manufacturer warranty and vendor support only |
| On-premise, SLA support | room, power, network, decisions on models, data and retention | agreed platform tasks within the service hours |
| EU data centre, dedicated | decisions on models, data, access and retention | power, cooling, physical security by the operator; agreed platform tasks |
| Hybrid with public APIs | the decision for each task and data class | the API provider runs its model; the query log shows what went out |
Deployment options as on our Private AI/ML page; the division of tasks is our example.
With our Private AI/ML service the models run on your hardware or on dedicated hardware in a Tier-3 data centre in Lithuania, and you choose at the assessment stage. Describe where your platform should run in the form below.
What a support SLA for an internal AI platform should define
Google’s SRE book defines an SLA as “an explicit or implicit contract with your users that includes consequences of meeting (or missing) the SLOs they contain”. An internal AI platform has two kinds of agreement. The company’s SLO for the assistant, for example answers without error during business hours, depends on components the company runs as well, such as the directory service, the network and the document sources. The supplier’s SLA covers the tasks the supplier owns. Write both down and keep them apart, so that a missed target can be traced to its cause. An SLA for platform support should define these points:
- The scope, meaning the servers, components and models that are covered and those that are not.
- The service hours, and whether a fault outside them is handled or waits for the next working day.
- Priority levels with the response time for each, in figures both parties agree for their case.
- Maintenance windows, notice periods and the change approval, with a rollback plan for each change.
- The provider’s accounts, the systems they reach, multi-factor sign-in and the logging of every session.
- Reports on incidents, changes and capacity, and a review meeting at a set interval.
- Incident notification from the provider to the company, and the vulnerability handling for the stack.
- The exit, with the handover of configuration, runbooks and documentation and the disposal of data the provider held.
Implementing Regulation (EU) 2024/2690 binds the providers listed in its Article 1, among them cloud computing, data centre and managed service providers; for other companies its Annex is a usable reference. Point 5.1.4 asks that contracts with suppliers and service providers specify, “where appropriate through service level agreements”, among other things incident notification without undue delay, the right to audit, the handling of vulnerabilities and obligations at the end of the contract. Point 5.1.7 adds that the entities “regularly monitor reports on the implementation of the service level agreements, where applicable”.
Vendor support behind a managed LLM platform
A support provider depends on the vendors behind the stack. For NVIDIA AI Enterprise, NVIDIA’s support page, as read in October 2026, describes a “Support, upgrades, and maintenance subscription, included with every NVIDIA AI Enterprise software license”, with experts available “during local business hours”. The same page lists NVIDIA-Certified Systems among the supported infrastructure options, so check the certification of the GPU servers. NVIDIA also sells Business Critical support, which it calls its premium level, for select products, with a one-hour response time for Severity Level 1 cases. Our guide to NVIDIA AI Enterprise licensing explains what the subscription covers.
vLLM is an open-source project without a support SLA of its own; commercial support for a vLLM-based runtime is sold by vendors such as Red Hat, whose AI Inference page says “Powered by vLLM and llm-d”. vLLM publishes security advisories on its GitHub security page, the most recent of them on 27 July 2026. For critical and high-severity issues its policy says fixes are developed in a private security fork before public disclosure. The operator of the platform has to follow these notices, judge whether a fix affects the deployed version and schedule the upgrade; the SLA should say who does that.
GPU servers are covered by the manufacturer warranty. The SLA should name who opens the warranty case, who runs the diagnostics the manufacturer asks for and who swaps the part on site.
Access, data protection and exit in a managed contract
A support provider with administrator rights on the GPU servers can read prompts, answers, the index and the query log, which contain personal data. Article 28(3) GDPR requires that processing by a processor “shall be governed by a contract or other legal act” that sets out the subject matter, duration, nature and purpose of the processing, the type of personal data, the categories of data subjects and the obligations and rights of the controller. Whether the provider acts as a processor for your data is a legal assessment for your legal department. Give the provider named accounts in the directory service, restrict them to the platform, and log every session.
Plan the exit when you sign. Configuration and deployment files belong in a repository the company owns, runbooks in the company’s documentation, and the provider’s credentials are revoked on the last day. NIST’s GOVERN 6.2 asks for “Contingency processes” to handle failures or incidents in third-party data or AI systems deemed to be high-risk; where a supplier runs the platform, a written handover plan for the end of the contract belongs among them.
What we do
Our Private AI/ML service builds and supports the platform and trains your team to run it and develop it further; afterwards the choice is yours, your own staff or our support under an agreed SLA. The models run on-premise on your servers or on dedicated hardware in a Tier-3 data centre in Lithuania, with logging of queries and answers and data and permissions management, and public APIs only where you enable them. Our engineering partner Vixen.UNO works in your environment under your access controls and takes no copies of production data out of it, with an NDA before technical detail and a data processing agreement on request, as our security and compliance page describes. Eurokommerz holds the contract and supplies the GPU servers, and a dedicated Eurokommerz project lead stays your single point of contact. The first call is free of charge, and the price of the technical assessment is fixed before work begins.
FAQ
What is managed private AI?
Who runs a private AI platform after the project ends?
What does a managed LLM platform include?
What should a private AI support SLA define?
Can a managed GPU server for AI run in an EU data centre?
Does NVIDIA AI Enterprise include support?
Send us the models and GPU servers you run or plan, the number of users, where the platform should run and which tasks your own staff will keep. We reply within one business day, and in the first call we work through your process and data with you, so you leave with 2 to 3 possible solution scenarios. The first call is free of charge.
Talk to an expertWe reply within one business day