BLOG · GUIDE ·

Managed private AI: who runs the LLM platform after go-live, and what the support SLA covers

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • After go-live a private LLM platform needs owners for model and serving-engine updates, GPU driver and firmware updates, capacity reviews, access and query-log management, incidents and user support
  • A company covers this with its own staff, with a supplier’s support under an agreed SLA, or with a mix in which it keeps the decisions on models, data and retention and the supplier carries out agreed technical tasks
  • The update load is steady: vLLM aims at a regular release every 2 weeks, and NVIDIA releases two production driver branches a year, each with bug fixes and security updates for up to 1 year
  • A managed platform can run on-premise on the company’s servers or on dedicated hardware in an EU data centre; in both cases the support provider should work under the company’s access controls
  • The support SLA defines scope, service hours, priorities and response times, maintenance windows and change approval, access logging, reporting and the exit, with figures the parties agree for their own case

Eurokommerz × Vixen.UNO: Private AI/ML  Talk to an expert →

Managed private AI: who runs the platform after go-live

After go-live, a private LLM platform needs an owner for model and serving-engine updates, GPU driver and firmware updates, capacity reviews, access and query-log management, incidents and user questions. A company can cover this with its own staff, with a supplier’s support under an agreed SLA, or with a mix in which it keeps the decisions and the supplier carries out agreed technical tasks. Managed private AI, or a private LLM managed service, usually means the second or third option, with the models still running on hardware dedicated to the company.

Take a platform for 1,200 employees on two GPU servers. It runs a chat model, an embedding model and a reranker on vLLM, a gateway with a query log, a chat front end and a vector index fed from six document sources. vLLM’s release document says “We aim to have a regular release every 2 weeks”. NVIDIA’s driver lifecycle page, updated on 9 September 2026, says that two production branches are released per year, each with bug fixes and security updates “for up to 1 year”. Each release has to be read, tested and applied in a maintenance window.

What has to be run after go-live

Our guide to operating a private LLM platform describes the day-2 work in depth: release cadences, canary rollouts, incident runbooks and access reviews. For the question of who does it, sort the work into three kinds.

Technical platform work covers the GPU servers, firmware and drivers, Kubernetes, the serving engine, the gateway, the front end and the index. It needs skills a mid-size IT team rarely has twice over, and it is the part a support provider can take on. Decisions about content and risk cover which model is served, which documents enter the index for which directory groups, how long the query log is kept and which tasks may use a public API. They stay with the company, because they rest on its data and its policies. User-facing work covers questions, access requests, training and feedback, and it fits the service desk the company already runs.

According to PeopleCert, ITIL 4 launched in 2019 with guidance for 34 management practices. Where IT already works with them, the AI platform joins the existing practices instead of getting its own. Faults go to incident management and problem management, access requests to service request management, and model and engine updates pass through change enablement and deployment management. Our AI rollout plan for 1,000 employees sets up the service desk category with three routes: access, platform and content.

Three operating models: own staff, support under an SLA, or a mix

With its own staff, the company holds every role. Our day-2 guide names four: a platform owner, a model owner, a data owner for each document source and security. One person can hold two roles, but each technical skill needs a second person for holidays and sick leave, so plan at least two people with GPU, Kubernetes and LLM serving skills for 1,200 users, beside their other work.

With support under an SLA, the supplier carries out the technical platform work within the hours and response times the agreement sets. The company still decides what goes in and who may read it. NIST’s AI Risk Management Framework (NIST AI 100-1, January 2023) asks in GOVERN 2.1 that roles, responsibilities and lines of communication for mapping, measuring and managing AI risks are documented and clear to individuals and teams throughout the organisation, and a support agreement does not replace that documentation.

The mix suits a company that has an IT team but no second person for each platform skill. The company keeps the model owner, the data owners, security and the first line at the service desk; the supplier is the second line for platform faults and carries out upgrades in agreed windows. The split can change once the company has trained its own platform owner.

After the build, our Private AI/ML service leaves the choice with you: your own staff, or our support under an agreed SLA. Tell us which tasks you want to keep in-house and who runs your systems today.

RACI for a private LLM platform: company and support provider

A RACI matrix gives each task one accountable party (A), the parties that carry it out (R), those consulted before (C) and those informed after (I). For the mix model the split can look like this.

TASKCOMPANYSUPPORT PROVIDER
Driver and firmware updatesA: approves the window and the changeR: tests on one node, applies node by node, reports
Serving engine upgradeA: model owner signs off the evaluation resultsR: upgrades, runs the evaluation set, keeps the rollback ready
Model change or refreshA and R: model owner chooses the checkpointC: memory, throughput and serving settings
Document sources and accessA and R: data owners and security decide groupsC: maps groups to sources in the index
Query log retentionA: security and legal set the periodR: configures automated deletion and access to the log
Capacity reviewA: decides on more GPUs or serversR: prepares busy-hour figures and options
Platform incidentA: service owner, I: users via service deskR: diagnosis and restore within the agreed hours
User questions and accessR: service desk, first lineC: second line for platform faults
Public API use (hybrid)A: decides per task and data classR: enables the route and logs what goes out

Example split for a platform of 500 to 2,000 users, not a description of a specific contract; roles follow our day-2 guide and NIST AI RMF GOVERN 2.1.

The capacity review rests on the serving metrics described in our guide to LLM monitoring with vLLM metrics, and the decision on the next server is the subject of our GPU capacity planning guide.

Managed AI server on-premise or in an EU data centre

The location changes who provides the room, power and physical access, but not who decides about data. On-premise, the servers stand in the company’s server room, and the support provider reaches them over a connection the company controls. In an EU data centre on dedicated hardware, the data-centre operator provides power, cooling, physical security and connectivity, and the platform is reached over a VPN or a private line. Our article on private LLM hosting in an EU data centre compares the hosting options in detail.

DEPLOYMENT OPTIONCOMPANY KEEPSMANAGED BY OTHERS
On-premise, own staffroom, power, network, servers and the whole software stackmanufacturer warranty and vendor support only
On-premise, SLA supportroom, power, network, decisions on models, data and retentionagreed platform tasks within the service hours
EU data centre, dedicateddecisions on models, data, access and retentionpower, cooling, physical security by the operator; agreed platform tasks
Hybrid with public APIsthe decision for each task and data classthe API provider runs its model; the query log shows what went out

Deployment options as on our Private AI/ML page; the division of tasks is our example.

With our Private AI/ML service the models run on your hardware or on dedicated hardware in a Tier-3 data centre in Lithuania, and you choose at the assessment stage. Describe where your platform should run in the form below.

What a support SLA for an internal AI platform should define

Google’s SRE book defines an SLA as “an explicit or implicit contract with your users that includes consequences of meeting (or missing) the SLOs they contain”. An internal AI platform has two kinds of agreement. The company’s SLO for the assistant, for example answers without error during business hours, depends on components the company runs as well, such as the directory service, the network and the document sources. The supplier’s SLA covers the tasks the supplier owns. Write both down and keep them apart, so that a missed target can be traced to its cause. An SLA for platform support should define these points:

  1. The scope, meaning the servers, components and models that are covered and those that are not.
  2. The service hours, and whether a fault outside them is handled or waits for the next working day.
  3. Priority levels with the response time for each, in figures both parties agree for their case.
  4. Maintenance windows, notice periods and the change approval, with a rollback plan for each change.
  5. The provider’s accounts, the systems they reach, multi-factor sign-in and the logging of every session.
  6. Reports on incidents, changes and capacity, and a review meeting at a set interval.
  7. Incident notification from the provider to the company, and the vulnerability handling for the stack.
  8. The exit, with the handover of configuration, runbooks and documentation and the disposal of data the provider held.

Implementing Regulation (EU) 2024/2690 binds the providers listed in its Article 1, among them cloud computing, data centre and managed service providers; for other companies its Annex is a usable reference. Point 5.1.4 asks that contracts with suppliers and service providers specify, “where appropriate through service level agreements”, among other things incident notification without undue delay, the right to audit, the handling of vulnerabilities and obligations at the end of the contract. Point 5.1.7 adds that the entities “regularly monitor reports on the implementation of the service level agreements, where applicable”.

Vendor support behind a managed LLM platform

A support provider depends on the vendors behind the stack. For NVIDIA AI Enterprise, NVIDIA’s support page, as read in October 2026, describes a “Support, upgrades, and maintenance subscription, included with every NVIDIA AI Enterprise software license”, with experts available “during local business hours”. The same page lists NVIDIA-Certified Systems among the supported infrastructure options, so check the certification of the GPU servers. NVIDIA also sells Business Critical support, which it calls its premium level, for select products, with a one-hour response time for Severity Level 1 cases. Our guide to NVIDIA AI Enterprise licensing explains what the subscription covers.

vLLM is an open-source project without a support SLA of its own; commercial support for a vLLM-based runtime is sold by vendors such as Red Hat, whose AI Inference page says “Powered by vLLM and llm-d”. vLLM publishes security advisories on its GitHub security page, the most recent of them on 27 July 2026. For critical and high-severity issues its policy says fixes are developed in a private security fork before public disclosure. The operator of the platform has to follow these notices, judge whether a fix affects the deployed version and schedule the upgrade; the SLA should say who does that.

GPU servers are covered by the manufacturer warranty. The SLA should name who opens the warranty case, who runs the diagnostics the manufacturer asks for and who swaps the part on site.

Access, data protection and exit in a managed contract

A support provider with administrator rights on the GPU servers can read prompts, answers, the index and the query log, which contain personal data. Article 28(3) GDPR requires that processing by a processor “shall be governed by a contract or other legal act” that sets out the subject matter, duration, nature and purpose of the processing, the type of personal data, the categories of data subjects and the obligations and rights of the controller. Whether the provider acts as a processor for your data is a legal assessment for your legal department. Give the provider named accounts in the directory service, restrict them to the platform, and log every session.

Plan the exit when you sign. Configuration and deployment files belong in a repository the company owns, runbooks in the company’s documentation, and the provider’s credentials are revoked on the last day. NIST’s GOVERN 6.2 asks for “Contingency processes” to handle failures or incidents in third-party data or AI systems deemed to be high-risk; where a supplier runs the platform, a written handover plan for the end of the contract belongs among them.

What we do

Our Private AI/ML service builds and supports the platform and trains your team to run it and develop it further; afterwards the choice is yours, your own staff or our support under an agreed SLA. The models run on-premise on your servers or on dedicated hardware in a Tier-3 data centre in Lithuania, with logging of queries and answers and data and permissions management, and public APIs only where you enable them. Our engineering partner Vixen.UNO works in your environment under your access controls and takes no copies of production data out of it, with an NDA before technical detail and a data processing agreement on request, as our security and compliance page describes. Eurokommerz holds the contract and supplies the GPU servers, and a dedicated Eurokommerz project lead stays your single point of contact. The first call is free of charge, and the price of the technical assessment is fixed before work begins.

FAQ

What is managed private AI?
Managed private AI is a private LLM platform on hardware dedicated to the company, with a supplier carrying out agreed operating tasks under a support agreement. The company keeps the decisions on models, data, access and query-log retention, while the supplier handles technical work such as upgrades and platform faults.
Who runs a private AI platform after the project ends?
Either the company’s own staff, a supplier under an agreed SLA, or a mix of both. At 500 to 2,000 users a mix can look like this: the company keeps the model owner, the data owners, security and the first line at the service desk, and the supplier is the second line for platform faults and upgrades.
What does a managed LLM platform include?
It covers the GPU servers, drivers and firmware, the serving engine, the gateway with its query log, the front end and the vector index, with updates, incident handling and capacity figures. Which of these the supplier carries out is written in the scope of the support agreement.
What should a private AI support SLA define?
Scope, service hours, priority levels with response times, maintenance windows and change approval, provider access and its logging, reporting, incident notification and the exit. The figures depend on which hours the business needs the assistant and are agreed for each case.
Can a managed GPU server for AI run in an EU data centre?
Yes, on dedicated hardware, where the data-centre operator provides power, cooling, physical security and connectivity and the platform is reached over a VPN or a private line. Decisions about models, data and access stay with the company in the same way as on-premise.
Does NVIDIA AI Enterprise include support?
NVIDIA’s support page describes a support, upgrades and maintenance subscription included with every NVIDIA AI Enterprise licence, with experts available during local business hours. NVIDIA lists NVIDIA-Certified Systems among the supported infrastructure options, and it sells Business Critical support, with a one-hour response for Severity Level 1 cases, separately for select products.

Send us the models and GPU servers you run or plan, the number of users, where the platform should run and which tasks your own staff will keep. We reply within one business day, and in the first call we work through your process and data with you, so you leave with 2 to 3 possible solution scenarios. The first call is free of charge.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna