Where to run private LLM inference in Europe: on-premise, an EU data centre or cloud GPUs
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- A private LLM for a RAG application can run on GPU servers in your own server room, on servers dedicated to you in an EU data centre, on GPU instances in the EU regions of global clouds or with a GPU-as-a-service provider; without a US provider in the data path, it runs on-premise or on dedicated servers where neither the data-centre operator nor any hosting provider is a US company or owned by one
- Prompts, retrieved passages, answers, the vector index and the query logs all carry document content; OWASP’s Top 10 for LLM Applications 2026 states that stored embeddings can be inverted to recover source text, and a cloud embedding API receives the text of every document during indexing
- A hyperscaler’s EU region fixes where stored data sits, not which jurisdiction the provider is subject to, and a model API called through a global endpoint or inference profile can process prompts in regions outside the EU
- On-premise needs a rack position that can feed and cool GPU servers; in an EU data centre the operator provides power, cooling and physical security, and the questions move to who holds root and BMC access and where data and logs are stored
- Where a host processes personal data on your behalf, Article 28 GDPR requires a processor contract under which subprocessors need your prior written authorisation and, at your choice, the data is deleted or returned at the end of the service, with existing copies deleted
Eurokommerz × Vixen.UNO: Private AI/ML Talk to an expert →
Where to run private LLM inference in Europe
A private LLM for a RAG application can run in four places in Europe: on GPU servers in your own server room, on servers dedicated to you in an EU data centre, on GPU instances that global cloud providers rent out in their EU regions, or with a GPU-as-a-service provider. If prompts and documents must not reach a US provider, run the model on-premise, where the whole data path stays in your network, or on dedicated servers in a data centre where neither the operator nor a hosting provider running the servers is a US company or owned by one. A GPU cloud run by such a company can serve as well, once the contract names who operates the hardware and where. An EU region of a global cloud fixes where data is stored, while the operator’s jurisdiction decides whether a US order can reach it, as our article on the CLOUD Act and EU data residency explains.
| OPTION | DATA PATH | WHO OPERATES | FITS |
|---|---|---|---|
| On-premise GPU servers | stays in your network unless you add a connection | your team, or a partner under your access controls | data that must stay in your network; a rack that can feed and cool the servers |
| EU data centre, dedicated | over a VPN or private line to hardware no one else uses | the operator runs power, cooling and physical access; you, a partner or the host run the servers | no suitable server room, users on several sites, data that may leave the building but not the EU |
| Cloud GPUs in an EU region | to the region you choose, on the provider’s multi-tenant infrastructure | the provider runs hardware and virtualisation; you run the OS, drivers and model stack | tests, peaks and short projects, with data your policy allows at that provider |
| GPU-as-a-service | to wherever the rented GPUs stand | the provider, at times on a partner’s hardware | training and fine-tuning runs, evaluations, bursts of demand |
Our summary of the sections below and of the provider documents they cite, read on 6 October 2026.
What a RAG application sends, and where it can leave
A user’s question travels from the browser to the chat front end, which searches the index for passages the user may read, adds them to the prompt and sends both to the model server; the answer streams back the same way. Each component on that path handles company data.
| COMPONENT | WHAT IT HANDLES | WHERE IT CAN LEAVE |
|---|---|---|
| Chat front end | accounts, questions, saved chats | when another provider hosts it |
| Index and retrieval | chunk text, embeddings, permissions | with replicas and backups of the index |
| Embedding model | every chunk at indexing, every question | when the embedding model is a cloud API |
| Model server | prompts with passages, answers, the KV cache | when it runs on another operator’s hardware |
| Query log | prompts, passages, answers, user names | wherever logs are shipped and kept |
| Public API connection | whatever its users send to it | at every request, by design |
| Telemetry, update checks | usage statistics, version checks | until switched off |
Data flows of a self-hosted stack as in our private ChatGPT alternative guide; OWASP Top 10 for LLM Applications 2026, LLM02 and LLM09 (resource page dated 3 August 2026).
An index for RAG keeps an embedding for each chunk of each document, and the chunk’s text in the index or in a store next to it, because that text goes into the prompt. The index is therefore a second copy of the document store and belongs where the documents may be. Embeddings without the text are no exception, since OWASP’s Top 10 for LLM Applications 2026 states under LLM09, vector and embedding weaknesses, that “Stored embeddings can be inverted to recover source text”. A cloud embedding API receives the text of every document during indexing, even when the LLM itself runs privately. Our guide to a private ChatGPT alternative lists the settings that keep a self-hosted stack from calling out. Downloading open model weights sends no prompts anywhere, and a cloud identity provider sees sign-ins, not conversations.
On-premise GPU servers: rack, power and operations
On-premise, the GPU server stands next to the systems that hold the documents, and prompts, index and logs stay in your network unless you add a connection. A server with several GPUs draws kilowatts under load, and the circuit rating, the PDU outlet type, the floor loading and the GPU power cable decide whether a delivered server can be switched on; our article on GPU rack power and cooling works through these checks and the airflow a rack position needs. What to specify for the server itself is in our guide to buying an AI server.
Operating it is the job of your team or a partner. Drivers, the serving engine and the front end need updates, failed parts need replacing, and one server is a single point of failure, so a production assistant needs a second server or a fallback the business accepts.
We build GPU servers for private LLMs and RAG to order, and we check the rack, power and airflow before we quote. Send us the model, your peak number of users and the power feed of the rack position through the form below.
Dedicated GPU servers in an EU data centre
Where the server room cannot take GPU servers, or users work from several sites, the servers can stand in an EU data centre and still serve only you. In colocation you own the servers and rent space, power and cooling; in dedicated hosting a provider supplies and hosts servers that no other customer uses. Either way the GPUs, memory and local disks hold only your data for as long as the contract runs.
The data centre takes over power, cooling, physical security and connectivity. Ask how much power per rack the contract includes, because a rack of GPU servers can draw more than one planned for classic servers was designed to carry. A certificate such as ISO 27001 counts for the site only if its scope statement names it. Settle as well who has root on the servers, who holds the credentials for the BMC, which controls the server below the operating system, and whether the data centre’s staff touch the hardware only at your request. The index, with the text of every chunk, and the logs then live in the data centre too, reached over a site-to-site VPN or a private line. The jurisdiction question covers the data-centre operator, any hosting provider between you and the hardware, and their owners.
Our Private AI/ML service can run the models on dedicated hardware in a Tier-3 data centre in Lithuania, where data stays in the EU. Tell us where the documents are stored today and which data classes the assistant may see.
Cloud GPUs and model APIs in EU regions
Global cloud providers rent GPU instances in their EU regions on demand, and the region you choose fixes where stored data sits. AWS, for example, says in its data privacy FAQ, as read on 6 October 2026, that it will not move or replicate customer content outside the chosen Regions “except as necessary to provide the services you initiated”, or to comply with the law or a binding order of a governmental body. Under its shared responsibility model, the customer manages the guest operating system and its patches, the software installed on an instance and the firewall rules, so the drivers, model server and front end are yours to run, as on-premise. Check that the GPU type you need is offered in the region and whether capacity has to be reserved.
The same providers sell models through APIs, where the way you call a model decides where a prompt is processed. Amazon Bedrock’s documentation, as read on 6 October 2026, describes inference profiles tied to a geography such as the EU, which send each request to a Region within that geography, and global profiles that can route requests to any supported commercial Region worldwide. Each cross-Region request is logged in the source Region, where the field additionalEventData. shows where it was processed. When an API costs less than a server of your own, and when it does not, is the subject of our comparison of a private LLM and a cloud API.
GPU-as-a-service: who runs the hardware and who shares the GPU
GPU-as-a-service providers, European GPU clouds among them, rent GPUs on demand or for fixed terms, as virtual machines, bare-metal servers or containers. The company you sign with does not always run the servers. NVIDIA, for example, describes its DGX Cloud Lepton platform as bringing together “a global network of NVIDIA Cloud Partners (NCPs), GPU marketplaces, cloud providers, and local environments”, as read on 6 October 2026. Have the operating company and the data centre written into the contract.
A whole GPU on bare metal or passed through to a virtual machine serves one tenant at a time, while a shared GPU is split between tenants, for example into MIG partitions or vGPU profiles. Where GPUs are shared between containers by time-slicing, NVIDIA’s GPU Operator documentation (updated 23 September 2026) notes that this gives up the memory and fault isolation that MIG provides. Containers on one host also share its kernel, so a kernel or runtime flaw can expose one tenant to another, and the host’s administrators can reach what runs in them, as they can in virtual machines without confidential computing. Model weights, indexes and logs can sit on local NVMe disks, so ask how those are wiped when an instance ends.
For training or fine-tuning, the EuroHPC Joint Undertaking’s website lists, as of October 2026, 19 AI Factories that the EU has established around its supercomputers, offering computing resources to European industry through access calls. Their access modes for industrial innovation are open to AI SMEs and start-ups, while other industrial applications can use pay-per-use commercial access. The Joint Undertaking describes AI Factories as hubs for developing generative AI models, so an assistant used every day needs a host of its own.
Latency, access and contract: what to settle before choosing
Network delay between sites in the EU is small next to the seconds a model needs to stream an answer of a few hundred tokens, so distance inside the EU rarely decides the site. Round trips add up where an agent calls the model many times in sequence, or where retrieval runs at one site and the model at another. Bandwidth matters for the first indexing of a large document store and for keeping the index current, and the link becomes a dependency, since users at that site lose the assistant when it fails.
Where a host processes personal data on your behalf, Article 28 GDPR requires a contract under which it engages another processor only with your prior specific or general written authorisation and, at your choice, deletes or returns the personal data at the end of the service and deletes existing copies, unless Union or Member State law requires their storage. How these rules apply to your data is a legal assessment for your legal department. Settle these points with each candidate host before you choose:
- The legal entity that operates the hardware, the data centre it stands in, and who owns both.
- Whether GPUs, memory and local disks are dedicated to you and, if they are shared, how tenants are separated and disks wiped.
- Who holds root, hypervisor and BMC access, from which countries, and how their sessions are logged.
- Where prompts, answers, the index, model weights, logs and backups are stored, and for how long.
- Which connections leave the platform: public APIs, embedding services, telemetry and update checks.
- The subprocessors named in the contract, and how the data comes back to you at the end.
- The power per rack for your own servers, the latency from your sites and the bandwidth for indexing.
What we do
Our Private AI/ML service deploys private LLMs and RAG assistants in one of three ways: on-premise on your servers, where data never leaves your network; on dedicated hardware in Baltneta’s Tier-3 data centres in Lithuania, with ISO 27001 and PCI DSS Level 1, where data stays in the EU; or hybrid, where public APIs are enabled only by your explicit decision and what goes to them is visible in the query log. You choose the option at the assessment stage, and a TCO calculation against cloud GPUs comes before any purchase. We build AI servers for inference and RAG to order, with 2 to 8 GPUs per node sized by model size and concurrent users. Eurokommerz holds the contract and supplies the hardware, engineering is by our partner Vixen.UNO, and our security and compliance page lists the documents we sign. The first call is free of charge, and the price of the technical assessment is fixed before work begins.
FAQ
Where can I run GPU inference for a RAG app in Europe without sending data to a US provider?
What is the difference between a dedicated GPU server and a cloud GPU instance?
Does an EU cloud region keep LLM prompts in the EU?
Can embeddings in a vector database reveal the original documents?
What does an on-premise GPU server for a private LLM need from the server room?
What should a contract with a GPU hosting provider in the EU cover?
Send us the model you plan to run, where the documents it should answer from are stored, your peak number of concurrent users and whether you would host on-premise or in an EU data centre. We reply within one business day to arrange a first call, from which you leave with two or three possible solution scenarios. The first call is free of charge.
Talk to an expertWe reply within one business day