BLOG · GUIDE ·

Where to run private LLM inference in Europe: on-premise, an EU data centre or cloud GPUs

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • A private LLM for a RAG application can run on GPU servers in your own server room, on servers dedicated to you in an EU data centre, on GPU instances in the EU regions of global clouds or with a GPU-as-a-service provider; without a US provider in the data path, it runs on-premise or on dedicated servers where neither the data-centre operator nor any hosting provider is a US company or owned by one
  • Prompts, retrieved passages, answers, the vector index and the query logs all carry document content; OWASP’s Top 10 for LLM Applications 2026 states that stored embeddings can be inverted to recover source text, and a cloud embedding API receives the text of every document during indexing
  • A hyperscaler’s EU region fixes where stored data sits, not which jurisdiction the provider is subject to, and a model API called through a global endpoint or inference profile can process prompts in regions outside the EU
  • On-premise needs a rack position that can feed and cool GPU servers; in an EU data centre the operator provides power, cooling and physical security, and the questions move to who holds root and BMC access and where data and logs are stored
  • Where a host processes personal data on your behalf, Article 28 GDPR requires a processor contract under which subprocessors need your prior written authorisation and, at your choice, the data is deleted or returned at the end of the service, with existing copies deleted

Eurokommerz × Vixen.UNO: Private AI/ML  Talk to an expert →

Where to run private LLM inference in Europe

A private LLM for a RAG application can run in four places in Europe: on GPU servers in your own server room, on servers dedicated to you in an EU data centre, on GPU instances that global cloud providers rent out in their EU regions, or with a GPU-as-a-service provider. If prompts and documents must not reach a US provider, run the model on-premise, where the whole data path stays in your network, or on dedicated servers in a data centre where neither the operator nor a hosting provider running the servers is a US company or owned by one. A GPU cloud run by such a company can serve as well, once the contract names who operates the hardware and where. An EU region of a global cloud fixes where data is stored, while the operator’s jurisdiction decides whether a US order can reach it, as our article on the CLOUD Act and EU data residency explains.

OPTIONDATA PATHWHO OPERATESFITS
On-premise GPU serversstays in your network unless you add a connectionyour team, or a partner under your access controlsdata that must stay in your network; a rack that can feed and cool the servers
EU data centre, dedicatedover a VPN or private line to hardware no one else usesthe operator runs power, cooling and physical access; you, a partner or the host run the serversno suitable server room, users on several sites, data that may leave the building but not the EU
Cloud GPUs in an EU regionto the region you choose, on the provider’s multi-tenant infrastructurethe provider runs hardware and virtualisation; you run the OS, drivers and model stacktests, peaks and short projects, with data your policy allows at that provider
GPU-as-a-serviceto wherever the rented GPUs standthe provider, at times on a partner’s hardwaretraining and fine-tuning runs, evaluations, bursts of demand

Our summary of the sections below and of the provider documents they cite, read on 6 October 2026.

What a RAG application sends, and where it can leave

A user’s question travels from the browser to the chat front end, which searches the index for passages the user may read, adds them to the prompt and sends both to the model server; the answer streams back the same way. Each component on that path handles company data.

COMPONENTWHAT IT HANDLESWHERE IT CAN LEAVE
Chat front endaccounts, questions, saved chatswhen another provider hosts it
Index and retrievalchunk text, embeddings, permissionswith replicas and backups of the index
Embedding modelevery chunk at indexing, every questionwhen the embedding model is a cloud API
Model serverprompts with passages, answers, the KV cachewhen it runs on another operator’s hardware
Query logprompts, passages, answers, user nameswherever logs are shipped and kept
Public API connectionwhatever its users send to itat every request, by design
Telemetry, update checksusage statistics, version checksuntil switched off

Data flows of a self-hosted stack as in our private ChatGPT alternative guide; OWASP Top 10 for LLM Applications 2026, LLM02 and LLM09 (resource page dated 3 August 2026).

An index for RAG keeps an embedding for each chunk of each document, and the chunk’s text in the index or in a store next to it, because that text goes into the prompt. The index is therefore a second copy of the document store and belongs where the documents may be. Embeddings without the text are no exception, since OWASP’s Top 10 for LLM Applications 2026 states under LLM09, vector and embedding weaknesses, that “Stored embeddings can be inverted to recover source text”. A cloud embedding API receives the text of every document during indexing, even when the LLM itself runs privately. Our guide to a private ChatGPT alternative lists the settings that keep a self-hosted stack from calling out. Downloading open model weights sends no prompts anywhere, and a cloud identity provider sees sign-ins, not conversations.

On-premise GPU servers: rack, power and operations

On-premise, the GPU server stands next to the systems that hold the documents, and prompts, index and logs stay in your network unless you add a connection. A server with several GPUs draws kilowatts under load, and the circuit rating, the PDU outlet type, the floor loading and the GPU power cable decide whether a delivered server can be switched on; our article on GPU rack power and cooling works through these checks and the airflow a rack position needs. What to specify for the server itself is in our guide to buying an AI server.

Operating it is the job of your team or a partner. Drivers, the serving engine and the front end need updates, failed parts need replacing, and one server is a single point of failure, so a production assistant needs a second server or a fallback the business accepts.

We build GPU servers for private LLMs and RAG to order, and we check the rack, power and airflow before we quote. Send us the model, your peak number of users and the power feed of the rack position through the form below.

Dedicated GPU servers in an EU data centre

Where the server room cannot take GPU servers, or users work from several sites, the servers can stand in an EU data centre and still serve only you. In colocation you own the servers and rent space, power and cooling; in dedicated hosting a provider supplies and hosts servers that no other customer uses. Either way the GPUs, memory and local disks hold only your data for as long as the contract runs.

The data centre takes over power, cooling, physical security and connectivity. Ask how much power per rack the contract includes, because a rack of GPU servers can draw more than one planned for classic servers was designed to carry. A certificate such as ISO 27001 counts for the site only if its scope statement names it. Settle as well who has root on the servers, who holds the credentials for the BMC, which controls the server below the operating system, and whether the data centre’s staff touch the hardware only at your request. The index, with the text of every chunk, and the logs then live in the data centre too, reached over a site-to-site VPN or a private line. The jurisdiction question covers the data-centre operator, any hosting provider between you and the hardware, and their owners.

Our Private AI/ML service can run the models on dedicated hardware in a Tier-3 data centre in Lithuania, where data stays in the EU. Tell us where the documents are stored today and which data classes the assistant may see.

Cloud GPUs and model APIs in EU regions

Global cloud providers rent GPU instances in their EU regions on demand, and the region you choose fixes where stored data sits. AWS, for example, says in its data privacy FAQ, as read on 6 October 2026, that it will not move or replicate customer content outside the chosen Regions “except as necessary to provide the services you initiated”, or to comply with the law or a binding order of a governmental body. Under its shared responsibility model, the customer manages the guest operating system and its patches, the software installed on an instance and the firewall rules, so the drivers, model server and front end are yours to run, as on-premise. Check that the GPU type you need is offered in the region and whether capacity has to be reserved.

The same providers sell models through APIs, where the way you call a model decides where a prompt is processed. Amazon Bedrock’s documentation, as read on 6 October 2026, describes inference profiles tied to a geography such as the EU, which send each request to a Region within that geography, and global profiles that can route requests to any supported commercial Region worldwide. Each cross-Region request is logged in the source Region, where the field additionalEventData.inferenceRegion shows where it was processed. When an API costs less than a server of your own, and when it does not, is the subject of our comparison of a private LLM and a cloud API.

GPU-as-a-service: who runs the hardware and who shares the GPU

GPU-as-a-service providers, European GPU clouds among them, rent GPUs on demand or for fixed terms, as virtual machines, bare-metal servers or containers. The company you sign with does not always run the servers. NVIDIA, for example, describes its DGX Cloud Lepton platform as bringing together “a global network of NVIDIA Cloud Partners (NCPs), GPU marketplaces, cloud providers, and local environments”, as read on 6 October 2026. Have the operating company and the data centre written into the contract.

A whole GPU on bare metal or passed through to a virtual machine serves one tenant at a time, while a shared GPU is split between tenants, for example into MIG partitions or vGPU profiles. Where GPUs are shared between containers by time-slicing, NVIDIA’s GPU Operator documentation (updated 23 September 2026) notes that this gives up the memory and fault isolation that MIG provides. Containers on one host also share its kernel, so a kernel or runtime flaw can expose one tenant to another, and the host’s administrators can reach what runs in them, as they can in virtual machines without confidential computing. Model weights, indexes and logs can sit on local NVMe disks, so ask how those are wiped when an instance ends.

For training or fine-tuning, the EuroHPC Joint Undertaking’s website lists, as of October 2026, 19 AI Factories that the EU has established around its supercomputers, offering computing resources to European industry through access calls. Their access modes for industrial innovation are open to AI SMEs and start-ups, while other industrial applications can use pay-per-use commercial access. The Joint Undertaking describes AI Factories as hubs for developing generative AI models, so an assistant used every day needs a host of its own.

Latency, access and contract: what to settle before choosing

Network delay between sites in the EU is small next to the seconds a model needs to stream an answer of a few hundred tokens, so distance inside the EU rarely decides the site. Round trips add up where an agent calls the model many times in sequence, or where retrieval runs at one site and the model at another. Bandwidth matters for the first indexing of a large document store and for keeping the index current, and the link becomes a dependency, since users at that site lose the assistant when it fails.

Where a host processes personal data on your behalf, Article 28 GDPR requires a contract under which it engages another processor only with your prior specific or general written authorisation and, at your choice, deletes or returns the personal data at the end of the service and deletes existing copies, unless Union or Member State law requires their storage. How these rules apply to your data is a legal assessment for your legal department. Settle these points with each candidate host before you choose:

  1. The legal entity that operates the hardware, the data centre it stands in, and who owns both.
  2. Whether GPUs, memory and local disks are dedicated to you and, if they are shared, how tenants are separated and disks wiped.
  3. Who holds root, hypervisor and BMC access, from which countries, and how their sessions are logged.
  4. Where prompts, answers, the index, model weights, logs and backups are stored, and for how long.
  5. Which connections leave the platform: public APIs, embedding services, telemetry and update checks.
  6. The subprocessors named in the contract, and how the data comes back to you at the end.
  7. The power per rack for your own servers, the latency from your sites and the bandwidth for indexing.

What we do

Our Private AI/ML service deploys private LLMs and RAG assistants in one of three ways: on-premise on your servers, where data never leaves your network; on dedicated hardware in Baltneta’s Tier-3 data centres in Lithuania, with ISO 27001 and PCI DSS Level 1, where data stays in the EU; or hybrid, where public APIs are enabled only by your explicit decision and what goes to them is visible in the query log. You choose the option at the assessment stage, and a TCO calculation against cloud GPUs comes before any purchase. We build AI servers for inference and RAG to order, with 2 to 8 GPUs per node sized by model size and concurrent users. Eurokommerz holds the contract and supplies the hardware, engineering is by our partner Vixen.UNO, and our security and compliance page lists the documents we sign. The first call is free of charge, and the price of the technical assessment is fixed before work begins.

FAQ

Where can I run GPU inference for a RAG app in Europe without sending data to a US provider?
On GPU servers in your own server room, or on servers dedicated to you in an EU data centre where neither the operator nor a hosting provider running the servers is a US company or owned by one; a GPU cloud run by such a company also works if the contract names who operates the hardware and where. Check the whole data path as well, because a cloud embedding API, a public API connection or a hosted front end can send documents to another provider while the model runs privately. A hyperscaler’s EU region fixes where stored data sits, but whether a US order can reach it depends on the provider’s jurisdiction, not on the region.
What is the difference between a dedicated GPU server and a cloud GPU instance?
A dedicated GPU server, your own in colocation or one rented from a hosting provider, serves only your workloads, and no other customer uses its GPUs, memory or local disks. A cloud GPU instance is usually a virtual machine rented on demand on infrastructure the provider runs for many customers, where the provider operates the hardware and virtualisation and you operate the operating system, drivers and model server. Instances suit tests and peaks, while dedicated servers suit steady production use and data that should stay on known hardware.
Does an EU cloud region keep LLM prompts in the EU?
For instances you run yourself, the region you choose fixes where stored data sits, within the exceptions the provider documents, such as compliance with a binding legal order. For model APIs it depends on how the model is called, because an endpoint or inference profile bound to the EU keeps processing in EU regions, while a global one can route a prompt to any region the provider supports worldwide. Where the data sits does not settle which jurisdiction the provider is subject to.
Can embeddings in a vector database reveal the original documents?
OWASP’s Top 10 for LLM Applications 2026 states under LLM09, vector and embedding weaknesses, that stored embeddings can be inverted to recover source text, and under LLM02, sensitive information disclosure, it calls a backup holding only embeddings a source-document breach. A RAG index, or a store next to it, also holds the text of each chunk, because that text goes into the prompt. Treat the index, its backups and the query logs as copies of the documents and keep them where the documents may be.
What does an on-premise GPU server for a private LLM need from the server room?
It needs a circuit and PDU outlets that carry its power draw, airflow or liquid cooling for the heat it produces, a floor that takes the weight and a network connection to the users and the document sources. Several GPU servers in one rack can draw more power than a rack planned for classic servers was designed for. Check the circuit rating, the outlets, the airflow and the GPU power cable before the order rather than at delivery.
What should a contract with a GPU hosting provider in the EU cover?
It should name the legal entity that runs the hardware, the data centre and its country, whether GPUs and disks are dedicated or shared, and who holds root, hypervisor and BMC access. Where the provider processes personal data for you, Article 28 GDPR requires a processor contract under which it engages subprocessors only with your prior specific or general written authorisation and, at your choice, deletes or returns the data at the end of the service and deletes existing copies. Add where logs and backups are kept, how disks are wiped and how the data comes back to you.

Send us the model you plan to run, where the documents it should answer from are stored, your peak number of concurrent users and whether you would host on-premise or in an EU data centre. We reply within one business day to arrange a first call, from which you leave with two or three possible solution scenarios. The first call is free of charge.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry – see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna