BLOG · GUIDE ·

Self-hosted LLM for a SaaS product: GPU servers for AI features, tenants and EU data residency

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • A SaaS company can serve open-weight models for its customer-facing AI features on its own GPU servers, in its own racks or in an EU data centre, so prompts and outputs stay in the EU and no model API provider joins its list of subprocessors
  • Size by requests in flight at the busiest minute per feature, the arrival rate times the time each request takes at its latency target, not by users or tenants; in our example 60 assistant and 30 summary requests a minute give about 55 in flight
  • For that peak on gpt-oss-120b at 32K, two servers with 4 RTX PRO 6000 or 2 H200 NVL each carry it with one server down and stay within an 80 per cent limit, by our estimates
  • Tenants on a shared model are kept apart by keys, quotas and logs in a gateway; vLLM’s cache salt, a random secret per tenant, closes the prefix cache timing channel but is not an isolation boundary; MIG gives up to 7 instances on an H200 NVL and 4 on an RTX PRO 6000, and the largest tenants get their own cards or server
  • A SaaS company that offers an AI feature under its own name can be the provider of that AI system under Article 3(3) of the AI Act, with Article 50 transparency duties where they fit, and where its product is a data processing service, the Data Act’s exportable data include output data generated by the customer’s use

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Self-hosted LLM for a SaaS product: when it fits

A European SaaS company that adds AI features for its customers can run a self-hosted LLM, serving open-weight models on its own GPU servers, on-premise or in an EU data centre, when customers require EU data residency and want no third-party model API in the data path. The sizing driver is the number of requests in flight at the busiest minute, per feature, held against a latency target for each feature, not the number of users or tenants.

Where a SaaS company processes personal data on its customers’ behalf, it is normally their processor, and Article 28(2) GDPR states that the processor “shall not engage another processor without prior specific or general written authorisation of the controller”. Under a general authorisation the processor informs the controller of an added processor, “thereby giving the controller the opportunity to object”. A model API provider that receives prompts containing personal data would normally be such an addition. Servers the SaaS company controls add no model API provider to the list of subprocessors; whether a data centre operator or a service provider with access to them is a processor depends on that access.

Sizing GPU servers for AI features: peak requests per feature

For each feature, multiply the arrival rate at the busiest minute by the time a request stays in flight, from arrival to its last token. That time follows from the latency target and the answer length, which differ per feature.

FEATURE TYPEEXAMPLE TARGET, P95REQUEST SHAPESERVING PATTERN
In-app assistantTTFT 1 s, TPOT 100 msshort prompts, answers of a few hundred tokensstreamed, shared model, prefix cache per tenant
Summary or draftTTFT 2.5 s, TPOT 100 msprompts of several thousand tokensstreamed, prefill-heavy, longer context
Agent or workflow stepTTFT 1 s, TPOT 50 msmany calls, growing contextfewer requests per card, time per task
Classification, extractionend-to-end time per calllong input, very short outputnot streamed, a smaller model may do
Bulk and backfill jobsdocuments per hourlong input, short outputoff-peak, own rate limit or own node

Example targets from our latency guide; request shapes and serving patterns are our examples, to be tested with your prompts.

Our guide to LLM latency targets explains time to first token (TTFT) and time per output token (TPOT). NVIDIA’s benchmarking guide says TTFT generally includes network latency, so the model servers belong close to the application servers that call them.

Worked example: three AI features on two GPU servers

The example is illustrative: a document management product with 300 business customers and 40,000 end users. At its busiest minute the assistant receives 60 requests, one per second. An answer of 400 tokens at the targets takes 1 + 399 × 0.1 s, about 41 s, so about 41 requests are in flight. The summary feature receives 30 requests a minute, and 250 tokens after a first token at 2.5 s take about 27 s, which adds about 14. A nightly extraction job runs outside working hours under its own rate limit. The interactive peak is about 55 requests in flight; taking the p95 targets as the time per request errs on the high side.

All features here run on gpt-oss-120b at a declared 32K context with a 16-bit KV cache. By the estimates in our sizing guides, that leaves room for about 19 conversations per RTX PRO 6000 Server Edition and 55 per H200 NVL, one copy of the model per card. Customers expect the feature to stay up during a server failure, so either server has to carry the peak alone, and our example policy lets the peak use at most 80 per cent of that capacity.

LAYOUTSESSIONS PER NODEPEAK, ONE NODE DOWNWITHIN 80 PER CENT
2 × 3 RTX PRO 60005796 per centno
2 × 4 RTX PRO 60007672 per centyes
2 × 1 H200 NVL55100 per centno
2 × 2 H200 NVL11050 per centyes

Illustrative example and our estimates: a peak of 55 requests in flight; sessions of gpt-oss-120b at 32K with a 16-bit KV cache, 19 per RTX PRO 6000 and 55 per H200 NVL, as in our capacity guides; 80 per cent is an example limit.

Memory sets one limit, and a load test with your own prompts at 55 requests in flight sets the other: the capacity of a server is the lower of the two.

We check the rack, power and airflow before we quote. Send us your features and their busiest-minute request rates through the form below.

Multi-tenant LLM serving: isolating tenants on shared GPUs

Most tenants share one model instance, and their isolation rests on the layers around it. A gateway in front of the model servers gives each tenant its own keys, token quotas and a tenant ID in every log entry, so one tenant’s bulk job cannot crowd out another tenant’s assistant. How keys, budgets and the query log work is set out in our guide to an LLM gateway with keys, budgets and logging.

A shared instance also shares its prefix cache. vLLM’s design documentation, dated 23 June 2026, describes a per-request cache_salt, “ensuring that only requests with the same salt can reuse cached KV blocks”, against timing attacks that infer cached content from latency differences. Its security guide, read on 10 October 2026, says to “Treat the salt as a secret” and to use random values “rather than predictable identifiers such as a user name or account ID”. Let the gateway hold a random secret per tenant, set it on every request and overwrite any salt a client sends. The guide also calls cache_salt “not an isolation boundary”, so tenants that must be isolated from each other get their own instance.

Where a tenant pays for a model adapted to its own terminology, vLLM serves LoRA adapters on one base model, and its model list marks gpt-oss as supported. Adapters “can be efficiently served on a per-request basis with minimal overhead”, and a request selects one through the model parameter. max_loras, the number of adapters in one batch, defaults to 1, so raise it to the number of tenant adapters in use at once. Loading adapters into a running server needs VLLM_ALLOW_RUNTIME_LORA_UPDATING, which vLLM says “should not be used in production unless it is an isolated, fully trusted environment”, so load tenant adapters at startup and roll them out like a model change.

NVIDIA’s MIG user guide, updated on 11 September 2026, lists up to 7 MIG instances on the H200 NVL and up to 4 on each edition of the RTX PRO 6000 Blackwell, each with “separate and isolated paths through the entire memory system”. The smallest profiles are 1g.24gb on the RTX PRO 6000 and 1g.18gb on the H200 NVL, which NVIDIA’s H200 page rates at 16.5 GB per instance, enough for one tenant’s embedding model or a model of a few billion parameters; gpt-oss-120b needs whole cards. All MIG instances use the host’s driver, so tenants that must not trust each other belong in separate virtual machines with a passthrough GPU or a vGPU. The L40S and L4 do not appear in NVIDIA’s list of MIG-capable GPUs.

TENANT PROFILEISOLATION METHODWHAT IT COSTS
Small tenants, standard useshared instance; keys, quotas, logs and a secret cache salt per tenantnothing beyond the shared pool
Tenant with its own adapterLoRA adapter on the shared base modeladapter memory; retraining at each base model change
Tenant with a small modelMIG instance on an H200 NVL or RTX PRO 6000a fixed slice, used or idle
Large tenant, heavy loadseparate vLLM instance on its own cardswhole cards for one tenant
Tenant needs own hardwarededicated server in its rack or an EU data centrea server for one tenant

vLLM security, prefix caching and LoRA documentation; NVIDIA MIG user guide and its supported GPUs page, read on 10 October 2026; the assignment to tenant profiles is our example.

Availability and model upgrades without breaking customers

Two nodes that each carry the peak, behind health checks, are the minimum for an AI feature that survives the loss of a server; two sites also cover the loss of a room. Our guide to LLM high availability with two GPU nodes covers the health checks and the update of one node at a time.

Customers build on a model’s output format and tone, so a model change is a release. vLLM’s --revision takes “a branch name, a tag name, or a commit id”, which pins the exact weights from the Hugging Face Hub, --tokenizer-revision does the same for the tokenizer, and --served-model-name sets “the model name(s) used in the API”. Give each model version its own served name and keep a moving alias only for tenants that opt into automatic upgrades.

  1. Pin the new model by commit, or by a versioned directory in a local model store, and run it as a separate instance under a new served name.
  2. Run each tenant’s evaluation set, its prompts and expected output formats, against the old and the new version.
  3. Retrain the tenants’ LoRA adapters on the new base model, since an adapter belongs to the base model it was trained on.
  4. Announce the change with a period in which both versions run, and move tenants over in batches through the gateway.
  5. Remove the old version once no tenant key calls it.

SaaS AI data residency: own racks or an EU data centre

Self-hosting decides where prompts, outputs, the KV cache, the query log and any fine-tuning data are kept. The servers can stand in racks the SaaS company runs, or as dedicated servers in an EU data centre. In the second case the questions move to who holds root and BMC access and where logs and backups are stored, as our guide to private LLM inference in an EU data centre sets out. Renting GPU instances is a third route, compared with own servers in our article on cloud GPU rental vs your own GPU server.

In our Private AI/ML service the models run on your servers or on dedicated hardware in a Tier-3 data centre in Lithuania, and nothing goes to public services unless you explicitly enable it. Tell us where your tenants require their data to stay.

AI Act and Data Act duties for a SaaS company with AI features

Under Article 3(3) of the AI Act, a provider develops an AI system, or has one developed, and places it on the market or puts it into service “under its own name or trademark, whether for payment or free of charge”. A SaaS company that offers an assistant built on an open-weight model under its own name can fit that definition. Its business customers using the feature are deployers under Article 3(4). Since 2 August 2026, Article 50 has required providers to design systems that interact directly with people so that they are informed they are interacting with an AI system, unless this is obvious, and to mark generated text in a machine-readable format, except for “an assistive function for standard editing”. For generative systems placed on the market before 2 August 2026, the AI Act as amended by Regulation (EU) 2026/1744 sets 2 December 2026 for that marking (Article 111(4)).

The Data Act’s switching chapter binds providers of data processing services, and its recital 81 names “Infrastructure as a Service (IaaS), Platform as a service (PaaS) and Software as a Service (SaaS)”. Exportable data under Article 2(38) is “the input and output data, including metadata” generated by the customer’s use, so summaries and extracted fields stored for a customer can fall within it, unless they are protected by intellectual property rights or are trade secrets of the provider or third parties. From 12 January 2027 providers may impose no switching charges (Article 29(1)), except for services mostly custom-built for one customer and not offered at broad commercial scale (Article 31(1)). As of 6 October 2026 the Commission’s Digital Omnibus proposal to amend this chapter had not been adopted. Which role and which duties apply to a given product and contract is a legal assessment for your legal department.

What we supply

We build AI servers to order with 2 to 8 GPUs per node, sized by model size and concurrent users, with manufacturer warranty on one EU contract and invoice. The cards for AI features in a SaaS product are the RTX PRO 6000 Server Edition with 96 GB, the H200 NVL with 141 GB of HBM3e and NVLink bridges, and the L40S and L4 for smaller models, from our range of professional NVIDIA GPUs. The platform on top, with private LLMs, a query log and hosting on dedicated hardware in a Tier-3 data centre in Lithuania, is our Private AI/ML service, with engineering by our partner Vixen.UNO and support under an agreed SLA.

FAQ

What does a self-hosted LLM for a SaaS product need?
It needs GPU servers sized by the peak requests in flight per AI feature at that feature’s latency target, a gateway that gives each tenant keys, quotas and logs, and at least two nodes that each carry the peak. A release process for model changes, with pinned model versions and tenant evaluation sets, keeps customer integrations working when the model changes.
How do you do multi-tenant LLM serving?
Most tenants share one model instance behind a gateway that enforces keys, quotas and a tenant ID in every log entry and sets vLLM’s per-request cache salt to a random secret per tenant, so tenants cannot reuse each other’s cached prompts. Tenants with their own adapted model get a LoRA adapter on the shared base model. vLLM’s security guide says a shared server process does not isolate tenants, so tenants that must be isolated get their own vLLM instance on a MIG instance or their own cards, or their own server.
How many GPUs do AI features in a SaaS product need?
Count requests in flight at the busiest minute: arrival rate times the time each request takes at its latency target. In our example, 60 assistant and 30 summary requests a minute give about 55 in flight, and on gpt-oss-120b at 32K two servers with 4 RTX PRO 6000 or 2 H200 NVL each carry that with one server down, by our estimates. A load test with your own prompts confirms the figure.
Can a SaaS company host an LLM for its customers in the EU?
Yes, on GPU servers in its own racks or on dedicated servers in an EU data centre, so prompts, outputs and logs stay in the EU. In a data centre, agree who holds root and BMC access and where logs and backups are stored. The model servers should sit close to the application servers, because network latency adds to the time to first token.
How does SaaS AI data residency work with a model API?
A model API provider that receives prompts containing personal data on your behalf is normally an additional processor, and under Article 28(2) GDPR customers who gave a general authorisation must be informed and can object. Self-hosting the model on servers you control keeps prompts, outputs and logs in the EU and adds no model API provider, although a data centre operator or service provider with access to the servers may be a processor. How this applies to your contracts is a legal assessment for your legal department.
Is a SaaS company a provider or a deployer under the AI Act?
A SaaS company that develops an AI feature and offers it under its own name can fit the AI Act’s definition of a provider in Article 3(3), and its business customers using the feature are deployers under Article 3(4). Since 2 August 2026, Article 50 requires providers to inform people that they are interacting with an AI system unless this is obvious, and to mark generated content in a machine-readable format, a duty the amended law applies from 2 December 2026 to generative systems placed on the market before 2 August 2026.

Send us the AI features you plan, the requests per feature at your busiest minute, the model and context length, the number of tenants and any tenant that needs its own hardware. We reply within one business day with a configuration and quote for the GPU servers those figures point to.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna