Private LLM for hospitals on-premise: clinical documentation, health data and GPU server sizing
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Hospitals keep LLM workloads such as discharge letter drafts, consultation transcripts, guideline search and coding support on-premise because prompts, documents and logs carry data concerning health, a special category under Article 9 of the GDPR
- Article 9(2)(h) of the GDPR covers processing for medical diagnosis and the provision of health care under the conditions of Article 9(3), and Article 35(3)(b) names large-scale processing of special categories among the cases that require a DPIA
- Software is a medical device under the MDR when its manufacturer intends it for a medical purpose such as diagnosis; Rule 11 puts software that informs diagnostic or therapeutic decisions in class IIa or higher, and MDCG 2025-6 treats such AI devices assessed by a notified body as high-risk under the AI Act
- With gpt-oss-120b at a declared 32K context and the example values of our sizing method, a hospital that gives 1,500 staff access reaches a peak of 60 requests in flight, carried by two servers with four RTX PRO 6000 or two H200 NVL each
- The LLM servers belong in a network zone of their own, apart from the imaging AI servers that take studies from the PACS, and the logs of prompts and answers carry health data, so access to them is restricted and their retention period defined
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Private LLM for hospitals: what runs on-premise
A hospital with 1,000 to 5,000 staff can run its LLM workloads on-premise, on its own GPU servers, which keeps every prompt, retrieved document and log entry inside its own network. Clinical text carries data concerning health, which the GDPR defines in Article 4(15) as personal data “related to the physical or mental health of a natural person”, and whose processing Article 9(1) prohibits unless an exception in Article 9(2) applies. The usual workloads are drafts of discharge letters and other documentation, transcription of consultations, guideline search and coding support. With the example values in this article, two servers with four RTX PRO 6000 or two H200 NVL each carry a hospital that gives 1,500 staff access, and either server can fail without stopping the service.
Clinical documentation workloads and model classes
A discharge letter draft starts from the notes, laboratory results and medication list of a stay, which the integration with the electronic health record (EHR) passes to the model; a physician edits and signs the draft. These prompts are long, so we declare a 32K context for them in the sizing below (example value). Transcription turns consultation audio into text with a speech recognition model, and the LLM structures it into a note; our guide to speech-to-text servers for Whisper and Parakeet sizes that part. Guideline search is retrieval-augmented generation (RAG): an embedding model and a reranker find passages in the guideline index, and the LLM answers with the source cited. Coding support suggests diagnosis and procedure codes from a finished letter for coding staff to review.
| WORKLOAD | MODEL CLASS | CONTEXT | GPU (OUR RANGE) |
|---|---|---|---|
| Discharge letter drafts | general open model, such as gpt-oss-120b | 32K (example) | RTX PRO 6000 Server Edition, H200 NVL |
| Consultation transcription | speech recognition, then the same LLM | audio, then 8K to 32K | L4, L40S for speech; LLM cards as above |
| Guideline search (RAG) | 0.6B embedding and reranker, plus the LLM | retrieved passages | L4 or a 24 GB MIG instance |
| Coding support | the same LLM, in batch | the finished letter | the same cards, outside busy hours |
| Health-tuned model trials | MedGemma 27B text, BF16 | up to 128K input | RTX PRO 6000 Server Edition |
gpt-oss-120b and MedGemma 27B model cards on Hugging Face, read on 10 October 2026; embedding models and MIG as in our company-size sizing guide; contexts and cards are our examples.
gpt-oss-120b is a 65.3 GB checkpoint that fits one 96 GB card. Google’s card for the health-tuned MedGemma 27B says the model “has been trained exclusively on medical text”, gives a total input length of 128K tokens and places its use under the Health AI Developer Foundations terms of use. In BF16 its 27B parameters take about 54 GB, by our arithmetic, which fits one RTX PRO 6000 with room for cache. The card states that its outputs “are not intended to directly inform clinical diagnosis, patient management decisions”, treatment recommendations or other direct clinical practice applications. Compare both on an evaluation set of anonymised letters before choosing.
Health data under GDPR Article 9 and the DPIA
Article 9(2)(h) of the GDPR allows processing that is necessary for purposes that include “medical diagnosis, the provision of health or social care or treatment”, and Article 9(3) ties this to data processed by or under the responsibility of “a professional subject to the obligation of professional secrecy”. Once a prompt carries patient information, the retrieved passages, answers, caches and gateway logs that go with it carry health data too, wherever they are stored.
Article 35(1) requires the controller to carry out a data protection impact assessment (DPIA) before processing that “is likely to result in a high risk to the rights and freedoms of natural persons”, and Article 35(3)(b) names “processing on a large scale of special categories of data referred to in Article 9(1)” among the cases. Article 32(1) lists measures “as appropriate”, including “the pseudonymisation and encryption of personal data” and “the ability to restore the availability and access to personal data in a timely manner”. IT supplies the data flows, where the model runs, retention and the access model to the DPIA; our guide to a DPIA for an internal LLM lists each input.
When software becomes a medical device under the MDR
Under Article 2(1) of the Medical Device Regulation (EU) 2017/745, software is a medical device when its manufacturer intends it for purposes that include “diagnosis, prevention, monitoring, prediction, prognosis, treatment or alleviation of disease”. By Article 2(12), the intended purpose is “the use for which a device is intended according to the data supplied by the manufacturer”. Recital 19 adds that “software for general purposes, even when used in a healthcare setting,” is not a medical device.
Rule 11 in Annex VIII, as reproduced in the guidance MDCG 2019-11 rev.1 of June 2025, classifies “Software intended to provide information which is used to take decisions with diagnosis or therapeutic purposes” as class IIa, as class IIb where such decisions may cause a serious deterioration of health or a surgical intervention, and as class III where they may cause death or an irreversible deterioration. It ends with “All other software is classified as class I.” The same guidance says that hospital information systems “are not in themselves qualified as medical devices” and that software performing only storage, archival, communication or “simple search” does not qualify. It adds that “modules integrated into or operating alongside EHR systems may qualify as MDSW if they fulfil a medical purpose”. The Commission’s proposal of 16 December 2025 to amend the MDR (COM(2025) 1023 final) adapts some classification rules, software among them, “resulting in lower risk classes for certain devices”. Arnold & Porter wrote on 7 August 2026 that the Parliament’s plenary vote on its position was expected “towards or shortly after the New Year”; the rule quoted here is the one in force.
Whether a given tool is a medical device, which class it falls in, whether the AI Act treats it as high-risk and on which legal basis health data is processed are assessments for the hospital’s legal and regulatory functions; the choice of server does not change them. For IT, this means recording the intended purpose of each tool and keeping a drafting assistant apart from any tool with a declared medical purpose.
AI Act risk classes and the European Health Data Space
Article 6(1) of the AI Act makes an AI system high-risk when it is, or is a safety component of, a product under the legislation in Annex I and that product needs a third-party conformity assessment. MDCG 2025-6, endorsed in June 2025 by the MDCG and the AI Board, applies this to AI medical devices where “the MDAI is subject to a third-party conformity assessment by a notified body in accordance with the MDR/IVDR”. Under Article 113 as amended by Regulation (EU) 2026/1744, these rules apply from 2 August 2028, and those for Annex III uses from 2 December 2027. Annex III, point 5(d), lists “emergency healthcare patient triage systems”. A drafting or search assistant outside these categories carries the AI literacy and transparency duties that our article on AI Act obligations for companies running LLMs explains.
The European Health Data Space regulation, Regulation (EU) 2025/327 of 11 February 2025, entered into force on 26 March 2025. The Commission’s EHDS page says it “introduces strict security and interoperability criteria for EHR systems”. By the page’s timeline, the exchange of patient summaries and ePrescriptions applies in all Member States from March 2029, and the exchange of “medical images, lab results, and hospital discharge reports” should be operational in all of them by March 2031. The assistant writes its drafts into the EHR, which stays the system of record. NIS2 lists healthcare providers in Annex I, and where a hospital falls within its scope, the Article 21 measures cover the AI servers as part of its network and information systems.
| REGULATION | WHAT THE TEXT SAYS | FOR THE LLM PLATFORM |
|---|---|---|
| GDPR Article 9 | data concerning health “shall be prohibited” unless 9(2) applies; 9(2)(h) for “medical diagnosis” and care | prompts, answers, caches and logs treated as health data |
| GDPR Article 35 | DPIA for “processing on a large scale of special categories of data” | data flows, model location and retention documented |
| GDPR Article 32 | “the pseudonymisation and encryption of personal data” | encrypted storage, access control, tested restore |
| MDR Art. 2(1), Rule 11 | software for “diagnosis, prevention, monitoring, prediction, prognosis, treatment” | intended purpose recorded per tool |
| AI Act Art. 6(1), Annex III | Annex I products from 2 August 2028; triage systems listed in 5(d) | logs and human oversight where a use is high-risk |
| NIS2 Article 21 | “access control policies” and “multi-factor authentication” | MFA at the gateway, backup of models and index |
Regulations (EU) 2016/679, 2017/745 and 2024/1689 (Article 113 as amended by 2026/1744) and Directive (EU) 2022/2555, Article 21(2)(i) and (j), on eur-lex.europa.eu as of October 2026; general information only.
Sizing the GPU servers for 1,000 to 5,000 staff
We size from the peak number of requests in flight, with the method and example values of our guide to private ChatGPT server sizing by company size: 40 per cent of the users with access active in the busiest hour, 6 requests each per hour, 30 seconds per request and a peak factor of 2. We assume that half of the staff get access (example value). By the estimate in that guide, gpt-oss-120b at a declared 32K context with a 16-bit cache holds about 19 conversations on one RTX PRO 6000 and about 55 on one H200 NVL.
| STAFF | WITH ACCESS | PEAK IN FLIGHT | EACH OF TWO SERVERS | SESSIONS PER SERVER |
|---|---|---|---|---|
| 1,000 | 500 | 20 | 2 × RTX PRO 6000 Server Edition, 1 × L4 | 38 |
| 3,000 | 1,500 | 60 | 4 × RTX PRO 6000, 1 × L4; or 2 × H200 NVL, 1 × L4 | 76 or 110 |
| 5,000 | 2,500 | 100 | 6 × RTX PRO 6000, 1 × L4; or 2 × H200 NVL, 1 × L4 | 114 or 110 |
Example values from our company-size sizing guide plus 50 per cent of staff with access (example value); sessions per server from its estimate for gpt-oss-120b at 32K; the L4 carries the embedding model and reranker; speech recognition cards come on top.
For 3,000 staff, 1,500 users with access give 600 in the busiest hour and 3,600 requests, one per second, so 30 are in flight on average and 60 at the example peak. Each server holds the whole peak, so the service survives the loss of one. With these example values, a 5,000-person hospital needs six RTX PRO 6000 per server, or two H200 NVL. Another model changes the per-card figures, so repeat the arithmetic with its weights and cache.
We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us your staff numbers, workloads and the models you are evaluating through the form below.
Separating the LLM platform from imaging AI and clinical networks
A radiology inference server takes 3D studies from the PACS, and its GPU memory follows the volume size, as our guide to the medical imaging AI server sets out. Separate servers let each be updated, tested and documented on its own, which matters where only one tool has a declared medical purpose.
Place the LLM servers in a network zone of their own. Users reach only the gateway, with single sign-on and multi-factor authentication. The integration engine is the one system that passes patient records to the model, and retrieval checks each user’s rights in the source system before a passage reaches the prompt. The model servers need no outbound internet access; checkpoints come in through a controlled import with checksums, and the management controllers sit on the management network.
The gateway log records the user, time, model, prompt and answer. It carries health data, so access to it is restricted and its retention is set in the DPIA. For a high-risk use, Article 26(6) of the AI Act asks deployers to keep the logs for a period “of at least six months”.
Our Private AI/ML service deploys private LLMs and RAG with query logging and each user’s access rights, with engineering by our partner Vixen.UNO. Describe your workloads and the network zone in the form below.
What we supply
We supply the cards that a hospital LLM platform uses, the H200 NVL, RTX PRO 6000 Server Edition, L40S and L4, as cards or in AI servers built to order, assembled and burn-in tested, with manufacturer warranty, on one EU contract and invoice; the GPU range lists each card. NVIDIA AI Enterprise and vGPU licences come on the same invoice. We check the rack, power and airflow before we quote. Deployment of models, RAG, logging and access control on top is our Private AI/ML service, on your servers or in a Tier-3 data centre in Lithuania, with engineering by our partner Vixen.UNO.
FAQ
Can a hospital run an LLM on premise?
Can health data be used in an LLM under the GDPR?
Is an LLM for clinical documentation a medical device?
Does a hospital LLM need a DPIA?
What AI server does a hospital need for a private LLM?
Is a hospital LLM high-risk under the EU AI Act?
Send us the number of staff who would use the assistant, the workloads you plan (letters, transcription, guideline search, coding support), the models you are evaluating and the network zone the servers would sit in. We reply within one business day with a configuration and quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day