BLOG · GUIDE ·

Private LLM for hospitals on-premise: clinical documentation, health data and GPU server sizing

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Hospitals keep LLM workloads such as discharge letter drafts, consultation transcripts, guideline search and coding support on-premise because prompts, documents and logs carry data concerning health, a special category under Article 9 of the GDPR
  • Article 9(2)(h) of the GDPR covers processing for medical diagnosis and the provision of health care under the conditions of Article 9(3), and Article 35(3)(b) names large-scale processing of special categories among the cases that require a DPIA
  • Software is a medical device under the MDR when its manufacturer intends it for a medical purpose such as diagnosis; Rule 11 puts software that informs diagnostic or therapeutic decisions in class IIa or higher, and MDCG 2025-6 treats such AI devices assessed by a notified body as high-risk under the AI Act
  • With gpt-oss-120b at a declared 32K context and the example values of our sizing method, a hospital that gives 1,500 staff access reaches a peak of 60 requests in flight, carried by two servers with four RTX PRO 6000 or two H200 NVL each
  • The LLM servers belong in a network zone of their own, apart from the imaging AI servers that take studies from the PACS, and the logs of prompts and answers carry health data, so access to them is restricted and their retention period defined

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Private LLM for hospitals: what runs on-premise

A hospital with 1,000 to 5,000 staff can run its LLM workloads on-premise, on its own GPU servers, which keeps every prompt, retrieved document and log entry inside its own network. Clinical text carries data concerning health, which the GDPR defines in Article 4(15) as personal data “related to the physical or mental health of a natural person”, and whose processing Article 9(1) prohibits unless an exception in Article 9(2) applies. The usual workloads are drafts of discharge letters and other documentation, transcription of consultations, guideline search and coding support. With the example values in this article, two servers with four RTX PRO 6000 or two H200 NVL each carry a hospital that gives 1,500 staff access, and either server can fail without stopping the service.

Clinical documentation workloads and model classes

A discharge letter draft starts from the notes, laboratory results and medication list of a stay, which the integration with the electronic health record (EHR) passes to the model; a physician edits and signs the draft. These prompts are long, so we declare a 32K context for them in the sizing below (example value). Transcription turns consultation audio into text with a speech recognition model, and the LLM structures it into a note; our guide to speech-to-text servers for Whisper and Parakeet sizes that part. Guideline search is retrieval-augmented generation (RAG): an embedding model and a reranker find passages in the guideline index, and the LLM answers with the source cited. Coding support suggests diagnosis and procedure codes from a finished letter for coding staff to review.

WORKLOADMODEL CLASSCONTEXTGPU (OUR RANGE)
Discharge letter draftsgeneral open model, such as gpt-oss-120b32K (example)RTX PRO 6000 Server Edition, H200 NVL
Consultation transcriptionspeech recognition, then the same LLMaudio, then 8K to 32KL4, L40S for speech; LLM cards as above
Guideline search (RAG)0.6B embedding and reranker, plus the LLMretrieved passagesL4 or a 24 GB MIG instance
Coding supportthe same LLM, in batchthe finished letterthe same cards, outside busy hours
Health-tuned model trialsMedGemma 27B text, BF16up to 128K inputRTX PRO 6000 Server Edition

gpt-oss-120b and MedGemma 27B model cards on Hugging Face, read on 10 October 2026; embedding models and MIG as in our company-size sizing guide; contexts and cards are our examples.

gpt-oss-120b is a 65.3 GB checkpoint that fits one 96 GB card. Google’s card for the health-tuned MedGemma 27B says the model “has been trained exclusively on medical text”, gives a total input length of 128K tokens and places its use under the Health AI Developer Foundations terms of use. In BF16 its 27B parameters take about 54 GB, by our arithmetic, which fits one RTX PRO 6000 with room for cache. The card states that its outputs “are not intended to directly inform clinical diagnosis, patient management decisions”, treatment recommendations or other direct clinical practice applications. Compare both on an evaluation set of anonymised letters before choosing.

Health data under GDPR Article 9 and the DPIA

Article 9(2)(h) of the GDPR allows processing that is necessary for purposes that include “medical diagnosis, the provision of health or social care or treatment”, and Article 9(3) ties this to data processed by or under the responsibility of “a professional subject to the obligation of professional secrecy”. Once a prompt carries patient information, the retrieved passages, answers, caches and gateway logs that go with it carry health data too, wherever they are stored.

Article 35(1) requires the controller to carry out a data protection impact assessment (DPIA) before processing that “is likely to result in a high risk to the rights and freedoms of natural persons”, and Article 35(3)(b) names “processing on a large scale of special categories of data referred to in Article 9(1)” among the cases. Article 32(1) lists measures “as appropriate”, including “the pseudonymisation and encryption of personal data” and “the ability to restore the availability and access to personal data in a timely manner”. IT supplies the data flows, where the model runs, retention and the access model to the DPIA; our guide to a DPIA for an internal LLM lists each input.

When software becomes a medical device under the MDR

Under Article 2(1) of the Medical Device Regulation (EU) 2017/745, software is a medical device when its manufacturer intends it for purposes that include “diagnosis, prevention, monitoring, prediction, prognosis, treatment or alleviation of disease”. By Article 2(12), the intended purpose is “the use for which a device is intended according to the data supplied by the manufacturer”. Recital 19 adds that “software for general purposes, even when used in a healthcare setting,” is not a medical device.

Rule 11 in Annex VIII, as reproduced in the guidance MDCG 2019-11 rev.1 of June 2025, classifies “Software intended to provide information which is used to take decisions with diagnosis or therapeutic purposes” as class IIa, as class IIb where such decisions may cause a serious deterioration of health or a surgical intervention, and as class III where they may cause death or an irreversible deterioration. It ends with “All other software is classified as class I.” The same guidance says that hospital information systems “are not in themselves qualified as medical devices” and that software performing only storage, archival, communication or “simple search” does not qualify. It adds that “modules integrated into or operating alongside EHR systems may qualify as MDSW if they fulfil a medical purpose”. The Commission’s proposal of 16 December 2025 to amend the MDR (COM(2025) 1023 final) adapts some classification rules, software among them, “resulting in lower risk classes for certain devices”. Arnold & Porter wrote on 7 August 2026 that the Parliament’s plenary vote on its position was expected “towards or shortly after the New Year”; the rule quoted here is the one in force.

Whether a given tool is a medical device, which class it falls in, whether the AI Act treats it as high-risk and on which legal basis health data is processed are assessments for the hospital’s legal and regulatory functions; the choice of server does not change them. For IT, this means recording the intended purpose of each tool and keeping a drafting assistant apart from any tool with a declared medical purpose.

AI Act risk classes and the European Health Data Space

Article 6(1) of the AI Act makes an AI system high-risk when it is, or is a safety component of, a product under the legislation in Annex I and that product needs a third-party conformity assessment. MDCG 2025-6, endorsed in June 2025 by the MDCG and the AI Board, applies this to AI medical devices where “the MDAI is subject to a third-party conformity assessment by a notified body in accordance with the MDR/IVDR”. Under Article 113 as amended by Regulation (EU) 2026/1744, these rules apply from 2 August 2028, and those for Annex III uses from 2 December 2027. Annex III, point 5(d), lists “emergency healthcare patient triage systems”. A drafting or search assistant outside these categories carries the AI literacy and transparency duties that our article on AI Act obligations for companies running LLMs explains.

The European Health Data Space regulation, Regulation (EU) 2025/327 of 11 February 2025, entered into force on 26 March 2025. The Commission’s EHDS page says it “introduces strict security and interoperability criteria for EHR systems”. By the page’s timeline, the exchange of patient summaries and ePrescriptions applies in all Member States from March 2029, and the exchange of “medical images, lab results, and hospital discharge reports” should be operational in all of them by March 2031. The assistant writes its drafts into the EHR, which stays the system of record. NIS2 lists healthcare providers in Annex I, and where a hospital falls within its scope, the Article 21 measures cover the AI servers as part of its network and information systems.

REGULATIONWHAT THE TEXT SAYSFOR THE LLM PLATFORM
GDPR Article 9data concerning health “shall be prohibited” unless 9(2) applies; 9(2)(h) for “medical diagnosis” and careprompts, answers, caches and logs treated as health data
GDPR Article 35DPIA for “processing on a large scale of special categories of data”data flows, model location and retention documented
GDPR Article 32“the pseudonymisation and encryption of personal data”encrypted storage, access control, tested restore
MDR Art. 2(1), Rule 11software for “diagnosis, prevention, monitoring, prediction, prognosis, treatment”intended purpose recorded per tool
AI Act Art. 6(1), Annex IIIAnnex I products from 2 August 2028; triage systems listed in 5(d)logs and human oversight where a use is high-risk
NIS2 Article 21“access control policies” and “multi-factor authentication”MFA at the gateway, backup of models and index

Regulations (EU) 2016/679, 2017/745 and 2024/1689 (Article 113 as amended by 2026/1744) and Directive (EU) 2022/2555, Article 21(2)(i) and (j), on eur-lex.europa.eu as of October 2026; general information only.

Sizing the GPU servers for 1,000 to 5,000 staff

We size from the peak number of requests in flight, with the method and example values of our guide to private ChatGPT server sizing by company size: 40 per cent of the users with access active in the busiest hour, 6 requests each per hour, 30 seconds per request and a peak factor of 2. We assume that half of the staff get access (example value). By the estimate in that guide, gpt-oss-120b at a declared 32K context with a 16-bit cache holds about 19 conversations on one RTX PRO 6000 and about 55 on one H200 NVL.

STAFFWITH ACCESSPEAK IN FLIGHTEACH OF TWO SERVERSSESSIONS PER SERVER
1,000500202 × RTX PRO 6000 Server Edition, 1 × L438
3,0001,500604 × RTX PRO 6000, 1 × L4; or 2 × H200 NVL, 1 × L476 or 110
5,0002,5001006 × RTX PRO 6000, 1 × L4; or 2 × H200 NVL, 1 × L4114 or 110

Example values from our company-size sizing guide plus 50 per cent of staff with access (example value); sessions per server from its estimate for gpt-oss-120b at 32K; the L4 carries the embedding model and reranker; speech recognition cards come on top.

For 3,000 staff, 1,500 users with access give 600 in the busiest hour and 3,600 requests, one per second, so 30 are in flight on average and 60 at the example peak. Each server holds the whole peak, so the service survives the loss of one. With these example values, a 5,000-person hospital needs six RTX PRO 6000 per server, or two H200 NVL. Another model changes the per-card figures, so repeat the arithmetic with its weights and cache.

We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us your staff numbers, workloads and the models you are evaluating through the form below.

Separating the LLM platform from imaging AI and clinical networks

A radiology inference server takes 3D studies from the PACS, and its GPU memory follows the volume size, as our guide to the medical imaging AI server sets out. Separate servers let each be updated, tested and documented on its own, which matters where only one tool has a declared medical purpose.

Place the LLM servers in a network zone of their own. Users reach only the gateway, with single sign-on and multi-factor authentication. The integration engine is the one system that passes patient records to the model, and retrieval checks each user’s rights in the source system before a passage reaches the prompt. The model servers need no outbound internet access; checkpoints come in through a controlled import with checksums, and the management controllers sit on the management network.

The gateway log records the user, time, model, prompt and answer. It carries health data, so access to it is restricted and its retention is set in the DPIA. For a high-risk use, Article 26(6) of the AI Act asks deployers to keep the logs for a period “of at least six months”.

Our Private AI/ML service deploys private LLMs and RAG with query logging and each user’s access rights, with engineering by our partner Vixen.UNO. Describe your workloads and the network zone in the form below.

What we supply

We supply the cards that a hospital LLM platform uses, the H200 NVL, RTX PRO 6000 Server Edition, L40S and L4, as cards or in AI servers built to order, assembled and burn-in tested, with manufacturer warranty, on one EU contract and invoice; the GPU range lists each card. NVIDIA AI Enterprise and vGPU licences come on the same invoice. We check the rack, power and airflow before we quote. Deployment of models, RAG, logging and access control on top is our Private AI/ML service, on your servers or in a Tier-3 data centre in Lithuania, with engineering by our partner Vixen.UNO.

FAQ

Can a hospital run an LLM on premise?
Yes. An open model such as gpt-oss-120b fits one 96 GB RTX PRO 6000 or one H200 NVL, and with the example values of our sizing method two servers with four RTX PRO 6000 each carry a hospital that gives 1,500 staff access, with one server able to fail. Keeping the platform on-premise keeps prompts, documents and logs, which carry health data, inside the hospital network.
Can health data be used in an LLM under the GDPR?
Data concerning health is a special category under Article 9(1) of the GDPR, and its processing is prohibited unless an exception in Article 9(2) applies. Point (h) covers processing necessary for medical diagnosis and the provision of health care, under the professional secrecy conditions of Article 9(3). Once prompts carry patient information, the retrieved passages, answers and gateway logs of a clinical assistant carry health data as well.
Is an LLM for clinical documentation a medical device?
It depends on the intended purpose its manufacturer states. Article 2(1) of the MDR covers software intended for purposes such as diagnosis, prediction, prognosis or treatment of disease, while recital 19 excludes software for general purposes used in a healthcare setting. MDCG 2019-11 rev.1 says hospital information systems are not medical devices in themselves, and modules alongside EHR systems may qualify if they fulfil a medical purpose.
Does a hospital LLM need a DPIA?
Article 35(1) of the GDPR requires a data protection impact assessment before processing that is likely to result in a high risk, and Article 35(3)(b) names large-scale processing of special categories such as health data among the cases. The IT team supplies the data flows, the location of the model, retention periods and the access model.
What AI server does a hospital need for a private LLM?
For 1,000 to 5,000 staff, two servers that each hold the whole peak: two to six RTX PRO 6000 Server Edition cards or two H200 NVL per server, plus an L4 for the embedding model and reranker, with the example values of our sizing method and gpt-oss-120b at 32K. Speech recognition for consultation transcripts comes on top and is sized from the hours of audio and live streams.
Is a hospital LLM high-risk under the EU AI Act?
Annex III does not list drafting or search assistants, while it lists emergency healthcare patient triage systems, high-risk from 2 December 2027 under the amended law. MDCG 2025-6 treats an AI medical device that needs a notified body as high-risk under Article 6(1), and Article 113 as amended by Regulation (EU) 2026/1744 applies those rules from 2 August 2028. Measures to support AI literacy (Article 4) apply already, and Article 50 on transparency where the use fits.

Send us the number of staff who would use the assistant, the workloads you plan (letters, transcription, guideline search, coding support), the models you are evaluating and the network zone the servers would sit in. We reply within one business day with a configuration and quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna