BLOG · GUIDE ·

GPU servers for public administration: on-premise LLMs for authorities and municipalities

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Public bodies run LLMs for drafting and summaries, triage of citizen enquiries, translation between EU languages, meeting transcription and records search; on-premise or dedicated hosting keeps citizen data on hardware the authority controls
  • Mistral Small 3.2 (24B, Apache 2.0) is a minor update of 3.1, whose card names German, Polish and Romanian among 25 languages but not Czech or Hungarian; EuroLLM-22B covers all 24 official EU languages, and Parakeet TDT 0.6B v3 transcribes 25 European languages
  • Under the AI Act as amended by Regulation (EU) 2026/1744, a public body that deploys an Annex III high-risk system, such as one evaluating eligibility for public assistance benefits (point 5(a)), performs a fundamental rights impact assessment under Article 27 from 2 December 2027
  • With Mistral Small 3.2 in BF16 and an FP8 cache, one RTX PRO 6000 Server Edition holds about 15 conversations at 32K and one H200 NVL about 31 by our estimate
  • At example usage values, 1,000 staff peak at about 40 requests in flight, served by two servers with four RTX PRO 6000 each; 2,000 staff peak at about 80, served by two servers with four H200 NVL or eight RTX PRO 6000 each

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

On-premise LLMs for authorities and municipalities

In public administration, authorities and municipalities use LLMs for five kinds of work: drafting and summarising letters, reports and minutes; sorting incoming citizen enquiries to the right department; translating between EU languages; transcribing meetings and hearings; and searching their own records. All five touch citizen data or internal files, which is why many public bodies prefer to run the models on their own servers or on dedicated hardware in an EU data centre rather than on a public service. A GPU server for this runs open-weight models with an engine such as vLLM, and the cards are sized by the number of staff who use it at the same time, the context length and the languages the models must handle.

For a body with 200 to 2,000 staff the result is one to two servers with two to eight cards each. If a use falls under Annex III of the AI Act, the authority also has legal duties as a deployer, set out further below.

Workloads, model classes and cards

The drafting model is the largest single item. A multilingual model of 20B to 30B parameters handles letters, summaries and the routing of enquiries in one deployment, so triage does not need a model of its own. Records search adds an embedding model and a reranker; Qwen3-Embedding-0.6B, for example, has 0.6B parameters, supports “100+ Languages” and a 32K context under Apache 2.0. Speech recognition is small as well: OpenAI gives about 10 GB of VRAM for Whisper large.

WORKLOADMODEL CLASSEXAMPLESCARD
Drafting and summariesmultilingual LLM, 20B to 30BMistral Small 3.2, EuroLLM-22BRTX PRO 6000 Server Edition, H200 NVL
Triage of enquiriesthe drafting LLM with a fixed prompt and categoriesas aboveshared with drafting
TranslationLLM or translation model, 9B to 24BEuroLLM-9B, EuroLLM-22B, Mistral Small 3.2L40S, RTX PRO 6000
Meeting transcriptionspeech recognition, under 2BWhisper large-v3, Parakeet TDT 0.6B v3L4, or a MIG instance
Records searchembedding and reranker, under 1BQwen3-Embedding-0.6B and its rerankerL4, or a MIG instance

Model cards on Hugging Face and OpenAI’s Whisper repository, read on 10 October 2026; card choice is our reading of the memory each model needs.

Small models do not need a whole card. NVIDIA’s MIG user guide, updated on 11 September 2026, lists up to four MIG instances on the RTX PRO 6000 Server Edition and seven on the H200 NVL, each with its own memory. The L4 has 24 GB at 72 W in a single-slot, low-profile card and the L40S 48 GB, and neither is in the guide’s list of MIG-capable GPUs. A card in MIG mode can carry the embedding model, the reranker and the speech model side by side, while the drafting model keeps whole cards.

Models for EU languages and their licences

The language list on the model card decides more in public administration than in most companies, since an authority in Central and Eastern Europe writes in its national language and answers citizens in others. The Mistral Small 3.1 card names 25 languages, among them German, Polish and Romanian but not Czech, Hungarian, Croatian, Slovak or Bulgarian, and the 3.2 card calls its model “a minor update” of 3.1. EuroLLM-22B lists all 24 official EU languages among 35, under Apache 2.0, with a sequence length of 32,768 tokens; its card states that it “has not been aligned to human preferences”. For speech, Parakeet TDT 0.6B v3 transcribes 25 European languages, every official EU language except Irish among them, under CC BY 4.0, while Whisper covers more languages under the MIT licence.

Our comparison of open LLMs for European languages sets out which model names which language and how much GPU memory each takes. Test two or three candidates with your own letters and enquiries in every language your staff use, graded by native speakers, before you fix the server size.

Check the licence as closely as the language list. Apache 2.0, MIT and CC BY 4.0 allow use by a public body, while other open-weight models come with acceptable use policies, revenue thresholds or non-commercial terms; our comparison of open LLM licences for company use quotes the clauses. Which terms an authority may accept is a legal assessment for its legal department.

Sizing a GPU server by authority size

A server is sized by the peak number of requests in flight, not by headcount. Our guide to private ChatGPT server sizing by company size explains the method. Its example values are 40 per cent of staff active in the busiest hour, six requests each, 30 seconds per request and a peak factor of 2, which put the peak at about 4 per cent of staff.

Mistral Small 3.2 serves as the example model, since public documents are long and its config.json allows 131,072 tokens, the 128K window the 3.1 card names. Its BF16 weights take 44.7 GiB, and by our sizing rule 38.3 GiB remain for the KV cache on a 96 GB RTX PRO 6000 and 78.6 GiB on a 141 GB H200 NVL. From its config.json (40 layers, 8 KV heads of dimension 128), a token takes 160 KiB of 16-bit cache, so one conversation at a declared 32K context takes 2.5 GiB in an FP8 cache. One RTX PRO 6000 then holds about 15 such conversations and one H200 NVL about 31.

AUTHORITY, STAFFPEAK IN FLIGHTCONFIGURATION32K CONVERSATIONS
Municipality, 30012one server, 2 × RTX PRO 6000 Server Edition: one runs the LLM, one in MIG for search and speech15
Regional office, 1,00040two servers, each 4 × RTX PRO 6000 Server Edition: three LLM copies, one card in MIG45 per server, 90 together
City or ministry, 2,00080two servers, each 4 × H200 NVL: three copies, one card in MIG; or each 8 × RTX PRO 6000: six copies, two for services93 or 90 per server

Our estimates: Mistral Small 3.2 in BF16 with an FP8 KV cache at a declared 32K context, one copy per card, 0.9 × the memory the driver reports less 3 GiB and the weights; peaks from the example values above. Layers and KV heads from the model’s config.json on Hugging Face, read on 10 October 2026.

For 1,000 and 2,000 staff, each of the two servers holds the whole example peak, so the service continues when one fails or is in a maintenance window. The municipality’s single server has no such reserve, and a second server of the same build adds it. With EuroLLM-22B in place of Mistral Small 3.2, one RTX PRO 6000 holds about 12 conversations at its full 32K context with an FP8 cache, and the copy counts rise accordingly. The model card’s own vLLM command spreads Mistral Small 3.2 over two GPUs with --tensor-parallel-size 2, while our count keeps one copy per 96 GB card; check the KV cache size vLLM reports at startup on the card you will use. Eight cards at up to 600 W each draw up to 4.8 kW before processors and fans, so check the rack feed and cooling for the eight-card builds before ordering.

We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users, and check the rack, power and airflow before we quote. Send us your staff numbers, workloads and languages through the form below.

Own server room or EU data centre

For an authority, data sovereignty depends on where the hardware stands, who holds administrative access and which contract governs the hosting. A server in the authority’s own server room settles these questions internally, but needs power, cooling and staff to run it. Dedicated hardware in an EU data centre keeps the data in the EU on servers no other customer uses, and the authority decides which records may go there, with its data protection officer involved.

Our Private AI/ML service runs models on your servers or on dedicated hardware in a Tier-3 data centre in Lithuania, with queries and answers logged. Tell us which records the assistant may reach and where they have to stay.

AI Act duties for public authorities: Article 27 and Annex III

A public body that uses an AI system under its authority is a deployer under Article 3(4) of the AI Act. Drafting, translation, transcription and records search are not among the purposes listed in Annex III, by our reading. Annex III point 5(a) does list systems intended to be used “by public authorities or on behalf of public authorities” to evaluate the eligibility of natural persons for essential public assistance benefits and services, including healthcare services, or to grant, reduce, revoke or reclaim them. Point 5(d) lists systems that evaluate and classify emergency calls by natural persons or dispatch emergency first response services. An enquiry router that starts evaluating eligibility for benefits, or a triage tool that classifies emergency calls, can therefore move into Annex III.

For such a high-risk use, Article 27(1) requires deployers “that are bodies governed by public law, or are private entities providing public services” to perform a fundamental rights impact assessment before deployment; systems for critical infrastructure under Annex III point 2 are excepted. Paragraph 2 applies the duty to the first use and requires an update when an element changes. The assessment covers the processes in which the system is used, the period and frequency of use, the categories of persons affected, the specific risks of harm, the human oversight measures and the steps if risks materialise, including “complaint mechanisms”. Under paragraph 3, the deployer notifies the market surveillance authority of the results and submits the filled-out template of paragraph 5 with the notification. Regulation (EU) 2026/1744 replaced paragraphs 4 and 5, so the deployer may now cross-reference the data protection impact assessment under the GDPR or include parts of it, and the AI Office develops the template. Article 26(8) adds that deployers of high-risk systems that are public authorities comply with the registration obligations of Article 49, and that they do not use a system they find unregistered in the EU database of Article 71 but inform its provider or distributor.

Article 27 sits in Chapter III, Section 3, which Article 113, as amended by Regulation (EU) 2026/1744, applies to Annex III systems from 2 December 2027. Article 50(4) has applied since 2 August 2026. A deployer discloses that text it publishes to inform the public on matters of public interest has been artificially generated or manipulated, unless the text has undergone human review or editorial control and someone holds editorial responsibility for its publication. Our guide to EU AI Act deployer obligations covers the timeline and Article 26. Whether a specific use falls under Annex III is a legal assessment for the authority’s legal department; the server can log prompts, answers and users so that the record exists when it is needed.

General information on EU law as of October 2026, from Regulation (EU) 2024/1689, Articles 3(4), 26(8), 27, 50(4) and 113 and Annex III, in the consolidated version as at 27 July 2026 on eur-lex.europa.eu, and Regulation (EU) 2026/1744, Article 1, point (13).

From pilot to tender and operation

Most authorities buy servers through a public tender, so the pilot also produces the technical specification. Our guide to the GPU server tender specification shows how to write measurable requirements instead of brand names.

  1. Choose one process with clear metrics, such as summarising case files or routing enquiries, and the languages it needs.
  2. Test two or three models with your own documents, and record the busy-hour requests and their duration from the gateway logs.
  3. Size the cards from the measured peak, the declared context and the models that stay loaded, including search and speech.
  4. Write the tender from these figures: memory per card and in total, MIG support, power per rack position, warranty and acceptance tests.
  5. Plan operation: who patches the server, where the logs are kept and for how long, and how a second server takes over.

What we supply

We build AI servers to order for this kind of platform, with the RTX PRO 6000 Server Edition, the H200 NVL with NVLink bridges, the L40S and the L4, assembled and burn-in tested, with manufacturer warranty on every component and delivery anywhere in the EU on one EU contract and invoice. The smaller professional NVIDIA GPUs we supply, such as the RTX PRO 2000 and 4000, suit transcription or search nodes at branch offices. NVIDIA AI Enterprise and vGPU licences come on the same invoice, and operating system, drivers, CUDA and a container runtime are installed on request. The models, records search and logging on top are our Private AI/ML service, with engineering by our partner Vixen.UNO and support under an agreed SLA.

FAQ

How can the public sector run AI on premise?
A public body runs open-weight models such as Mistral Small 3.2 or EuroLLM on its own GPU server, or on dedicated hardware in an EU data centre, with a serving engine such as vLLM. The drafting model takes whole cards such as the RTX PRO 6000 Server Edition or the H200 NVL, while speech recognition and search models share a card through MIG. Prompts, answers and documents then stay on hardware the authority controls.
Which LLM suits public administration?
Public administration needs a multilingual open-weight model under a permissive licence, tested with the authority’s own documents in every language it uses. Mistral Small 3.2 (24B, Apache 2.0) is a minor update of 3.1, whose card names 25 languages including German, Polish and Romanian, while EuroLLM-22B (Apache 2.0) covers all 24 official EU languages with a 32,768-token context. Native speakers should grade the answers of two or three candidates before the server is sized.
What does a sovereign AI server for government mean?
It means that the models, the documents and the logs run on hardware whose location, administrative access and hosting contract the authority controls, either in its own server room or on dedicated servers in an EU data centre. The model weights stay on the server, and no prompt leaves it unless someone enables an external service. The authority decides, with its data protection officer involved, which records may go to which location.
How big a GPU server does a municipality need?
A municipality with about 300 staff peaks at around 12 requests in flight at example usage values, which one RTX PRO 6000 Server Edition serves with Mistral Small 3.2 at a 32K context by our estimate. A second card in MIG mode carries the search and speech recognition models. A second server of the same build adds a reserve if one fails.
Can a public authority use a private LLM instead of a cloud service?
Yes, open-weight models under Apache 2.0, MIT or CC BY 4.0 licences can be deployed on the authority’s own servers, and the terms of other licences are a legal assessment for its legal department. The authority then sizes the GPUs itself, by peak requests in flight, context length and the models it keeps loaded. For 1,000 staff, two servers with four RTX PRO 6000 each hold the example peak of 40 requests even with one server down.
What does the AI Act require of public authorities?
Every public body that uses an AI system under its authority is a deployer (Article 3(4)) and, since 2 August 2026, discloses AI-generated text published to inform the public on matters of public interest, unless the text has undergone human review or editorial control and someone holds editorial responsibility for it (Article 50(4)). For high-risk uses listed in Annex III, such as evaluating eligibility for public assistance benefits under point 5(a), Article 27 requires a fundamental rights impact assessment before first use and Article 26(8) compliance with the registration obligations of Article 49. Under Article 113 as amended by Regulation (EU) 2026/1744, these high-risk duties apply from 2 December 2027.

Send us the number of staff, the workloads you plan (drafting, enquiries, translation, transcription, records search), the languages and where the server will run. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna