AI infrastructure for banks and insurers: GPU servers for on-premise LLMs, isolation and DORA
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Banks and insurers run internal assistants with RAG, claims, KYC and AML document review, call transcription and code assistants on premise; each is a different model type, and the card follows from the model and the peak number of requests in flight
- With the example values of our company-size guide, 2,000 staff send about 80 requests in flight at the peak; each of two servers with eight RTX PRO 6000 Server Edition can carry gpt-oss-120b for that peak plus a coding model, a document model and four 24 GB MIG instances, by our estimate
- DORA has applied since 17 January 2025 to credit institutions and insurance undertakings among others, and its RTS 2024/1774 asks for encryption of data in use where necessary, network segmentation, least-privilege access and logs protected against tampering
- AI Act Annex III point 5 lists creditworthiness of natural persons and life and health insurance pricing as high-risk uses; under the amended Article 113 those rules apply from 2 December 2027, and financial institutions keep the logs within their financial-services documentation
- MIG splits an H200 NVL into up to 7 instances and an RTX PRO 6000 into up to 4, with their own memory and cache; NVIDIA lists confidential computing for the H200 NVL, and its confidential containers page lists single-GPU passthrough for the RTX PRO 6000 Server Edition and the H200, with every GPU on the host assigned to one confidential VM
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
AI infrastructure for banks and insurers: what runs on the GPUs
On-premise LLMs in a bank or insurer can serve five kinds of work: an internal assistant with RAG on policies and procedures, review of claims, KYC and AML documents, transcription and summaries of customer calls, code assistants for their developers and batch summaries of archives. Each is a different model type with its own memory profile, so AI infrastructure for banks and insurers is planned from the list of workloads and the peak number of requests in flight, not from headcount. The models run on servers the institution controls, in its own data centre or a hosted rack.
Sizing per model is covered in our guide to sizing a private ChatGPT server by company size, which converts headcount into requests in flight. This article adds the points that matter in the financial sector: the EU rules that frame the infrastructure, isolation between business units, confidential computing, offline installation and logs.
Workloads, model types and GPU cards
The table maps the common workloads to model types and to the cards we supply. Figures come from the vendors’ model pages and from the estimates in our sizing guides, which state their method.
| WORKLOAD | MODEL TYPE | CARD AND LAYOUT |
|---|---|---|
| Internal assistant and RAG | general LLM such as gpt-oss-120b, plus embedding and reranker models | one RTX PRO 6000 or H200 NVL per copy of the LLM; retrieval models on a 24 GB MIG instance or an L4 |
| Claims, KYC and AML review | OCR models of 0.9B to 8.3B, or a vision-language model such as Qwen3.8-27B | small OCR models on an L4 or RTX PRO 4000; a 27B model on an RTX PRO 6000 or H200 NVL |
| Call transcription | Whisper large-v3 or Parakeet, plus an LLM for summaries | Whisper large needs about 10 GB, so a MIG instance, an L4 or an L40S |
| Code assistant | coding model such as Qwen3-Coder-30B-A3B | one RTX PRO 6000 holds its FP8 weights and about 19 sessions of 64K with an FP8 cache |
Whisper VRAM from OpenAI’s Whisper repository; OCR and Qwen3.8-27B sizes from our document AI guide; Qwen3-Coder sessions from our coding assistant guide; gpt-oss-120b from our company-size guide; MIG profiles from NVIDIA’s MIG user guide (11 September 2026). All session counts are our estimates.
Document review is usually a batch job with a deadline, while the assistant and the code assistant are interactive and sized for the busy hour. Keeping them on separate cards stops a large batch of claims documents from slowing the assistant during office hours.
A worked example: two GPU servers for 2,000 staff
Take a bank or insurer with 2,000 employees and the example values of our company-size guide: 40 per cent of staff use the assistant in the busiest hour, each sends 6 requests an hour, a request is in flight for 30 seconds and the peak is twice the hourly average. That gives about 80 requests in flight. With gpt-oss-120b at a declared 32K context and a 16-bit KV cache, one RTX PRO 6000 holds about 19 conversations by our estimate, so five cards with one copy each hold 95.
Each of two servers with eight RTX PRO 6000 Server Edition can then carry every workload of the service on its own:
- Cards 1 to 5: one copy of gpt-oss-120b each, 95 conversations at 32K, above the peak of 80.
- Card 6: Qwen3-Coder-30B-A3B in FP8 for the code assistant, about 19 sessions of 64K with an FP8 cache.
- Card 7: a vision-language model such as Qwen3.8-27B (30.9 GB in FP8) for claims and KYC documents.
- Card 8: split by MIG into four 24 GB instances for the embedding model, the reranker, Whisper and a small OCR model.
Because each server holds the whole peak, the service keeps running when one server is down for a failure or a driver update. With the H200 NVL, the same guide estimates about 55 conversations per card at 32K, so two cards per server carry the assistant and the other workloads need cards of their own. Eight 600 W cards draw 4.8 kW before processors and fans, which calls for three-phase feeds at each rack position.
The two servers can stand in two rooms or at two sites, which spreads the risk of a power or cooling fault. Whether the second site also serves as a recovery site under the institution’s continuity plan is part of its DORA planning, which our guide to DORA backup and resilience testing covers.
We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us your headcount, the workloads and the models on your shortlist through the form below, and we return a configuration and quote within one business day.
DORA and the AI Act: the EU frame for AI infrastructure
Regulation (EU) 2022/2554, the Digital Operational Resilience Act (DORA), lists credit institutions and insurance and reinsurance undertakings among the entities it covers in Article 2(1), and its Article 64 states: “It shall apply from 17 January 2025.” Article 6(1) requires “a sound, comprehensive and well-documented ICT risk management framework”, and Commission Delegated Regulation (EU) 2024/1774 of 13 March 2024, the RTS on ICT risk management, sets out its policies and tools in Title II, from encryption and logging to access control; entities under DORA’s simplified framework of Article 16(1) follow Title III instead. DORA’s definition of ICT services in Article 3(21) covers “hardware as a service and hardware services”, including technical support via software or firmware updates by the hardware provider, so a support contract for GPU servers can come within it.
The AI Act adds rules for specific uses rather than for infrastructure. Annex III point 5 lists AI systems “intended to be used to evaluate the creditworthiness of natural persons or establish their credit score”, with an exception for detecting financial fraud, and systems for risk assessment and pricing “in relation to natural persons in the case of life and health insurance”. Under Article 113 as amended by Regulation (EU) 2026/1744, the Digital Omnibus on AI, the rules for Annex III systems apply from 2 December 2027. For deployers of such high-risk systems that are financial institutions, Article 26(6) says they “shall maintain the logs as part of the documentation kept pursuant to the relevant Union financial service law”. An internal assistant that drafts and summarises is not among the Annex III purposes, and our guide to the EU AI Act for companies using LLMs covers the deployer duties. Whether a specific system is high-risk, and whether a support contract is an ICT service for a critical or important function in the sense of DORA Article 3(22), is a legal assessment for the institution’s legal and compliance departments.
Isolating business units with MIG and separate cards
A bank or insurer may need to keep business units apart whose data must not mix, for example retail lending, asset management and internal audit. On a GPU server there are three levels of separation, and the choice follows from the data and the model size.
The first is the application layer: one model serves several units, and the RAG platform filters documents by each user’s access rights before they reach the prompt. The second is MIG, which NVIDIA describes as a way to partition a GPU “into up to seven separate GPU Instances”. Its user guide states that “L2 cache banks, memory controllers, and DRAM address busses are all assigned uniquely to an individual instance” and that MIG gives “a defined quality of service (QoS) with fault isolation for different clients”. As of 11 September 2026, NVIDIA’s supported-GPU table lists 7 instances for the H200 NVL and 4 for every RTX PRO 6000 edition, so an RTX PRO 6000 gives units up to four 24 GB instances. The third level is a whole card or a whole server per unit, which large models need anyway: gpt-oss-120b’s weights take about 60.8 GiB of a 96 GB card, so MIG suits the small retrieval, speech and OCR models, not the main LLM.
Units that may share a model but not documents work at the first level. Units whose data may not share memory or cache with others get their own MIG instance or card. MIG instances on one card still share the host and its GPU driver, so each unit’s instance also runs in its own virtual machine or container, with access rules recorded in the platform’s documentation.
Confidential computing, offline installation and logs
The RTS asks for “the encryption of data in use, where necessary”, and where that is not possible, for processing “in a separated and protected environment, or take equivalent measures”. On GPUs, data in use is protected by confidential computing, which needs a host CPU with AMD SEV-SNP or Intel TDX and a GPU in confidential computing mode. NVIDIA’s H200 product page lists confidential computing as “Supported” for both the H200 SXM and the H200 NVL. Its confidential containers page, updated 22 September 2026, lists the RTX PRO 6000 Blackwell Server Edition and the “NVIDIA H200”, without naming the edition, for single-GPU passthrough. A note on that page states that “Configuring only some GPUs on a node for Confidential Computing is not supported”, and that all GPUs on the host must be assigned to one confidential container virtual machine. By our reading, a host in that stack runs one confidential virtual machine with one of these cards, so confidential workloads go on servers of their own. A model such as gpt-oss-120b, which runs on one card, fits that layout. The page lists the H200 for multi-GPU passthrough only in Protected PCIe mode, which our guide to confidential computing on the H200 NVL and RTX PRO 6000 ties to HGX 8-GPU boards; it also covers attestation and the supported modes.
Where inference servers run without internet access, drivers, container images and model weights are fetched at fixed versions on a connected staging host, checked against a SHA-256 manifest and installed from internal mirrors, as our guide to installing an air-gapped LLM server offline describes. The recorded hashes then give the asset inventory a verifiable record of each model version.
For logs, the RTS lists events from access control and identity management, capacity management, change management, ICT operations and network traffic, and asks for “measures to protect logging systems and log information against tampering, deletion, and unauthorised access”. For an LLM service the useful record is at the gateway: user identity, model and version, time and, where policy allows, prompts and answers. A gateway that authenticates each user also supports the RTS provision on user accountability, which limits “the use of generic and shared user accounts”.
Our Private AI/ML service, with engineering by our partner Vixen.UNO, includes logging of queries and answers and data and permissions management. Describe which business units will share the platform in the form below, and what each may see.
Control expectations and technical measures
The table pairs expectations from the EU texts with technical measures on a GPU platform. A measure supports a control, and whether the institution meets the text is for its own assessment.
| SOURCE | EXPECTATION | TECHNICAL MEASURE |
|---|---|---|
| RTS Art. 6(2)(b) | encryption of data in use, where necessary | confidential computing on dedicated servers, or a separate card or MIG instance with its own VM |
| RTS Art. 13(a) | segregation and segmentation of systems and networks | GPU servers and their BMCs in separate segments; inference API reachable only from the gateway |
| RTS Art. 21 | need-to-know, least privilege, identifiable users | single sign-on at the gateway, no shared API keys, dedicated administrator accounts |
| RTS Art. 12 | logging, protected against tampering and deletion | gateway records shipped to a central log store with restricted write access |
| RTS Art. 9 | capacity and performance management | GPU memory, utilisation and requests in flight per model, compared with the sized peak |
| AI Act Art. 26(6) | logs of high-risk systems, at least six months | retention per system, kept with the financial-services documentation |
Commission Delegated Regulation (EU) 2024/1774 (RTS on ICT risk management, Title II) and Regulation (EU) 2024/1689 as consolidated on 27 July 2026, read in the Official Journal text and on the Commission’s AI Act Service Desk on 10 October 2026. The measures are technical examples, not a statement of compliance.
What we supply
We supply the RTX PRO 6000 Server Edition, the H200 NVL with NVLink bridges, the L40S, the L4 and the smaller RTX PRO cards, as cards or in AI servers built to order with 2 to 8 GPUs per node, burn-in tested, with manufacturer warranty, on one EU contract and invoice. NVIDIA AI Enterprise and vGPU licences come on the same invoice, and the card line-up is on our professional GPUs page. Operating system, drivers, CUDA and a container runtime are installed on request. The platform on top, private LLMs, RAG with access rights, query logging and MLOps, is our Private AI/ML service, with engineering by our partner Vixen.UNO, on your servers or in a Tier-3 data centre in Lithuania.
FAQ
What AI infrastructure does a bank need for on-premise LLMs?
How many GPUs does an insurance company with 2,000 staff need for an LLM?
Does DORA apply to AI infrastructure in banks?
Is credit scoring with AI high-risk under the EU AI Act?
How do you isolate business units on one GPU server?
Does confidential computing work on the RTX PRO 6000 and H200?
Send us your headcount, the workloads you plan (assistant, document review, call transcription, code), the models on your shortlist and the sites the servers will stand in. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day