BLOG · GUIDE ·

AI infrastructure for banks and insurers: GPU servers for on-premise LLMs, isolation and DORA

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Banks and insurers run internal assistants with RAG, claims, KYC and AML document review, call transcription and code assistants on premise; each is a different model type, and the card follows from the model and the peak number of requests in flight
  • With the example values of our company-size guide, 2,000 staff send about 80 requests in flight at the peak; each of two servers with eight RTX PRO 6000 Server Edition can carry gpt-oss-120b for that peak plus a coding model, a document model and four 24 GB MIG instances, by our estimate
  • DORA has applied since 17 January 2025 to credit institutions and insurance undertakings among others, and its RTS 2024/1774 asks for encryption of data in use where necessary, network segmentation, least-privilege access and logs protected against tampering
  • AI Act Annex III point 5 lists creditworthiness of natural persons and life and health insurance pricing as high-risk uses; under the amended Article 113 those rules apply from 2 December 2027, and financial institutions keep the logs within their financial-services documentation
  • MIG splits an H200 NVL into up to 7 instances and an RTX PRO 6000 into up to 4, with their own memory and cache; NVIDIA lists confidential computing for the H200 NVL, and its confidential containers page lists single-GPU passthrough for the RTX PRO 6000 Server Edition and the H200, with every GPU on the host assigned to one confidential VM

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

AI infrastructure for banks and insurers: what runs on the GPUs

On-premise LLMs in a bank or insurer can serve five kinds of work: an internal assistant with RAG on policies and procedures, review of claims, KYC and AML documents, transcription and summaries of customer calls, code assistants for their developers and batch summaries of archives. Each is a different model type with its own memory profile, so AI infrastructure for banks and insurers is planned from the list of workloads and the peak number of requests in flight, not from headcount. The models run on servers the institution controls, in its own data centre or a hosted rack.

Sizing per model is covered in our guide to sizing a private ChatGPT server by company size, which converts headcount into requests in flight. This article adds the points that matter in the financial sector: the EU rules that frame the infrastructure, isolation between business units, confidential computing, offline installation and logs.

Workloads, model types and GPU cards

The table maps the common workloads to model types and to the cards we supply. Figures come from the vendors’ model pages and from the estimates in our sizing guides, which state their method.

WORKLOADMODEL TYPECARD AND LAYOUT
Internal assistant and RAGgeneral LLM such as gpt-oss-120b, plus embedding and reranker modelsone RTX PRO 6000 or H200 NVL per copy of the LLM; retrieval models on a 24 GB MIG instance or an L4
Claims, KYC and AML reviewOCR models of 0.9B to 8.3B, or a vision-language model such as Qwen3.8-27Bsmall OCR models on an L4 or RTX PRO 4000; a 27B model on an RTX PRO 6000 or H200 NVL
Call transcriptionWhisper large-v3 or Parakeet, plus an LLM for summariesWhisper large needs about 10 GB, so a MIG instance, an L4 or an L40S
Code assistantcoding model such as Qwen3-Coder-30B-A3Bone RTX PRO 6000 holds its FP8 weights and about 19 sessions of 64K with an FP8 cache

Whisper VRAM from OpenAI’s Whisper repository; OCR and Qwen3.8-27B sizes from our document AI guide; Qwen3-Coder sessions from our coding assistant guide; gpt-oss-120b from our company-size guide; MIG profiles from NVIDIA’s MIG user guide (11 September 2026). All session counts are our estimates.

Document review is usually a batch job with a deadline, while the assistant and the code assistant are interactive and sized for the busy hour. Keeping them on separate cards stops a large batch of claims documents from slowing the assistant during office hours.

A worked example: two GPU servers for 2,000 staff

Take a bank or insurer with 2,000 employees and the example values of our company-size guide: 40 per cent of staff use the assistant in the busiest hour, each sends 6 requests an hour, a request is in flight for 30 seconds and the peak is twice the hourly average. That gives about 80 requests in flight. With gpt-oss-120b at a declared 32K context and a 16-bit KV cache, one RTX PRO 6000 holds about 19 conversations by our estimate, so five cards with one copy each hold 95.

Each of two servers with eight RTX PRO 6000 Server Edition can then carry every workload of the service on its own:

  1. Cards 1 to 5: one copy of gpt-oss-120b each, 95 conversations at 32K, above the peak of 80.
  2. Card 6: Qwen3-Coder-30B-A3B in FP8 for the code assistant, about 19 sessions of 64K with an FP8 cache.
  3. Card 7: a vision-language model such as Qwen3.8-27B (30.9 GB in FP8) for claims and KYC documents.
  4. Card 8: split by MIG into four 24 GB instances for the embedding model, the reranker, Whisper and a small OCR model.

Because each server holds the whole peak, the service keeps running when one server is down for a failure or a driver update. With the H200 NVL, the same guide estimates about 55 conversations per card at 32K, so two cards per server carry the assistant and the other workloads need cards of their own. Eight 600 W cards draw 4.8 kW before processors and fans, which calls for three-phase feeds at each rack position.

The two servers can stand in two rooms or at two sites, which spreads the risk of a power or cooling fault. Whether the second site also serves as a recovery site under the institution’s continuity plan is part of its DORA planning, which our guide to DORA backup and resilience testing covers.

We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us your headcount, the workloads and the models on your shortlist through the form below, and we return a configuration and quote within one business day.

DORA and the AI Act: the EU frame for AI infrastructure

Regulation (EU) 2022/2554, the Digital Operational Resilience Act (DORA), lists credit institutions and insurance and reinsurance undertakings among the entities it covers in Article 2(1), and its Article 64 states: “It shall apply from 17 January 2025.” Article 6(1) requires “a sound, comprehensive and well-documented ICT risk management framework”, and Commission Delegated Regulation (EU) 2024/1774 of 13 March 2024, the RTS on ICT risk management, sets out its policies and tools in Title II, from encryption and logging to access control; entities under DORA’s simplified framework of Article 16(1) follow Title III instead. DORA’s definition of ICT services in Article 3(21) covers “hardware as a service and hardware services”, including technical support via software or firmware updates by the hardware provider, so a support contract for GPU servers can come within it.

The AI Act adds rules for specific uses rather than for infrastructure. Annex III point 5 lists AI systems “intended to be used to evaluate the creditworthiness of natural persons or establish their credit score”, with an exception for detecting financial fraud, and systems for risk assessment and pricing “in relation to natural persons in the case of life and health insurance”. Under Article 113 as amended by Regulation (EU) 2026/1744, the Digital Omnibus on AI, the rules for Annex III systems apply from 2 December 2027. For deployers of such high-risk systems that are financial institutions, Article 26(6) says they “shall maintain the logs as part of the documentation kept pursuant to the relevant Union financial service law”. An internal assistant that drafts and summarises is not among the Annex III purposes, and our guide to the EU AI Act for companies using LLMs covers the deployer duties. Whether a specific system is high-risk, and whether a support contract is an ICT service for a critical or important function in the sense of DORA Article 3(22), is a legal assessment for the institution’s legal and compliance departments.

Isolating business units with MIG and separate cards

A bank or insurer may need to keep business units apart whose data must not mix, for example retail lending, asset management and internal audit. On a GPU server there are three levels of separation, and the choice follows from the data and the model size.

The first is the application layer: one model serves several units, and the RAG platform filters documents by each user’s access rights before they reach the prompt. The second is MIG, which NVIDIA describes as a way to partition a GPU “into up to seven separate GPU Instances”. Its user guide states that “L2 cache banks, memory controllers, and DRAM address busses are all assigned uniquely to an individual instance” and that MIG gives “a defined quality of service (QoS) with fault isolation for different clients”. As of 11 September 2026, NVIDIA’s supported-GPU table lists 7 instances for the H200 NVL and 4 for every RTX PRO 6000 edition, so an RTX PRO 6000 gives units up to four 24 GB instances. The third level is a whole card or a whole server per unit, which large models need anyway: gpt-oss-120b’s weights take about 60.8 GiB of a 96 GB card, so MIG suits the small retrieval, speech and OCR models, not the main LLM.

Units that may share a model but not documents work at the first level. Units whose data may not share memory or cache with others get their own MIG instance or card. MIG instances on one card still share the host and its GPU driver, so each unit’s instance also runs in its own virtual machine or container, with access rules recorded in the platform’s documentation.

Confidential computing, offline installation and logs

The RTS asks for “the encryption of data in use, where necessary”, and where that is not possible, for processing “in a separated and protected environment, or take equivalent measures”. On GPUs, data in use is protected by confidential computing, which needs a host CPU with AMD SEV-SNP or Intel TDX and a GPU in confidential computing mode. NVIDIA’s H200 product page lists confidential computing as “Supported” for both the H200 SXM and the H200 NVL. Its confidential containers page, updated 22 September 2026, lists the RTX PRO 6000 Blackwell Server Edition and the “NVIDIA H200”, without naming the edition, for single-GPU passthrough. A note on that page states that “Configuring only some GPUs on a node for Confidential Computing is not supported”, and that all GPUs on the host must be assigned to one confidential container virtual machine. By our reading, a host in that stack runs one confidential virtual machine with one of these cards, so confidential workloads go on servers of their own. A model such as gpt-oss-120b, which runs on one card, fits that layout. The page lists the H200 for multi-GPU passthrough only in Protected PCIe mode, which our guide to confidential computing on the H200 NVL and RTX PRO 6000 ties to HGX 8-GPU boards; it also covers attestation and the supported modes.

Where inference servers run without internet access, drivers, container images and model weights are fetched at fixed versions on a connected staging host, checked against a SHA-256 manifest and installed from internal mirrors, as our guide to installing an air-gapped LLM server offline describes. The recorded hashes then give the asset inventory a verifiable record of each model version.

For logs, the RTS lists events from access control and identity management, capacity management, change management, ICT operations and network traffic, and asks for “measures to protect logging systems and log information against tampering, deletion, and unauthorised access”. For an LLM service the useful record is at the gateway: user identity, model and version, time and, where policy allows, prompts and answers. A gateway that authenticates each user also supports the RTS provision on user accountability, which limits “the use of generic and shared user accounts”.

Our Private AI/ML service, with engineering by our partner Vixen.UNO, includes logging of queries and answers and data and permissions management. Describe which business units will share the platform in the form below, and what each may see.

Control expectations and technical measures

The table pairs expectations from the EU texts with technical measures on a GPU platform. A measure supports a control, and whether the institution meets the text is for its own assessment.

SOURCEEXPECTATIONTECHNICAL MEASURE
RTS Art. 6(2)(b)encryption of data in use, where necessaryconfidential computing on dedicated servers, or a separate card or MIG instance with its own VM
RTS Art. 13(a)segregation and segmentation of systems and networksGPU servers and their BMCs in separate segments; inference API reachable only from the gateway
RTS Art. 21need-to-know, least privilege, identifiable userssingle sign-on at the gateway, no shared API keys, dedicated administrator accounts
RTS Art. 12logging, protected against tampering and deletiongateway records shipped to a central log store with restricted write access
RTS Art. 9capacity and performance managementGPU memory, utilisation and requests in flight per model, compared with the sized peak
AI Act Art. 26(6)logs of high-risk systems, at least six monthsretention per system, kept with the financial-services documentation

Commission Delegated Regulation (EU) 2024/1774 (RTS on ICT risk management, Title II) and Regulation (EU) 2024/1689 as consolidated on 27 July 2026, read in the Official Journal text and on the Commission’s AI Act Service Desk on 10 October 2026. The measures are technical examples, not a statement of compliance.

What we supply

We supply the RTX PRO 6000 Server Edition, the H200 NVL with NVLink bridges, the L40S, the L4 and the smaller RTX PRO cards, as cards or in AI servers built to order with 2 to 8 GPUs per node, burn-in tested, with manufacturer warranty, on one EU contract and invoice. NVIDIA AI Enterprise and vGPU licences come on the same invoice, and the card line-up is on our professional GPUs page. Operating system, drivers, CUDA and a container runtime are installed on request. The platform on top, private LLMs, RAG with access rights, query logging and MLOps, is our Private AI/ML service, with engineering by our partner Vixen.UNO, on your servers or in a Tier-3 data centre in Lithuania.

FAQ

What AI infrastructure does a bank need for on-premise LLMs?
It needs GPU servers sized for the peak number of requests in flight across its workloads: the internal assistant, document review, call transcription and code assistants. A large model such as gpt-oss-120b takes one RTX PRO 6000 or H200 NVL per copy, and small embedding, speech and OCR models fit 24 GB MIG instances or an L4. Two servers that each carry the whole peak keep the service running when one fails.
How many GPUs does an insurance company with 2,000 staff need for an LLM?
With example values of 40 per cent busy-hour users, 6 requests an hour, 30 seconds per request and a peak factor of 2, 2,000 staff produce about 80 requests in flight. By our estimate one RTX PRO 6000 holds about 19 conversations of gpt-oss-120b at 32K, so five cards hold 95, and one H200 NVL holds about 55. A pilot replaces the example values with figures from its own gateway logs.
Does DORA apply to AI infrastructure in banks?
DORA, Regulation (EU) 2022/2554, has applied since 17 January 2025 to credit institutions, insurance undertakings and other financial entities, and its ICT risk management rules, detailed in RTS 2024/1774, are written for ICT systems and assets in general. Requirements such as encryption of data in use where necessary, network segmentation, least-privilege access and protected logs map to GPU servers as to any other system. Entities under DORA’s simplified ICT risk management framework of Article 16(1) follow Title III of the same RTS instead.
Is credit scoring with AI high-risk under the EU AI Act?
Annex III point 5(b) lists AI systems intended to evaluate the creditworthiness of natural persons or establish their credit score as high-risk, with an exception for systems that detect financial fraud. Point 5(c) adds risk assessment and pricing for natural persons in life and health insurance. Under the amended Article 113, these rules apply from 2 December 2027.
How do you isolate business units on one GPU server?
Units that may share a model but not documents are separated at the application layer, where the RAG platform filters documents by access rights. Units whose data may not share GPU memory get their own MIG instance, with dedicated memory controllers and cache, or a whole card. NVIDIA lists up to 7 MIG instances on the H200 NVL and up to 4 on every RTX PRO 6000 edition.
Does confidential computing work on the RTX PRO 6000 and H200?
NVIDIA’s confidential containers page, updated 22 September 2026, lists the RTX PRO 6000 Blackwell Server Edition and the H200 for single-GPU passthrough, with AMD SEV-SNP or Intel TDX on the host, and NVIDIA’s H200 product page lists confidential computing as supported on the H200 NVL. The confidential containers page also requires every GPU on the host in confidential computing mode, all assigned to one confidential virtual machine. By our reading that is one card per host in this stack, so confidential workloads go on servers of their own.

Send us your headcount, the workloads you plan (assistant, document review, call transcription, code), the models on your shortlist and the sites the servers will stand in. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna