BLOG · GUIDE ·

Confidential computing on the GPU: H200 NVL and RTX PRO 6000 Server Edition requirements

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • NVIDIA’s R595 TRD1 release notes (April 2026) list the H200 NVL and the RTX PRO 6000 Blackwell Server Edition for confidential computing in single-GPU passthrough only, with driver 595.58.03 and CUDA 13.2
  • The host needs Intel TDX (Emerald Rapids, Granite Rapids) on Ubuntu 25.10 with kernel 6.17 or AMD SEV-SNP (Milan, Genoa, Turin) on Ubuntu 25.04 with kernel 6.14, with KVM/QEMU and an Ubuntu 24.04 guest
  • The GPU accepts no work until the confidential VM sets its ready state, normally after attestation with NVIDIA’s local verifier or the NRAS cloud service
  • In confidential mode NVIDIA’s Secure AI does not support MIG, MPS, GPUDirect RDMA or CUDA forward compatibility, so each confidential VM gets a whole card, and in NVIDIA’s Kubernetes stack all GPUs on a host go to one confidential VM
  • An arXiv preprint of 2024 that tested an H100 NVL and an H200 NVL with vLLM v0.5.4 reports overhead below 7 per cent for most typical LLM queries, mainly from encrypted PCIe transfers

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Confidential computing on the H200 NVL and RTX PRO 6000

NVIDIA confidential computing on a GPU works on the H200 NVL and the RTX PRO 6000 Server Edition, both in single-GPU passthrough mode only. NVIDIA’s Trusted Computing Solutions release notes for R595 TRD1 (RN-12817-001_v02, April 2026) list “NVIDIA H200NVL” and the “RTX PRO 6000 Blackwell Server Edition” under “Single GPU Passthrough (SPT CC)”, paired with data-centre driver 595.58.03. Neither card appears under a multi-GPU mode. The card works inside a confidential virtual machine (CVM) on a host whose CPU provides AMD SEV-SNP or Intel TDX, and it takes work only after the guest sets its ready state, normally after verifying the card by attestation.

Confidential computing protects code, data and model weights while they are in use. NVIDIA’s whitepaper on Secure AI with Blackwell and Hopper GPUs (August 2025) states the aim as to “protect all application code and data in the VM instance from being read by the host”, which covers the host operating system, the hypervisor and the people who administer them. General hardening of the server is covered in our guide to securing a GPU server.

NVIDIA’s confidential containers page, updated 22 September 2026, lists “NVIDIA H200” for single-GPU passthrough without naming the edition. The release notes name the NVL card and list the HGX H200 8-GPU board as a separate line. Of the RTX PRO 6000 family, the release notes list only the Server Edition, air-cooled and liquid-cooled; the Workstation and Max-Q editions are not listed.

Requirements: CPU TEE, firmware, host and guest

NVIDIA’s Confidential Computing Deployment Guide (DU-12302-001_v7.1, April 2026) asks for Intel processors with TDX, naming Emerald Rapids and Granite Rapids, or AMD processors with SEV-SNP, naming Milan 7xx3, Genoa 9xx4 and Turin. Intel describes TDX as hardware-isolated virtual machines called trust domains, and AMD says SEV-SNP “adds strong memory integrity protection to help prevent malicious hypervisor-based attacks”. The CPU TEE protects the VM’s memory, and the GPU protects the workload on its side and encrypts the transfers between CPU and GPU.

COMPONENTREQUIREMENTSOURCE
GPUH200 NVL or RTX PRO 6000 Server Edition, single-GPU passthroughR595 TRD1 release notes
CPU and TEEIntel TDX or AMD SEV-SNP; SEV-SNP firmware later than 1.51:1; TDX module 1.x on Emerald Rapids, 2.x on Granite RapidsDeployment guide v7.1
BIOSSEV-SNP, IOMMU and SNP memory coverage on; or TDX on, memory integrity and the 46-bit CPU PA limit offDeployment guide v7.1
HostKVM/QEMU; Ubuntu 25.10 with kernel 6.17.0+ (Intel), Ubuntu 25.04 with 6.14+ (AMD)R595 TRD1 release notes
Guest VMUbuntu 24.04, open kernel driver 595.58.03, CUDA 13.2, LKCA enabled, persistence mode onRelease notes, deployment guide
GPU modeset per card with nvidia_gpu_tools.py and --set-cc-mode=on, host Secure Boot off while switchingDeployment guide v7.1
AttestationLocal GPU Verifier 2.6.0 or later, or NVIDIA’s NRAS serviceRelease notes, NVIDIA attestation docs
KubernetesConfidential Containers: Ubuntu 25.10 or 26.04 host, kernel 6.17+, all GPUs on the host in CC mode and in one CVMConfidential containers page

NVIDIA Trusted Computing Solutions release notes R595 TRD1 (April 2026), Confidential Computing Deployment Guide v7.1 (April 2026), Confidential Containers supported platforms (22 September 2026) and attestation documentation (6 October 2026), all read on 10 October 2026.

The guide states that “Setting the CC modes for the GPUs is not possible if the host is configured in Secure Boot mode”, so the mode is set during commissioning and checked afterwards with nvidia-smi conf-compute -f, which reports “CC status: ON”. It also warns that “a reboot command terminates the VM” because of CPU confidential computing limits, which matters for maintenance runbooks. The confidential containers page lists AMD Genoa and Milan and Intel Emerald Rapids and Granite Rapids for its Kubernetes stack, so Turin hosts are documented for VMs but not, on that page, for containers.

Attestation before the GPU accepts work

The deployment guide defines attestation as “the process of challenging the GPU where measurements are collected and signed by the GPU, and these measurements are compared to known-good, golden measurements.” Until that has happened, “The GPU will not accept any work until an enlightened CVM user sets the ReadyState.” A successful attestation run as root sets the ready state, and so does nvidia-smi conf-compute -srs 1, which sets it without an attestation run. A runbook that sets the ready state by command therefore skips the check.

NVIDIA offers two ways to verify. The local path runs inside the CVM, and the release notes require the Local GPU Verifier in version 2.6.0 or later for single-GPU mode. NVIDIA’s C++ attestation SDK, which it calls “the successor of the Python-based guest tools in nvTrust”, verifies locally with nvattest attest and --verifier local. The remote path sends the evidence to NRAS, which NVIDIA describes as “a cloud-based service that verifies the integrity and authenticity of NVIDIA GPUs and platforms” and which checks it against reference values from NVIDIA’s RIM service. NVIDIA groups RIM, NRAS and OCSP under its cloud services, and the pages we read do not say whether verification works on a host without internet access. Offline sites should settle this first.

Attestation can also gate the release of a secret. Our guide to securing a GPU server describes NVIDIA’s reference architecture of June 2026, in which model keys are released only against valid attestation evidence. For Kubernetes, NVIDIA’s confidential containers documentation sets up a development Trustee as the attestation service “for evaluation only” and refers to the upstream Confidential Containers documentation for production.

Single-GPU mode: sizing for 500 to 2,000 users

Single-GPU passthrough gives each confidential VM one whole card, so a model and its KV cache must fit 141 GB on an H200 NVL or 96 GB on an RTX PRO 6000. The release notes say that in this mode “one GPU can be passed through for each Confidential VM (CVM)”. The confidential containers page adds that all GPUs on the host must be in confidential mode and “assigned to one Confidential Container virtual machine”, so by our reading a Kubernetes node in that stack runs one of these cards. For confidential VMs outside Kubernetes, the documents we read do not say whether several single-GPU CVMs may share one host, so this guide counts one card per host.

That suits a common layout for private assistants, one model copy per card behind a load balancer. Our guide for GPU servers in banks and insurers takes 2,000 staff with example values to about 80 requests in flight at the peak. By its estimate, gpt-oss-120b at 32K with a 16-bit cache holds about 19 conversations on one RTX PRO 6000 and about 55 on one H200 NVL.

  1. Five hosts with one RTX PRO 6000 Server Edition each, every card in one CVM with a copy of gpt-oss-120b, hold about 95 conversations, above the peak of 80. A sixth host keeps 95 while one is down for a driver update or a fault.
  2. Two hosts with one H200 NVL each hold about 110, and a third keeps 110 with one host down.

Models that need several cards, such as DeepSeek-V3.2 on eight H200 NVL, have no documented confidential mode on these cards, and the release notes list no confidential mode that uses the NVLink bridges of the H200 NVL. The notes’ multi-GPU modes cover HGX baseboards only, B200 and B300 boards, and Hopper 8-GPU boards in Protected PCIe, where “GPU-GPU communications over the NVLink or NVSwitch interconnect are not encrypted”.

We build AI servers to order with H200 NVL or RTX PRO 6000 Server Edition cards and AMD EPYC or Intel Xeon processors. Tell us which models must run in confidential VMs and how many requests each serves at the peak.

Features that stop working in confidential mode

NVIDIA’s Secure AI Operations Guide (DU-12609-001_v01, November 2025) states that “NVIDIA Secure AI does not support the following features”, and lists CUDA minor version compatibility, CUDA forward compatibility, GPUDirect RDMA, CBL, Multi-Process Service (MPS) and Multi-Instance GPU (MIG). The H200 NVL can otherwise be split into up to seven MIG instances and the RTX PRO 6000 into up to four, but a card in confidential mode serves one VM as a whole. Embedding, reranking and speech models that would share a MIG-partitioned card go on servers outside the confidential set.

The deployment guide sets the mode per card, while the confidential containers page states that “Configuring only some GPUs on a node for Confidential Computing is not supported.” In NVIDIA’s Kubernetes stack, confidential workloads therefore go on servers of their own. Without forward compatibility, the CUDA version in each container must suit the guest driver, and our guide to NVIDIA driver branches and CUDA versions explains how branches and CUDA releases pair. The TRD release fixes a validated driver, 595.58.03, while the production branch has moved on to later builds such as 595.91.07, so check NVIDIA’s Secure AI compatibility matrix, which lists VBIOS, CUDA driver and confidential computing mode per card, before updating a confidential host.

What confidential computing protects and what it does not

THREATIN CC MODESOURCE
Host admin reads VM memoryprotectedSecure AI whitepaper
PCIe data, CPU to GPUencrypted, 256-bit AES-GCM bounce buffersSecure AI whitepaper
Modified GPU firmwarereported in signed measurements, checked if attestation gates the ready stateWhitepaper, deployment guide
GPU performance countersblocked in CC-On, open in CC-DevToolsSecure AI whitepaper
Physical attack on hardwaresophisticated attacks out of scopeSecure AI whitepaper
Hypervisor stops the VMout of scopeSecure AI whitepaper
Flaw in the applicationnot addressed: code inside the CVM sees the dataour reading of the whitepaper
Hopper NVLink in PPCIe modenot encryptedR595 TRD1 release notes

NVIDIA Secure AI with Blackwell and Hopper GPUs whitepaper (WP-12554-001_v1.3, August 2025), Confidential Computing Deployment Guide v7.1 and R595 TRD1 release notes, read on 10 October 2026.

The whitepaper excludes “Sophisticated physical attacks” and “Denial of service attacks”, and names as the only availability attack out of scope “one where a malicious hypervisor prevents access to a CVM.” The host can therefore stop a confidential workload but cannot read it. A user with valid access, a prompt injection or a flaw in the serving code works inside the trusted VM, so access control and gateway logging stay as necessary as on any server. The DevTools mode unblocks performance counters “for developer use”, so production hosts run in CC-On.

Performance effects in published tests

The NVIDIA documents we read give no overhead figures. A preprint on arXiv (2409.03992, submitted on 6 September 2024, version of 5 November 2024) tested an H100 NVL on an AMD SEV-SNP host and an H200 NVL on an Intel TDX host with vLLM v0.5.4. It ran Llama 3.1 8B, a 14B model and Llama 3.1 70B in 4-bit, each on one GPU. Its abstract reports that “the overhead remains below 7%” for most typical LLM queries, with “nearly zero overhead” for larger models and longer sequences.

The authors attribute the penalty mainly to “CPU-GPU data transfers via PCIe”, and the paper gives an average overhead below 9 per cent. Those runs used the software of 2024, not the R595 driver or a current vLLM release. NVIDIA’s operations guide adds that OpenSSL 3.5.0 and later use AVX512 for faster encryption and that architectures based on bounce buffers “benefit most”. We found no published figures for the RTX PRO 6000 Server Edition in confidential mode, so plan a test run on the delivered server before fixing user numbers.

Banks, health data and hosted servers

Confidential computing addresses whether those who run the host can read the workload. That question arises for a bank whose encryption policy under DORA covers “the encryption of data in use, where necessary” (Delegated Regulation (EU) 2024/1774, Article 6(2)(b)), a hospital that runs models on patient records and any company whose GPU servers stand in a data centre administered by others. Our guide on running a private LLM in an EU data centre compares colocation and dedicated hosting, and our article on the CLOUD Act and EU data residency explains that without it a VM on a provider’s hosts is decrypted by those hosts while it runs. Whether a given regulation requires confidential computing is a legal assessment for the company’s legal department.

We supply both cards in servers built to order, with manufacturer warranty, on one EU contract and invoice. Describe in the form below where the servers will stand and who administers the hosts.

What we supply

We build AI servers to order with the H200 NVL or the RTX PRO 6000 Blackwell Server Edition, the two cards in our range that NVIDIA lists for confidential computing, on AMD EPYC or Intel Xeon platforms, sized per workload. Each server is assembled and burn-in tested, with manufacturer warranty on every component and delivery anywhere in the EU. We check the rack, power and airflow before we quote, and operating system, drivers, CUDA and a container runtime are installed on request. The full card range, including the L40S and L4 for workloads outside the confidential set, is on our professional GPUs page.

FAQ

Which NVIDIA GPUs support confidential computing?
NVIDIA’s R595 TRD1 release notes of April 2026 list, among others, the H100 PCIe, H100 NVL and H200 NVL cards, HGX boards with H100, H200, H20, B200 and B300, and the RTX PRO 6000 Blackwell Server Edition. The H200 NVL and the RTX PRO 6000 Server Edition are listed for single-GPU passthrough only. Multi-GPU modes are listed for HGX boards.
Does the H200 support confidential computing?
Yes. The release notes list “NVIDIA H200NVL” under single-GPU passthrough, and the H200 NVL product page shows confidential computing as supported. NVIDIA’s confidential containers page lists “NVIDIA H200” for single-GPU passthrough without naming the edition, and the HGX H200 8-GPU board is listed separately, for single-GPU passthrough and for Protected PCIe.
Does the RTX PRO 6000 support confidential computing?
The RTX PRO 6000 Blackwell Server Edition does, in its air-cooled and liquid-cooled versions, in single-GPU passthrough with driver 595.58.03 per NVIDIA’s R595 TRD1 release notes. The Workstation and Max-Q editions are not listed. NVIDIA’s Secure AI operations guide of November 2025 lists MIG, which otherwise splits the card into up to four instances, among the features not supported in confidential mode.
What does NVIDIA confidential computing need on the host?
It needs a CPU with Intel TDX or AMD SEV-SNP, the matching BIOS settings, KVM/QEMU on Ubuntu 25.10 with kernel 6.17 for Intel or Ubuntu 25.04 with kernel 6.14 for AMD, and an Ubuntu 24.04 guest with the 595.58.03 open driver. The GPU is switched into confidential mode per card with NVIDIA’s GPU admin tools while host Secure Boot is off.
What is a GPU TEE and how is it attested?
A GPU TEE extends the confidential VM of the CPU to the GPU, so the host cannot read the workload and traffic over PCIe is encrypted. The GPU signs measurements of its firmware and configuration, which the local verifier or NVIDIA’s NRAS cloud service compares with reference values. The GPU accepts no work until the confidential VM sets its ready state, normally after a successful attestation.
How much does confidential AI slow down LLM inference?
The NVIDIA documents we read give no overhead figures. An arXiv preprint of 2024, tested on an H100 NVL and an H200 NVL with vLLM v0.5.4, reports overhead below 7 per cent for most typical queries and an average below 9 per cent, mainly from encrypted CPU to GPU transfers, with almost none for large models and long sequences. We found no published figures for the RTX PRO 6000 Server Edition or for the R595 driver.

Send us the models that must run in confidential VMs, their peak requests in flight, the host CPU platform you prefer and whether the servers stand on-premise or in a data centre. We reply within one business day with a configuration and quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna