BLOG · COMPARISON ·

H200 NVL and RTX PRO 6000 in one platform: which workloads go on which servers

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • In a platform of two to four GPU servers, the H200 NVL takes models above 96 GB, long contexts with many users, fine-tuning and FP64 work; the RTX PRO 6000 Server Edition takes mid-size models, NVFP4 checkpoints, embeddings in MIG slices and virtual desktops
  • The H200 NVL has 141 GB of HBM3e at 4.8 TB/s, NVLink bridges for 2 or 4 cards and 30 TFLOPS of FP64; the RTX PRO 6000 Server Edition has 96 GB of GDDR7, FP4 Tensor Cores, RT cores, video encoders and MIG with graphics
  • By our estimate, Qwen3-235B-A22B-Instruct-2507 in FP8 (236.4 GB) with an FP8 cache holds about 93 conversations of 32K on four bridged H200 NVL, 38 on four RTX PRO 6000 and 9 on two H200 NVL
  • Both cards advertise the same Kubernetes resource, nvidia.com/gpu, so workloads are pinned to a pool by GPU Feature Discovery labels, node affinity and a taint on the H200 NVL nodes that the GPU Operator’s own pods must tolerate
  • NVIDIA’s GPU Operator lists both cards with driver 595.91.07 as default; each H200 NVL includes a five-year NVIDIA AI Enterprise subscription, and of the two only the RTX PRO 6000 Server Edition is on NVIDIA’s vGPU list for virtual desktops

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Which workloads go on the H200 NVL and which on the RTX PRO 6000

In a mixed GPU cluster of two to four servers, put the H200 NVL where one job needs a large amount of fast memory in one place: models larger than one 96 GB card, long contexts with many concurrent users, fine-tuning and double-precision computing. Put the RTX PRO 6000 Blackwell Server Edition where the platform runs many mid-size models at once: chat and coding models that fit one card, vision-language models, embedding and reranking services, NVFP4 checkpoints, virtual desktops and rendering.

Both come as dual-slot, air-cooled PCIe Gen5 cards rated up to 600 W, configurable, according to NVIDIA’s product pages; the RTX PRO 6000 Server Edition is also listed as a single-slot liquid-cooled card. Power and cooling are therefore planned the same way per card in both pools, and the two pools can share one driver version, one Kubernetes cluster and one operating procedure. Most of the planning work lies in keeping each workload on the pool it was sized for.

The card-by-card specifications and published benchmarks are in our comparison of the RTX PRO 6000 and H200 NVL for LLM inference. This article covers how the two divide the work when a company runs both.

The card properties that decide placement

Memory and bandwidth come first. The H200 NVL carries 141 GB of HBM3e at 4.8 TB/s, and the RTX PRO 6000 Server Edition 96 GB of GDDR7 at 1,597 GB/s. The H200 NVL also joins 2 or 4 cards with NVLink bridges at 900 GB/s per GPU, against 128 GB/s for PCIe Gen5, which counts once a model is split across cards. NVIDIA rates it at 30 TFLOPS of FP64 for simulation and other double-precision code. The RTX PRO 6000 Server Edition page gives no FP64 figure.

The RTX PRO 6000 brings what the Hopper card lacks. NVIDIA lists 4 PFLOPS of FP4 Tensor performance for it, while the H200 NVL’s table stops at FP8 and INT8, so FP4 arithmetic runs natively only in the Blackwell pool. FP4 weights can still be served on Hopper without FP4 arithmetic: OpenAI’s gpt-oss-120b, with MXFP4 expert weights, names the H100 among the GPUs it fits. Our guide to FP8, NVFP4 and MXFP4 explains the formats. The card has 188 RT cores and four video encode and four decode engines. NVIDIA’s MIG guide adds +gfx profiles, new with this generation, which “enables graphics support in MIG instances”. The H200 NVL lists decoders only, seven NVDEC and seven JPEG.

Both cards take MIG, in different sizes. The RTX PRO 6000 splits into up to four instances of 24 GB. The H200 NVL splits into up to seven of 16.5 GB, by NVIDIA’s product page, a profile the MIG guide’s table for the H200 141GB names 1g.18gb.

Workload placement: which card for which job

WORKLOADPOOLREASON
Model above 96 GB in FP8H200 NVL, 2 or 4 bridgedweights and cache split over NVLink at 900 GB/s per GPU
Long context, many usersH200 NVL141 GB per card leaves more room for the KV cache
Fine-tuning and trainingH200 NVLmemory per card, NVLink for sharded weights
FP64 simulationH200 NVL30 TFLOPS of FP64
Chat model up to 96 GBRTX PRO 6000one copy per card, scaled by adding copies
NVFP4 checkpointsRTX PRO 6000FP4 Tensor Cores; Hopper has none
Embeddings, rerankersRTX PRO 6000 with MIGfour isolated 24 GB slices per card
Virtual desktops, renderingRTX PRO 6000on NVIDIA’s vGPU list, RT cores, encoders, MIG with graphics

NVIDIA H200 and RTX PRO 6000 Blackwell Server Edition product pages, NVIDIA MIG user guide (11 September 2026) and vGPU supported GPUs list (2 October 2026), read on 10 October 2026; the placement is our reading of these properties.

Qwen3-235B-A22B-Instruct-2507 in FP8 shows why large models go to the H200 pool. Qwen’s FP8 checkpoint of it takes 236.4 GB (220.2 GiB) in its Hugging Face file list, and its model card serves it with --tensor-parallel-size 4. Its config.json lists 94 layers and 4 KV heads of dimension 128, so one conversation of 32K tokens takes 2.94 GiB in an FP8 cache (--kv-cache-dtype fp8 in vLLM). Our sizing rule gives the cache 0.9 × the memory the driver reports (95.6 GiB on the RTX PRO 6000, 140.4 GiB on the H200 NVL), minus 3 GiB per card, minus the weights.

By that rule, four bridged H200 NVL hold about 93 such conversations beside the weights. Four RTX PRO 6000 hold about 38, with the traffic between cards on PCIe, and two H200 NVL about 9. With tensor parallelism over four cards each card keeps one of the four KV heads, since vLLM’s blog of 7 August 2026 states that “TP splits the KV cache by those heads first”. Over eight cards the same blog warns that “once TP exceeds the number of KV heads, the cache starts duplicating across GPUs”.

A model that fits one card runs in either pool. OpenAI’s model card places gpt-oss-120b in use cases that “fit into a single 80GB GPU”, so one copy per RTX PRO 6000 serves it, while the H200 NVL pool remains available for the jobs only it can take. Single-user generation is bounded by memory bandwidth, three times higher on the H200 NVL; where that latency is part of a service level, the model belongs there instead. Fine-tuning is compared in detail in our guide to the H200 NVL and RTX PRO 6000 for fine-tuning.

We supply both cards, the H200 NVL with its two-way and four-way NVLink bridges, as cards or in AI servers built to order. Send us your list of models and workloads through the form below, and we reply with a configuration.

Example fleets of two to four servers

FLEETH200 NVL SERVERSRTX PRO 6000 SERVERSWHAT RUNS WHERE
Two servers1 × 4 cards, bridged1 × 4 cardslarge model on the H200 server; mid-size models, embeddings in MIG on the RTX server
Three servers1 × 4 cards, bridged2 × 4 cardsas above, with chat copies spread over two RTX hosts so one can fail
Four servers2 × 4 cards, bridged2 × 8 cardsone copy of the large model per H200 server, fine-tuning in agreed windows; chat, coding, MIG and virtual desktops on the RTX hosts

Example layouts, not a sizing; counts depend on the models, context length and peak concurrency. Bridging per NVIDIA’s H200 product page: 2 or 4 cards per NVLink domain.

In the two-server fleet, the large model stops when its one server stops, for a failure, a reboot or a driver update. The three-server fleet keeps the mid-size models running with one RTX host down, but not the large model. In the four-server fleet, each H200 server holds one copy of a model such as Qwen3-235B-A22B-Instruct-2507, about 93 conversations of 32K each by the estimate above, so one host down halves that capacity instead of stopping the service. Our article on one 8-GPU server or two 4-GPU servers covers that failure-domain choice in depth.

Fine-tuning and serving on the same H200 cards compete for memory. In the four-server fleet, a training run takes one H200 server in an agreed window while the other serves the large model alone, and the run has to fit the window.

Scheduling a mixed, heterogeneous GPU cluster in Kubernetes

Both cards appear to Kubernetes as the same extended resource, nvidia.com/gpu, so a pod that asks for one GPU lands on any node with an unused GPU unless something constrains it. GPU Feature Discovery, which NVIDIA’s GPU Operator installs, labels each node with the card model in nvidia.com/gpu.product, its memory in MiB in nvidia.com/gpu.memory and the compute capability in nvidia.com/gpu.compute.major. NVIDIA’s compute capability table puts the H200 at 9.0 and the RTX PRO 6000 Server Edition at 12.0, so that label separates the pools.

  1. Install the GPU Operator on every GPU node; GPU Feature Discovery labels each node with its card model, memory and compute capability.
  2. Taint the H200 NVL nodes, for example with gpu-pool=h200:NoSchedule, so only pods with a matching toleration are placed there; the Kubernetes documentation uses GPU nodes as its example of nodes with special hardware. Add the same toleration to the GPU Operator’s daemonsets.tolerations and to its Node Feature Discovery worker, whose defaults tolerate only the nvidia.com/gpu key; otherwise the driver, device plugin and labelling pods are not scheduled on those nodes.
  3. Give every GPU workload a node affinity on one of the labels, as a required rule (requiredDuringScheduling​IgnoredDuringExecution), under which “The scheduler can’t schedule the Pod unless the rule is met.”
  4. Keep MIG on the RTX PRO 6000 nodes that serve small models, and whole cards on the nodes that serve large ones; with the mixed MIG strategy the slices appear as their own resource, nvidia.com/mig-1g.24gb, not as nvidia.com/gpu.
  5. Test each container image in both pools before it is allowed to run in both.

The taint keeps general GPU work off the H200 pool, and the affinity keeps an NVFP4 model or a large tensor-parallel deployment on the right cards. MIG layouts and the sharing methods in each pool are covered in our guide to GPU sharing in Kubernetes.

One driver version and container images for both pools

The two cards do not need separate driver branches. The platform support page of NVIDIA’s GPU Operator, updated on 23 September 2026, lists both the H200 NVL and the RTX PRO 6000 Blackwell Server Edition and names driver 595.91.07 as recommended and default for the 26.7 releases, beside drivers of the 580, 610 and 615 branches. Its notes for the RTX PRO 6000 Server Edition say that it needs driver 575.57.08 or later, that MIG is not supported on 575.57.08 itself, and that Heterogeneous Memory Management in UVM may have to be disabled where CUDA initialisation fails.

One driver version means one test and one rollout procedure for the whole cluster. The rollout still runs pool by pool in maintenance windows, so the large model and the chat copies are never down at the same time.

Container images need more care than the driver. NVIDIA’s CUDA 12.8 release notes add “compiler support” for SM_120, the RTX PRO 6000’s architecture, while the H200 is SM_90. An image meant for both pools has to carry compiled code or PTX for both architectures, and a kernel built for FP4 runs only on the Blackwell pool. Pin such deployments to nvidia.com/gpu.compute.major 12.

Virtual machines and virtual desktops on a mixed platform

Virtual desktops and virtual workstations belong in the RTX PRO 6000 pool. NVIDIA’s list of GPUs supported by its vGPU software, updated on 2 October 2026, includes the RTX PRO 6000 Blackwell Server Edition from release 19.0 and does not list the H200 NVL. Both cards appear among the supported discrete GPUs of the NVIDIA AI Enterprise 8.2 support matrix of 2 September 2026, which covers bare-metal and virtualised deployments but does not map them to each card. A company can therefore keep desktops on RTX PRO 6000 hosts and run the H200 NVL for compute under NVIDIA AI Enterprise; before placing H200 NVL cards in virtual machines, confirm the mode in the documentation of the release you deploy.

Licences across both pools

NVIDIA’s licensing guide, updated on 2 September 2026, states that NVIDIA AI Enterprise “is licensed on a per-GPU basis” and that “Each NVIDIA H200 NVL Tensor Core GPU includes a five-year NVIDIA AI Enterprise subscription”, which has to be activated. The guide names no included subscription for the RTX PRO 6000 Server Edition. Where the RTX pool runs software that needs NVIDIA AI Enterprise, each of its cards needs a licence of its own, while each H200 NVL brings its own five-year subscription.

NVIDIA AI Enterprise and vGPU licences come on the same invoice as the hardware. Tell us which pool runs which software, and we include the licences each card needs in the quote.

What we supply

We supply the H200 NVL with its two-way and four-way NVLink bridges and the RTX PRO 6000 Blackwell Server Edition, as GPUs for servers you already run or in AI servers built to order, on one EU contract and invoice with manufacturer warranty. We build training and fine-tuning nodes with H200 NVL cards and NVLink bridges or with the RTX PRO 6000 Server Edition, and inference nodes with 2 to 8 GPUs, assembled and burn-in tested, with the operating system, drivers, CUDA and a container runtime installed on request. We check the rack, power and airflow before we quote, and send the configuration and quote within one business day. The platform on top, models, RAG and MLOps, is our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

H200 NVL or RTX PRO 6000: which should a company buy?
The RTX PRO 6000 Server Edition, with 96 GB and MIG in four slices, suits many mid-size models, embeddings, NVFP4 checkpoints and virtual desktops. The H200 NVL, with 141 GB, 4.8 TB/s and NVLink bridges, suits models above 96 GB, long contexts with many users, fine-tuning and FP64 work. A platform of two to four servers often uses both.
Can H200 NVL and RTX PRO 6000 run in one Kubernetes cluster?
Yes. NVIDIA’s GPU Operator lists both cards as supported, with driver 595.91.07 as the default. GPU Feature Discovery labels each node with its card model, memory and compute capability, 9 for the H200 and 12 for the RTX PRO 6000, so workloads can be pinned to a pool.
How do I keep workloads off the H200 NVL nodes in a mixed GPU cluster?
Taint the H200 NVL nodes with NoSchedule, so only pods with a matching toleration are placed there, and add that toleration to the GPU Operator’s own pods. Then give each GPU workload a required node affinity on a GPU Feature Discovery label such as nvidia.com/gpu.compute.major. Both cards advertise the same nvidia.com/gpu resource, so without these rules a pod can land on either.
Can the H200 NVL and RTX PRO 6000 use the same NVIDIA driver?
Yes. NVIDIA’s GPU Operator platform page lists both cards and names 595.91.07 as its default driver for the 26.7 releases. The RTX PRO 6000 Server Edition needs driver 575.57.08 or later, and MIG on it is not supported on 575.57.08 itself.
Can the H200 NVL run NVFP4 models?
Not with FP4 arithmetic. NVIDIA lists FP4 Tensor performance for the RTX PRO 6000 Server Edition, while the H200 NVL’s specifications stop at FP8 and INT8, so FP4 weights run there only without FP4 arithmetic, as with gpt-oss-120b, which OpenAI sizes for a single 80 GB GPU. In a mixed fleet, NVFP4 checkpoints go to the Blackwell pool, and FP8 models run in either.
Does the RTX PRO 6000 include NVIDIA AI Enterprise like the H200 NVL?
No. NVIDIA’s licensing guide states that each H200 NVL includes a five-year NVIDIA AI Enterprise subscription, which has to be activated, and names no included subscription for the RTX PRO 6000 Server Edition. NVIDIA AI Enterprise is licensed per GPU, so RTX cards that run software needing it need licences of their own.

Send us the models you plan to run, their precision and context length, the workloads beside them (fine-tuning, embeddings, virtual desktops) and the number of servers your racks can take. We reply within one business day with a configuration that splits the work between H200 NVL and RTX PRO 6000 servers, and a quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna