Mixing Ada and Blackwell GPUs: one driver branch, different GPUs in one server and in one cluster
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- NVIDIA’s data-centre drivers 580.178.04 and 595.91.07 of 3 August 2026 list the RTX Ada cards, the L40S and the L4 beside the RTX PRO 6000, 5000, 4500, 4000 and 2000 Blackwell, so one installed driver runs both generations
- FP8 Tensor Cores exist on both generations; FP4 Tensor Cores exist only on the Blackwell cards, MIG only on the RTX PRO 6000, 5000 and 4500, and NVIDIA’s MIG guide lists no Ada GPU
- Container images need code for compute capability 8.9 (Ada) and 12.0 (RTX PRO Blackwell); binaries that carry only cubins for older architectures have to be rebuilt for Blackwell
- By our reading, a model split with tensor parallelism over an L40S and an RTX PRO 6000 is held to the 48 GB and 864 GB/s of the L40S, so each pool serves its own models
- NVIDIA’s vGPU guide describes XenServer hosts with several GPU types, grouped by type, but requires a GPU of the same type on the destination host for VM migration, so each GPU type is its own migration domain
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Mixing Ada and Blackwell GPUs at the host and driver level
Ada and Blackwell GPUs can be mixed at the level of the host and the driver. NVIDIA’s data-centre drivers 580.178.04 and 595.91.07, both released on 3 August 2026, list the RTX 6000, 5880, 5000, 4500 and 4000 Ada Generation cards, the L40S and the L4 beside the RTX PRO 6000 Blackwell in all three editions and the RTX PRO 5000, 4500, 4000 and 2000. One installed driver therefore runs both generations, in one server or across a fleet.
Several features do not span the two generations. FP4 Tensor Cores exist only on the Blackwell cards and MIG only on the larger ones, and by our reading a model split across both generations runs at the pace of the slower card. A virtual machine with vGPU migrates only to a host with the same GPU type, and container images need code for both architectures.
This guide is for a company that runs 4 to 8 Ada cards, such as two servers with four L40S each, and now adds RTX PRO Blackwell cards. The pattern we recommend for such a fleet is one GPU model per host, both generations as separate pools in one cluster, and a list of which workloads stay on Ada. Mixing the H200 NVL with the RTX PRO 6000 is covered in our article on running H200 NVL and RTX PRO 6000 in one platform, and the cards themselves in our tier-by-tier comparison of RTX Ada and RTX PRO Blackwell.
One NVIDIA driver for Ada and Blackwell
The product lists of R580 and R595 contain the cards named above, with two exceptions among the cards we supply: the RTX PRO 4500 Blackwell Server Edition appears only in the R595 list, and the RTX PRO 5500 appears in neither yet. NVIDIA’s GPU Operator platform page (23 September 2026) names 595.91.07 as the recommended and default driver of release 26.7.1. It lists the L40S and the L4 under Ada and the RTX PRO 6000 and RTX PRO 4500 Blackwell Server Edition under Blackwell. We found the RTX 6000 Ada on that page only in its list of GPUs for vGPU with KubeVirt.
A host with a Blackwell card needs the open kernel modules. NVIDIA wrote on 17 July 2024 that for Blackwell “you must use the open-source GPU kernel modules”, and it recommended the same modules for Turing, Ampere, Ada Lovelace and Hopper GPUs. A host that ran the proprietary flavour for its Ada cards switches to the open modules with the first Blackwell card, and the Ada cards run on them too.
Branch lifecycles, CUDA compatibility and pinning the branch are in our guide to NVIDIA driver branches and CUDA versions.
What does not span both generations
| FEATURE | L40S (ADA) | RTX PRO 6000 SE | EFFECT OF MIXING |
|---|---|---|---|
| Compute capability | 8.9 | 12.0 | images need code for both |
| First CUDA toolkit | 11.8 | 12.8 | older builds lack Blackwell code |
| Memory, bandwidth | 48 GB GDDR6, 864 GB/s | 96 GB GDDR7, 1,597 GB/s | a split model is held to the L40S |
| Host interface | PCIe Gen4 x16 | PCIe Gen5 | traffic to the L40S at Gen4 rate |
| FP8 Tensor Cores | yes | yes | FP8 is the shared format |
| FP4 Tensor Cores | no | yes | NVFP4 models on Blackwell only |
| MIG | no | up to 4 instances | MIG layouts on Blackwell nodes only |
| Kernel modules | open or proprietary | open only | a mixed host runs the open modules |
| vGPU, first release | 16.1 | 19.0 | VMs migrate only to the same GPU type |
| Maximum power | 350 W | up to 600 W, configurable | power and airflow planned per card |
NVIDIA’s L40S and RTX PRO 6000 Server Edition pages, CUDA GPUs list, CUDA 11.8 and 12.8 release notes, MIG guide, vGPU GPU list and blog of 17 July 2024, read on 10 October 2026; the effects are our reading. The RTX 6000 Ada has the same 48 GB, PCIe Gen 4 and compute capability, at 300 W.
FP8 is the format both generations share. NVIDIA lists FP8 Tensor Core performance for the L40S and the RTX PRO 6000 Server Edition, and FP8 Tensor Cores for the RTX 6000 Ada. FP4 Tensor Core performance appears only on the Blackwell page; the L40S table stops at INT8 and INT4. NVFP4 checkpoints therefore belong on the Blackwell cards; FP8 and 16-bit models run in either pool if they fit.
NVIDIA’s MIG guide, updated on 11 September 2026, lists the RTX PRO 6000 with up to four instances and the RTX PRO 5000 and 4500 with two, and no Ada Lovelace GPU; the L40S page states MIG support as “No”.
Container images need the most care. Ada support came with CUDA 11.8 and compiler support for SM_120, the RTX PRO Blackwell architecture, with CUDA 12.8. NVIDIA’s Blackwell compatibility guide, updated on 13 September 2026, states that binaries which “only include cubins” for older architectures “need to be rebuilt to run on the Blackwell GPUs”, while binaries that include PTX “should work as-is”. An image built for the Ada pool with only sm_89 code fails on the Blackwell cards. Images meant for both pools carry one -gencode entry per architecture, for example -gencode=arch=compute_ next to the one for sm_89.
Splitting one model across Ada and Blackwell cards
Tensor parallelism divides every layer into equal shares, one per card, and the cards exchange partial results within each layer. By our reading, each step therefore waits for the slowest card, and each card can use only as much memory as the smallest one has. Two L40S and two RTX PRO 6000 Server Edition under tensor parallelism of four behave like four 48 GB cards: 192 GB usable of the 288 GB installed, at the memory bandwidth of the L40S, with the traffic to the L40S on PCIe Gen4.
We found no statement on mixing GPU models in vLLM’s parallelism documentation (6 May 2026). It asks that “every node provides an identical execution environment, including the model path and Python packages”, and for GPUs that “do not have NVLINK interconnect (e.g. L40S)” it suggests pipeline parallelism rather than tensor parallelism for throughput. The workable layout keeps each model inside one pool, scaled by adding copies, with the router in front of the inference engines choosing the pool.
Different GPUs in the same server
The driver accepts an L40S and an RTX PRO 6000 in one chassis, but server makers define power cables, fan sets and maximum quantities per card model. Our article on planning a GPU server for growth found no Dell or Lenovo rule on mixing models and recommends growing with the model a server started with.
CUDA enumerates devices “from fastest to slowest using a simple heuristic” by default, so in a mixed server device 0 may be the Blackwell card whatever its slot. Setting CUDA_DEVICE_ORDER=PCI_BUS_ID makes CUDA number the cards “by PCI bus ID in ascending order”, and CUDA_VISIBLE_DEVICES then pins each job to the card meant for it.
In Kubernetes the device plugin advertises every card as nvidia.com/gpu, and GPU Feature Discovery labels the node, not each card; its documentation does not describe a node with two models. A pod that asks for one GPU on a mixed node can receive either card. NVIDIA’s DRA driver, which allocates each device on its own, states that its GPU allocation features “are not yet officially supported”. Mixed cards suit a development workstation where each card runs its own job; shared production servers keep one GPU model per host.
Ada and Blackwell node pools in Kubernetes
- Keep one GPU model per node, with the Ada and Blackwell nodes as two pools on one driver version.
- Let GPU Feature Discovery label each node:
nvidia.com/gpu.productwith the card model andnvidia.com/gpu.compute.majorwith 8 on Ada and 12 on RTX PRO Blackwell;nvidia.com/gpu.compute.minor9 separates Ada from any Ampere nodes still in the cluster. - Give NVFP4 models and images built only for
sm_120a required node affinity on compute major 12, and images built only forsm_89an affinity on the Ada labels. - Configure MIG on the Blackwell nodes only; the Ada nodes keep whole cards or time-slicing.
- Test each image on one node of each pool before it is allowed to run in both.
A taint on the Blackwell nodes keeps general GPU work off them. The GPU Operator’s own pods then need the matching toleration, or the driver and device plugin are not scheduled there, as our H200 NVL and RTX PRO 6000 article describes. The sharing methods are in our guide to GPU sharing in Kubernetes with MIG, time-slicing and MPS.
vGPU hosts with mixed GPU types
One vGPU release covers both generations. NVIDIA’s list of supported GPUs, updated on 2 October 2026, gives the first release for each card, for example 15.2 for the RTX 6000 Ada, 16.1 for the L40S and 19.0 for the RTX PRO 6000 Blackwell Server Edition. Hosts of both kinds can run the same vGPU Manager of release 19 or 20; the RTX PRO 4500 Blackwell Server Edition needs release 20.0 or later, and the RTX PRO 6000 Workstation and Max-Q editions are not on the list.
On each physical GPU, the vGPU user guide (29 September 2026) states that “By default, a GPU or GPU instance supports only vGPUs with the same amount of frame buffer”, and that profiles of different sizes need the GPU to be “put into mixed-size mode”.
Different GPU types in one vGPU host are documented for XenServer. NVIDIA’s user guide says XenServer creates GPU groups at startup “to represent the distinct types of physical GPU present on the platform”, each “a collection of physical GPUs, all of the same type”, and its example host carries three GPU models. We found no such statement for the other hypervisors.
VM migration stays within one GPU type. For every hypervisor the guide lists “The NVIDIA GPUs on both host machines must be of the same type”, and its general prerequisites add that the GPU topologies on both hosts must be identical. A virtual desktop on an L40S can move only to a host with an L40S that has room for its profile. Each GPU type is therefore its own migration domain and needs spare room for the VMs of a host in maintenance. Two L40S hosts and two RTX PRO 6000 hosts make two domains of two hosts, not one of four.
NVIDIA AI Enterprise and vGPU licences come on the same invoice as the hardware. Tell us which hypervisor and vGPU profiles you run on the Ada hosts today, and we include the licences the new cards need in the quote.
Which workloads stay on Ada, and mixing patterns
| PATTERN | WHEN IT FITS | WHAT TO WATCH |
|---|---|---|
| Separate servers per model | production servers, the default | failover sized within each pool |
| Node pools, one cluster | several teams on one platform | labels, affinity, images for both |
| vGPU cluster per GPU type | desktops on Ada and Blackwell | migration only within a type |
| Mixed cards, one workstation | development, one job per card | device order, power, airflow |
| One model over both | not advised | smaller, slower card sets the pace |
Our assessment, based on NVIDIA’s vGPU user guide (29 September 2026), GPU Feature Discovery and CUDA documentation and the first table.
The Ada cards keep work that fits 48 GB per copy and needs no FP4 or MIG: embeddings and rerankers, speech-to-text, chat models whose FP8 weights and cache fit one or two cards, video analytics on the L4 and virtual desktops on L40S or RTX 6000 Ada hosts. The Blackwell cards take NVFP4 checkpoints, models that need 96 GB per card and isolated MIG slices.
Take two servers with four L40S each and one new server with four RTX PRO 6000 Server Edition. The Ada servers run embeddings, a reranker and copies of an FP8 chat model, so one can fail while the other keeps serving. The Blackwell server runs a larger model over its four cards, 384 GB in total, which has no second host until a second Blackwell server joins.
We supply both generations, as cards for servers you already run or in AI servers built to order. Describe your current Ada servers in the form below, with the workloads you want to move, and we send a configuration for the Blackwell side.
What we supply
We supply the RTX 6000, 5880 and 5000 Ada, the L40S and the L4, and the RTX PRO 6000 in all three editions, the RTX PRO 5000 and the RTX PRO 4500 and 4500 Server Edition, as GPUs for the servers and workstations you run or in AI servers built to order. Everything comes on one EU contract and invoice with manufacturer warranty. For servers built to order we install the operating system, drivers, CUDA and a container runtime on request, on the driver version your Ada hosts run. We check the rack, power and airflow before we quote, and send the configuration and quote within one business day.
FAQ
Can you mix Ada and Blackwell GPUs?
Can different GPUs be in the same server?
Can an RTX 6000 Ada and an RTX PRO 6000 run in the same system?
Which NVIDIA driver supports both Ada and Blackwell GPUs?
How do I run mixed GPU generations in Kubernetes?
Can a vGPU host have different GPU types, and can VMs migrate between them?
Send us the Ada cards you run today, the servers or workstations they sit in, the driver branch, whether you run Kubernetes or vGPU, and the workloads you want to move. We reply within one business day with a configuration for the Blackwell cards or servers and a quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day