NVIDIA L40S vs RTX PRO 6000 Blackwell Server Edition: where Ada still wins
- Same job, different era: L40S is Ada with 48 GB GDDR6 at 864 GB/s and 350 W; RTX PRO 6000 Server Edition is Blackwell with 96 GB GDDR7 at 1,597 GB/s and 400–600 W configurable
- Blackwell adds what Ada never had: FP4 Tensor Cores, MIG (four 24 GB instances), a fourth NVENC/NVDEC engine and PCIe Gen5
- L40S still wins on fit: 350 W lets four cards sit where only two 600 W cards can, it needs far less airflow, and it is qualified in almost every Gen4 and Gen5 chassis on the market
- Neither card has NVLink; multi-GPU traffic is PCIe on both. Neither ships with an NVIDIA AI Enterprise subscription
- Rule of thumb: if the model or the vGPU estate fits in 48 GB and the chassis is already bought, L40S; if 96 GB, FP4 or MIG changes the design, RTX PRO 6000
Two cards, one slot, one purpose
Both are the “universal” passive card of their generation: the one a server vendor qualifies first, the one that runs inference, graphics, video and virtual desktops from the same slot. NVIDIA positioned the L40S as exactly that in 2023, and the RTX PRO 6000 Blackwell Server Edition took over the role in 2025. The question we get in almost every configuration call is whether the older card is now the wrong order. Often it is not, and the reasons are mechanical rather than glamorous.
The specification table
| PARAMETER | L40S | RTX PRO 6000 SERVER EDITION |
|---|---|---|
| Architecture | Ada Lovelace | Blackwell |
| CUDA cores | 18,176 | 24,064 |
| Memory | 48 GB GDDR6 with ECC | 96 GB GDDR7 with ECC, 512-bit |
| Bandwidth | 864 GB/s | 1,597 GB/s |
| FP8 Tensor (dense | sparse) | 733 | 1,466 TFLOPS | ~1 | 2 PFLOPS |
| FP4 Tensor | none | 4 PFLOPS (with sparsity) |
| FP32 | 91.6 TFLOPS | 120 TFLOPS |
| Board power | 350 W default (cap adjustable downward) | 400–600 W configurable; 450 W when the power cable signals it |
| Cooling and slots | Passive, dual-slot FHFL | Passive, dual-slot FHFL (air) or single-slot FHXL (liquid) |
| PCIe | Gen4 x16 | Gen5 x16 |
| MIG | No | Up to 4 instances of 24 GB |
| vGPU | Yes, up to 32 vGPUs per card (1Q profiles) | Yes, up to 48 graphics vGPUs per card with MIG (vGPU 19.0+); 12 AI Enterprise compute vGPUs |
| NVENC | NVDEC | 3 | 3, AV1 | 4 | 4, plus 4 JPEG engines |
| Display outputs | 4× DP 1.4a, off by default | 4× DP 2.1, off by default |
| NVLink | No | No |
| Thermal | passive, 350 W of heat | passive, up to 600 W of heat; fan upgrades recommended in some chassis |
| NVIDIA AI Enterprise | not bundled | not bundled |
Sources: NVIDIA product pages and product briefs for L40S (PB-11470) and RTX PRO 6000 Server Edition (SP-12355), Lenovo product guides LP1812 and LP2263. Blackwell PFLOPS figures are NVIDIA’s headline numbers; the workstation datasheet marks them as sparsity figures, so treat dense as roughly half.
Two things in the table matter more than the rest. The first is 96 GB against 48 GB: a 70B model in FP8 is about 68 GiB of weights and does not fit the L40S at all, while it fits the Blackwell card with room for a working KV cache. The second is 1,597 against 864 GB/s: for single-stream generation, which is bandwidth-bound, that is 1.85× more tokens per second before any Tensor Core arithmetic is counted.
What Blackwell adds
FP4. Ada has no 4-bit Tensor path; on the L40S a 4-bit model is weight-only compression that is unpacked to FP8 or FP16 for the arithmetic. Blackwell runs NVFP4 natively, which is why NVIDIA quotes 4 PFLOPS for the format and why a 70B model in NVFP4 (about 40 GB) becomes a one-card deployment with real headroom.
MIG. The L40S cannot be partitioned in hardware; you divide it with time-sliced vGPU and accept that one noisy tenant affects the others. The RTX PRO 6000 Server Edition splits into four isolated 24 GB instances, each with its own memory and compute, which is what makes it the first card of this class that a hosting team can hand to four departments with a straight face.
Density of seats. vGPU 19 lets one Blackwell card with MIG carry up to 48 virtual machines; the L40S tops out at 32 with the smallest 1 GB profile. For VDI and CAD estates, that is the number that sets the card count.
Video. Four encoders, four decoders and four JPEG engines against three and three. If the workload is transcoding or a vision pipeline with many camera streams, the fourth engine is a third more throughput from the same slot.
PCIe Gen5. Doubling the link matters when several cards exchange tensors without NVLink, and neither of these cards has NVLink. On a Gen4 platform the Blackwell card falls back to Gen4 x16 and works, but the advantage is gone.
Where the L40S still wins
Power envelope. 350 W against 600 W by default. A Lenovo SR650a V4 takes four L40S cards but only two RTX PRO 6000 Server Edition at 600 W, or four capped to 450 W. Across a rack of existing 2U hosts, that is the difference between adding cards and replacing servers.
Airflow. Both cards are passive and live on chassis fans. At the same inlet temperature a 600 W card needs roughly 70 per cent more air than a 350 W card, which is why HPE recommends a fan upgrade for 600 W cards in the DL380a Gen12 and why Lenovo limits some configurations to a 30 °C inlet. In an edge room, a factory cabinet or a hall that runs warm, the older card has margin the new one does not.
Qualification breadth. NVIDIA’s certified-systems list carries the L40S in Dell PowerEdge R750xa, R760 and R760xa, HPE DL380a and DL385 Gen11, Lenovo SR650 V3, SR655 V3, SR665 V3 and SR675 V3, and the Supermicro 2U and 4U GPU lines. If the chassis is already on the floor, the L40S is the card it was built for.
Homogeneity. An estate of L40S hosts grows most cheaply with more L40S: one vGPU profile set, one power budget, one spare pool. The driver is not the problem (vGPU 19 and 20 support both cards from the same host driver), but a multi-GPU job across mixed cards runs at the pace of the slower one.
Enough is enough. Diffusion models, most vision models, CAD and VDI seats, video pipelines and language models up to roughly 30B in FP8 all fit comfortably in 48 GB. For those, the extra memory is unused capacity.
Where the RTX PRO 6000 is the only sensible order
Any model above roughly 30B parameters that has to run on one card. Any deployment where four isolated tenants per card is the requirement rather than a nice-to-have. Any new 4U or 5U GPU server: Dell XE7745, HPE DL380a Gen12 (ten), Lenovo SR675 V3 and the Supermicro 4U and 5U lines all take eight or more of the Blackwell cards, and building a new eight-slot machine around 2023 silicon makes little sense. And any workload where NVFP4 is on the roadmap: NVIDIA’s own comparison material puts LLM inference on the new card at up to five times the L40S, and while vendor numbers are vendor numbers, the direction is not in doubt.
Lifecycle and support
Neither card is on an end-of-life list. NVIDIA’s vGPU software lifecycle page, updated in August 2026, lists the L40S, the L40 and the RTX PRO 6000 Server Edition as fully supported, and the L40S remains in current vendor guides including the Xeon 6 generation of Lenovo servers. What changes over time is not support but the workload: the models people ask us to size have grown faster than the cards.
Our engineering partner Vixen.UNO sizes both from the same worksheet: model size, precision, concurrency, context length, then the chassis. The card falls out of that arithmetic; it is rarely the first decision.
FAQ
Is the L40S discontinued?
Can I put an RTX PRO 6000 Server Edition into a PCIe Gen4 server?
Does either card support MIG?
How many VMs does each card carry?
Which one for a 70B model?
Can I mix L40S and RTX PRO 6000 in one server?
Tell us the model, the precision and the chassis you already own, and we will say which card it needs. We reply within one business day.
Talk to an expertWe reply within one business day