BLOG · COMPARISON ·

RTX 6000 Ada vs L40S: the same Ada cores and 48 GB, built for a workstation or a server

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • The RTX 6000 Ada and the L40S have the same Ada Lovelace core configuration (18,176 CUDA cores, 568 Tensor Cores, 142 RT cores) and 48 GB of GDDR6 with ECC on a dual-slot, 10.5 inch board with PCIe 4.0 x16
  • The RTX 6000 Ada is an actively cooled 300 W workstation card with 960 GB/s; the L40S is a passive 350 W data-centre card with 864 GB/s that depends on server airflow
  • The L40S ships with its display outputs off, the mode NVIDIA vGPU requires; the RTX 6000 Ada ships with them on and must be switched before vGPU use; neither card has MIG or NVLink
  • HPE’s accelerator QuickSpecs of 8 September 2026 list the L40S for ProLiant Gen11 and Gen12 servers and do not list the RTX 6000 Ada; Supermicro’s 2U GPU server lists only the L40S of the two, its 4U tower lists both
  • For 4 to 8 cards in rack servers, order the L40S; the RTX 6000 Ada belongs in deskside workstations or in tower systems whose maker lists it

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

RTX 6000 Ada vs L40S: the difference in short

The RTX 6000 Ada and the L40S carry the same Ada Lovelace core configuration and the same 48 GB of GDDR6 with ECC, so the same models fit on both and their compute peaks are within one per cent. They differ in where they are meant to run. The RTX 6000 Ada is a workstation card with its own fan, a 300 W board power, 960 GB/s of memory bandwidth and its display outputs switched on. The L40S is a passive data-centre card at 350 W and 864 GB/s that takes its cooling from the server’s fans, ships with its displays off and appears on server makers’ GPU support lists. For a rack server with 4 to 8 cards, that last point decides the choice more often than any specification.

Both cards are on NVIDIA’s current support lists for vGPU and NVIDIA AI Enterprise, and neither has MIG, NVLink or FP4 Tensor Cores. The rest of this article goes through the specifications, the cooling, the display mode, the server lists and the choice for a deployment of several servers or workstations.

Specifications and support side by side

SPECIFICATIONRTX 6000 ADAL40S
CUDA, Tensor, RT cores18,176, 568, 14218,176, 568, 142
Memory48 GB GDDR6 with ECC48 GB GDDR6 with ECC
Memory bandwidth960 GB/s864 GB/s
FP3291.1 TFLOPS91.6 TFLOPS
FP8 Tensor, sparsity1,457 TFLOPS1,466 TFLOPS
Board power300 W350 W default and maximum
Coolingactivepassive, airflow in either direction
Power connectorone 16-pin (PCIe CEM5)one 16-pin, cable strapped for 450 or 600 W
Size and slot4.4 × 10.5 in, dual slot4.4 × 10.5 in, dual slot
System interfacePCIe 4.0 x16PCIe 4.0 x16
Display outputs4× DP 1.4a, on by default4× DP 1.4a, off by default
Encode, decode3 and 3, with AV13 and 3, with AV1
MIG, NVLinkno, nono, no
Secure boot, NEBSnot listedroot of trust, NEBS Level 3
vGPU on vSpherevWS, vPC, vAppsvWS, vPC, vApps
NVIDIA AI Enterprise 7.8on the support matrixon the support matrix

NVIDIA RTX 6000 Ada datasheet (2647623, February 2023) and product page; NVIDIA L40S product page and product brief PB-11470-001_v02 (August 2023); NVIDIA NVENC support matrix (updated 31 July 2026); NVIDIA vGPU R595 release notes for vSphere (updated 29 September 2026); NVIDIA AI Enterprise support matrix 7.8 (updated 12 August 2026); all read on 10 October 2026.

NVIDIA names only the Ada Lovelace architecture on both product pages and does not name the chip. Its Ada architecture whitepaper lists the L40, the earlier data-centre card with the same 18,176 CUDA cores, as an AD102 GPU with 142 streaming multiprocessors, out of 144 on the full AD102. For the video engines, the RTX 6000 Ada product page speaks of “two encode and two decode engines”, while its datasheet lists three of each. NVIDIA’s NVENC support matrix, updated 31 July 2026, also counts three encoders for the card, as for the L40S, so the table follows the datasheet.

Memory bandwidth, compute and power for LLM inference

The RTX 6000 Ada reads its memory at 960 GB/s and the L40S at 864 GB/s. In compute the L40S is slightly ahead, at 91.6 TFLOPS in FP32 against 91.1, and it may draw up to 350 W against 300 W. Generating tokens for one user of a dense language model reads all weights once per token, so the speed of a single stream follows memory bandwidth. By our arithmetic the RTX 6000 Ada has a ceiling about 11 per cent higher for the same model. Reading long prompts and serving many requests in a batch depend more on Tensor Core throughput, where NVIDIA’s FP8 figures with sparsity, 1,457 and 1,466 TFLOPS, are within one per cent of each other.

Memory capacity is the same, so the same models fit. The 48 GB GPUs comparison lists what 48 GB holds, and the measured MLPerf and NIM results for the L40S are in our L40S benchmarks article. We quote no benchmark here, since a gap of this size depends on the engine, the precision and the batch more than on the card. Neither card has FP4 Tensor Cores or NVLink, so a model split across cards exchanges data over PCIe Gen4 on both.

The power connector needs attention on the L40S. NVIDIA’s product brief states that the 16-pin cable “must be strapped for the 450 W or 600 W power level for the L40S to boot”, so the GPU power cable in the server has to be the one its maker lists for the card. The RTX 6000 Ada takes one PCIe CEM5 16-pin connector at a 300 W board power.

Passive L40S and actively cooled RTX 6000 Ada in a server

The L40S has no fan. NVIDIA describes a bidirectional heat sink that “accepts airflow either left-to-right or right-to-left”, so the card relies entirely on the air the server’s fans push through it. The RTX 6000 Ada carries its own fan, which NVIDIA lists as an active thermal solution. That design is made for a workstation case.

In a rack server, whether the server maker has tested and listed the card matters as much as its temperature. Server makers qualify each GPU option with a thermal kit, a riser and a power cable for a given chassis. Lenovo’s product guide for the RTX 6000 Ada as a ThinkSystem option says its part numbers “are for thermal kits and include other components needed to install the GPU”, and HPE’s accelerator QuickSpecs refer to each platform’s QuickSpecs “for configuration rules including enablement kits”. A card that is not on the list has no qualified kit or cable from the maker for that server, which makes it a question for the support contract as much as for temperatures.

Display outputs, default mode and vGPU

Both cards have four DisplayPort 1.4a outputs. The L40S ships in display-off mode, and its product brief states that this default “is required to run NVIDIA Virtual GPU software”. The RTX 6000 Ada ships with its outputs on, and its datasheet footnote says “Display ports are not active when using vGPU software”. NVIDIA’s vGPU release notes for vSphere list the L40S as display-off and the RTX 6000 Ada as display-enabled from the factory, state that such GPUs “must be used in NVIDIA vGPU software deployments in display-off mode” and name the displaymodeselector tool for changing the mode. A fleet of RTX 6000 Ada cards bound for vGPU hosts therefore needs that step on every card before the hosts go into production.

Once in display-off mode, the two cards have the same vGPU support. The R595 release notes for vSphere, updated 29 September 2026, list both for vSphere 8.0 Update 3 and 9.0 with RTX Virtual Workstation, Virtual PC and Virtual Applications, and on both a virtual machine can take several Q-series vGPUs. Neither card supports MIG: NVIDIA’s MIG user guide, updated 11 September 2026, lists no Ada Lovelace GPU, so both are shared between virtual machines in time slices.

Server makers’ GPU support lists

The lists we read on 10 October 2026 point the same way. HPE’s QuickSpecs for NVIDIA accelerators, version 83 of 8 September 2026, list the L40S for the ProLiant DL320, DL380 and DL380a in Gen11 and Gen12, the DL145 Gen11, the DL385 Gen11, the ML350 Gen12 and the DL345 Gen12, among other HPE systems, and contain no RTX 6000 Ada option. Supermicro’s 2U GPU server SYS-221GE-NR takes up to four double-width cards and lists the L40S among them but not the RTX 6000 Ada.

Supermicro’s SYS-741GE-TNRT, a tower that converts to a 4U rack system, lists both cards among up to four double-width GPUs, with an optional rear fan kit for passive GPU cooling. For any other model, the maker’s current GPU support list for that chassis decides. How many cards a given chassis takes, by lanes, power and air, is covered in how many GPUs fit in one server.

We supply both cards and build on your chassis after a compatibility check of the platform, power and cooling. Tell us the server model and how many cards it should take in the form below.

NVIDIA AI Enterprise and NVIDIA-Certified Systems

The NVIDIA AI Enterprise support matrix for release 7.8, updated 12 August 2026, lists both the “NVIDIA L40S” and the “NVIDIA RTX 6000 Ada Generation” under Ada Lovelace. It qualifies that list with the condition that the software “is supported on the following NVIDIA GPUs with compatible third-party servers listed on the NVIDIA-Certified Systems page”. The RTX 6000 Ada product page names NVIDIA AI Enterprise under AI software support. NVIDIA’s certified systems page covers NVIDIA-Certified Workstations as well as servers and points to its Qualified System Catalog for the listed models.

For NVIDIA AI Enterprise support, the matrix ties the card to a system on that page, so ask for the listing of the exact server or workstation model with the card before the order. NVIDIA’s licensing guide states that the software “is licensed on a per-GPU basis”, which applies to both cards.

Choosing for 4 to 8 cards: rack servers or deskside

At the scale of several departments, the choice follows the place the cards will run. Four L40S cards draw up to 1,400 W of board power and eight up to 2,800 W, before processors and fans. Four RTX 6000 Ada cards draw up to 1,200 W. Supermicro’s SYS-741GE-TNRT, which lists the card, takes up to four double-width GPUs on two redundant 2,000 W power supplies.

DEPLOYMENTCARDREASON
Rack server, 4 to 8 cardsL40Spassive, on the server makers’ lists, displays off
vGPU hosts for VDIL40Sdisplay-off default, listed servers
Deskside workstationsRTX 6000 Adaown fan, displays on, 300 W
Tower or tower-rack systemeither, as the maker lists itboth listed for some towers
Existing server, not listedneither until checkedno qualified kit or cable
Models above 48 GBRTX PRO 600096 GB against 48 GB on both Ada cards

Server lists from HPE QuickSpecs version 83 (8 September 2026) and Supermicro product pages, read on 10 October 2026; board power from NVIDIA.

An engineering firm that gives eight designers one RTX 6000 Ada each in a deskside workstation runs eight 300 W cards with their own fans and displays. The same firm moving those users to virtual workstations would order two rack servers with four L40S each, on a server model that lists the card, with vGPU licences for the virtual workstations. Where a model outgrows 48 GB, our comparison of the L40S and the RTX PRO 6000 Server Edition covers the next step in servers, and RTX PRO 6000 Blackwell vs RTX 6000 Ada the step in workstations.

Send us the number of users, the workloads and where the machines will stand through the form below, and we reply with a configuration and quote within one business day.

What we supply

We supply the RTX 6000 Ada and the L40S as professional NVIDIA GPUs, and AI servers built to order with the L40S for rack deployments, with manufacturer warranty on one EU contract and invoice. For servers you already run, we build on your chassis after a compatibility check of the platform, power and cooling. NVIDIA AI Enterprise and vGPU licences come on the same invoice as the hardware. Configuration and quote follow within one business day.

FAQ

RTX 6000 Ada vs L40S: what is the difference?
Both have the same Ada Lovelace core configuration and 48 GB of GDDR6 with ECC. The RTX 6000 Ada is an actively cooled 300 W workstation card with 960 GB/s and its displays on, while the L40S is a passive 350 W data-centre card with 864 GB/s and its displays off by default. HPE’s accelerator QuickSpecs and Supermicro’s 2U GPU server list the L40S and not the RTX 6000 Ada.
L40S vs RTX 6000 Ada for LLM inference: which is faster?
For one user of a dense model the RTX 6000 Ada has a ceiling about 11 per cent higher, because token generation follows memory bandwidth, 960 against 864 GB/s. For long prompts and large batches NVIDIA’s FP8 figures are within one per cent of each other. Both hold the same models in 48 GB, so the server or workstation they go into decides more than speed.
Can I put an RTX 6000 Ada in a rack server?
Only where the server maker lists it, because each GPU option is qualified with a thermal kit, riser and power cable for that chassis. HPE’s accelerator QuickSpecs of 8 September 2026 contain no RTX 6000 Ada option, and Supermicro’s 2U GPU server lists the L40S but not the RTX 6000 Ada. Supermicro lists both cards for its 4U tower-rackmount system.
L40S passive vs RTX 6000 Ada blower: what does the cooling change?
The L40S has no fan and takes its cooling from the server’s fans, with a heat sink that accepts airflow in either direction. The RTX 6000 Ada carries its own fan, which NVIDIA lists as an active thermal solution, and suits a workstation case. In a server, the maker’s list of qualified cards matters more than the cooling method.
Do the RTX 6000 Ada and L40S support vGPU and NVIDIA AI Enterprise?
Yes. NVIDIA’s vGPU release notes for vSphere list both with RTX Virtual Workstation, Virtual PC and Virtual Applications, and the AI Enterprise 7.8 support matrix lists both under Ada Lovelace, supported with servers on the NVIDIA-Certified Systems page. The RTX 6000 Ada has to be switched to display-off mode before vGPU use, and neither card supports MIG.
Can the L40S go into a workstation?
Only into a system built to push air through a passive card. Supermicro’s SYS-741GE-TNRT, a tower that converts to a 4U rack system, lists the L40S and offers a rear fan kit for passive GPU cooling. In an ordinary workstation case without such airflow, the actively cooled RTX 6000 Ada is the card designed for it.

Send us the workload, the number of cards and whether they go into rack servers you already run or into deskside workstations. We reply within one business day with the card that fits, a configuration and a quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna