L40S vs RTX PRO 4500 Server Edition: 48 GB at 350 W or 32 GB at 165 W per card
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- The L40S has 48 GB of GDDR6 at 864 GB/s on a 350 W dual-slot board; the RTX PRO 4500 Blackwell Server Edition has 32 GB of GDDR7 at 800 GB/s on a 165 W single-slot board, and both are passive
- The Server Edition adds FP4 Tensor Cores, MIG with two 16 GB instances, PCIe 5.0 and 4:2:2 H.264 and HEVC encoding; the L40S has 73 per cent more CUDA cores and 91.6 against 51 TFLOPS of FP32
- By our sizing rule a 32B model in FP8 fits one L40S with cache for about five 8K conversations with ECC on, eight with it off, and does not fit 32 GB, while Qwen3-14B in FP8 fits both cards, for about 35 and 13 conversations with ECC on
- In a 2U layout of four double-wide or eight single-wide GPU slots, Lenovo’s SR650a V4 as an example, eight Server Edition cards give 256 GB of GPU memory at 1,320 W of board power, four L40S give 192 GB at 1,400 W
- NVIDIA sets the RTX PRO 6000 Server Edition against the L40S and the RTX PRO 4500 Server Edition against the L4, and its vGPU list of 2 October 2026 shows the L40S and the RTX PRO 4500 Server Edition as fully supported
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
L40S vs RTX PRO 4500 Server Edition: the short answer
The NVIDIA L40S has 48 GB of GDDR6 at 864 GB/s on a dual-slot, 350 W board, and the RTX PRO 4500 Blackwell Server Edition has 32 GB of GDDR7 at 800 GB/s on a single-slot, 165 W board. Choose the L40S when one model with its cache needs more than about 24 GiB on a card, a 32B model in FP8 for example, or when you extend an existing L40S environment. Choose the RTX PRO 4500 Server Edition when the models fit in 32 GB and the server has single-width slots: it adds FP4 Tensor Cores, two 16 GB MIG instances, PCIe 5.0 and 4:2:2 video encoding, and a 2U server with eight single-width slots can take twice as many of them, where its maker lists the card for those slots.
L40S and RTX PRO 4500 Server Edition specifications
| SPECIFICATION | L40S | RTX PRO 4500 SE |
|---|---|---|
| Architecture | Ada Lovelace | Blackwell |
| CUDA cores | 18,176 | 10,496 |
| GPU memory | 48 GB GDDR6 with ECC | 32 GB GDDR7 with ECC, 256-bit |
| Memory bandwidth | 864 GB/s | 800 GB/s |
| FP8 Tensor | 733 TFLOPS dense, 1,466 with sparsity | 811 TFLOPS, basis not stated |
| FP4 Tensor | none | 1.6 PFLOPS, basis not stated |
| FP32 | 91.6 TFLOPS | 51 TFLOPS |
| Maximum power | 350 W | 165 W |
| Board | dual slot, 4.4 by 10.5 in, passive | single slot, full height, full length, passive |
| Host interface | PCIe 4.0 x16 | PCIe 5.0 x16 |
| MIG | no | up to 2 instances of 16 GB |
| vGPU, first release | 16.1 | 20.0 |
| NVENC and NVDEC | 3 and 3, with AV1 | 3 and 3, with 4:2:2 H.264 and HEVC encoding |
| NVLink | no | no |
NVIDIA L40S product page; NVIDIA RTX PRO 4500 Blackwell Server Edition product page, datasheet (5192801, April 2026) and product brief PB-12790-001_02 (20 April 2026); NVIDIA’s list of GPUs supported by vGPU, updated 2 October 2026.
The L40S is the larger chip, with 73 per cent more CUDA cores and 1.8 times the FP32 rate, but only 8 per cent more memory bandwidth. NVIDIA does not state whether the Server Edition’s 811 TFLOPS of FP8 include sparsity; on the L40S page, 733 is the dense figure and 1,466 the figure with sparsity. If NVIDIA’s Blackwell figure is a sparse one, the L40S has the higher FP8 rate per card. Per server the order reverses: eight Server Edition cards add up to 6,488 TFLOPS, against 2,932 dense or 5,864 sparse for four L40S, by our arithmetic.
Our RTX PRO 4500 Server Edition guide covers the card on its own, and our L40S benchmark review the published measurements of the older one.
Which models fit in 32 GB and in 48 GB
We size with the rule from our guide to how much VRAM an LLM needs: 90 per cent of the card’s memory, less about 3 GiB for activations and the runtime, holds the weights and the KV cache. Both cards have ECC memory, which on GDDR cards reduces the available memory by 6.25 per cent, according to NVIDIA’s CUDA C++ Best Practices Guide. With ECC enabled that leaves about 37.3 GiB for weights and cache on the L40S and 23.9 GiB on the RTX PRO 4500 Server Edition, with ECC off 40.0 and 25.7 GiB; nvidia-smi -q shows the ECC mode of the delivered card. The cache of one conversation follows from the model’s config.json, and with an FP8 cache at a full 8,192-token context it takes 0.5625 GiB for Qwen3-8B, 0.625 GiB for Qwen3-14B and 1 GiB for Qwen3-32B.
| MODEL AND FORMAT | WEIGHTS | L40S | RTX PRO 4500 SE |
|---|---|---|---|
| Qwen3-8B, FP8 | 8.8 GiB | about 50 (55) | about 26 (30) |
| Qwen3-14B, FP8 | 15.2 GiB | about 35 (39) | about 13 (16) |
| Qwen3-32B, NVFP4 | 19.3 GiB | no FP4 Tensor Cores | about 4 (6) |
| Qwen3-32B, FP8 | 32.0 GiB | about 5 (8) | does not fit |
| Llama 3.3 70B, NVFP4 | 39.8 GiB | no room for cache | does not fit |
Concurrent 8,192-token conversations per card with an FP8 KV cache, by our arithmetic, first with ECC enabled, in brackets with ECC off; card memory is our estimate of what the driver reports. Weights: NVIDIA’s FP8 and NVFP4 checkpoints and Qwen’s FP8 release of Qwen3-32B on Hugging Face, read on 10 October 2026; layers and KV heads from each config.json.
The dividing line is a 32B model in FP8, which fits the L40S with room for about five conversations (eight with ECC off) and does not fit the Server Edition. Qwen’s FP8 release scales its weights in blocks of 128 by 128, a format TensorRT-LLM’s support matrix does not list for Ada, as our L40S benchmark review notes, so confirm that your serving engine runs it on the L40S. In NVFP4 the same model fits the Blackwell card for about four conversations (six with ECC off). Ada has no FP4 Tensor Cores, so on the L40S vLLM runs NVFP4 only as weight-only compression, and other 4-bit models run as INT4 through AWQ or GPTQ. Models up to 14B in FP8 fit both, and the L40S holds more than twice the Qwen3-14B conversations per card. Speed for one user is close, because a dense model reads its weights once per token: for Qwen3-14B in FP8, with 16.3 GB of weights, the ceiling is about 53 tokens per second on the L40S and 49 on the Server Edition by our arithmetic.
Per-server totals in a 2U server with four or eight GPU slots
Lenovo Press describes the ThinkSystem SR650a V4, a 2U server, with “Support for up to 4x double-wide GPUs or 8x single-wide GPUs, installed in the front slots” (LP2128, updated 5 October 2026), and lists “4x NVIDIA L40S 350W GPUs” among its features. Lenovo’s guide to the RTX PRO 4500 Server Edition (LP2391) notes in its update of 28 July 2026 that “The SR650a V4 now supports the GPU”. Lenovo’s own eight-card example is “8x NVIDIA L4 single-wide GPUs”, and neither guide we could read states how many RTX PRO 4500 Server Edition cards the server takes. The table uses this layout as an example; the card count per server comes from the maker’s configurator.
| PER 2U SERVER | 4 × L40S | 8 × RTX PRO 4500 SE |
|---|---|---|
| GPU memory | 192 GB | 256 GB |
| Board power, GPUs | 1,400 W | 1,320 W |
| Bandwidth, sum | 3,456 GB/s | 6,400 GB/s |
| FP32, sum | 366 TFLOPS | 408 TFLOPS |
| NVENC and NVDEC | 12 and 12 | 24 and 24 |
| MIG instances | none | 16 of 16 GB |
| Qwen3-14B FP8, 8K | about 140 (156) | about 104 (128) |
| Qwen3-32B FP8, 8K | about 20 (32) | does not fit |
| vPC users, maximum | 128 with 1 GB profiles | 128 with 2 GB profiles |
Example layout, not a configuration: slot counts from Lenovo Press LP2128; per-card figures from NVIDIA’s product pages; conversations from the table above, with ECC on and, in brackets, off; vPC users per board from NVIDIA’s vPC sizing guide, updated 19 August 2026. Sums are our arithmetic.
The eight-card server has a third more GPU memory at slightly lower board power, and nearly twice the summed bandwidth and twice the video engines. Each card holds less, so the Qwen3-14B count falls from about 140 to 104 with ECC on, and a 32B model in FP8 has to be split across two cards by tensor parallelism, over PCIe, since neither card has NVLink. With 8B models the order changes again, about 208 conversations on eight Server Edition cards against 200 on four L40S with ECC on.
We build AI servers to order with either card and check the rack, power and airflow before we quote. Tell us the models, the peak conversations and the server or rack position, and we reply with the card count per server.
MIG and vGPU on the L40S and the RTX PRO 4500 Server Edition
NVIDIA’s MIG user guide, updated 11 September 2026, lists up to two MIG 1g.16gb instances per RTX PRO 4500 Blackwell card, each with half the memory, half the SMs and one NVENC, one NVDEC and one JPEG engine, and 2g.32gb for the whole card. The L40S has no MIG and shares memory and compute in time slices. By our sizing rule each 16 GB instance holds Qwen3-8B in FP8 with cache for about four 8K conversations with ECC off, or two with it on, so one eight-card server can run 16 separate small models with hardware isolation between them.
For compute VMs, NVIDIA AI Enterprise’s vGPU reference (release 8.2, updated 2 September 2026) lists L40S types from L40S-48C down to L40S-4C, twelve 4 GB VMs per card. The smallest compute type for the RTX PRO 4500 Server Edition is 8 GB, four per card, time-sliced as DC-8C or MIG-backed as DC-2-8C. In the 2U example that makes 48 compute VMs of 4 GB on four L40S and 32 of 8 GB on eight Server Edition cards.
For virtual desktops, NVIDIA’s vPC sizing guide gives the Server Edition 16 users per board with 2 GB profiles and marks 1 GB profiles “Not supported on Blackwell”; the L40S reaches 32 users per board with 1 GB profiles. The guide adds that its table “reflects maximum supported density rather than a recommended deployment point”. Four L40S and eight Server Edition cards therefore reach the same 128 users, and each Blackwell seat has twice the frame buffer.
NVIDIA’s licensing guide states that NVIDIA AI Enterprise “is licensed on a per-GPU basis” and that a licence “is required for every GPU installed on the server or workstation” that hosts its software, so an eight-card server that runs it needs eight licences where a four-card server needs four. Virtual desktops are licensed differently: Lenovo’s guide for the card lists vPC, vApps and RTX vWS licences per concurrent user (CCU).
We supply both cards with NVIDIA vGPU and AI Enterprise licences on one EU contract and invoice. Describe your seat count, profile sizes and hypervisor in the form below, and we size the cards per host.
Video engines, power and airflow
Each card has three NVENC and three NVDEC engines, so eight cards bring 24 of each and four bring 12. NVIDIA states that the Blackwell encoders “add new support for 4:2:2 H.264 and HEVC encoding”, and the L40S page lists AV1 encode and decode for its engines.
Our guide to GPU rack power and cooling plans about 160 CFM of airflow per kW at an 11 °C rise, which by our arithmetic is about 26 CFM for each Server Edition card and 56 CFM for each L40S, or 211 and 224 CFM per server before processors, memory and drives. NVIDIA’s product brief PB-12790-001_02 gives the Server Edition an ambient operating range of 10 to 45 °C and a power limit that can be set from 165 W down to 100 W. Both cards take one 16-pin power connector, so the eight-card server needs eight GPU power cables, and Lenovo’s card guide lists an auxiliary power cable for the SR650a V4.
L40S successor and Blackwell alternatives: what NVIDIA says
The L40S product page names no successor. NVIDIA’s RTX PRO 6000 Blackwell Server Edition page states that the card “delivers a massive leap in performance versus the previous-generation NVIDIA L40S”, without test conditions on the page. The RTX PRO 4500 Server Edition page compares the card with the L4 instead, claiming “over 5x the performance of the previous-generation L4 GPU”, also without conditions. For virtual PCs its vPC sizing guide calls the RTX PRO 4500 Server Edition “a strong modernization path for customers moving from earlier generations of data center GPUs”, naming Ada Lovelace (L-series) among them. Our comparison of the L40S and the RTX PRO 6000 Server Edition covers that pairing.
The RTX PRO 4500 Server Edition becomes an L40S alternative when the server design changes with it: more 32 GB cards per server, partitioned with MIG, at less than half the power per card. NVIDIA’s vGPU support list, updated 2 October 2026, shows both cards with “Full Support” and the status “Active”.
When to choose the L40S or the RTX PRO 4500 Server Edition
Choose the L40S when a model with its cache needs more than about 24 GiB on one card, such as a 32B model in FP8, when the work leans on FP32 arithmetic, or when the servers have dual-width slots qualified for 350 W cards.
Choose the RTX PRO 4500 Server Edition for many models up to 14B in FP8 or 32B in NVFP4, for embedding and reranking models next to an assistant, for departments that need hardware isolation through MIG, for VDI with 2 GB profiles or larger and for video work that needs 4:2:2 encoding. It suits racks limited by power per rack position and servers with eight single-width slots.
What we supply
We supply the NVIDIA L40S and the RTX PRO 4500 Blackwell Server Edition as cards for servers you run, after a compatibility check of the platform, power and cooling, or in AI servers built to order, on one EU contract and invoice with manufacturer warranty. For models that outgrow 32 or 48 GB per card, we supply the RTX PRO 6000 Server Edition and the H200 NVL. NVIDIA AI Enterprise and vGPU licences come on the same invoice as the hardware. The full range of cards is on our professional GPU page.
FAQ
L40S vs RTX PRO 4500 Server Edition: which is the right card for inference?
What is the successor of the NVIDIA L40S?
Is the RTX PRO 4500 Server Edition an L40S alternative on Blackwell?
How many RTX PRO 4500 Server Edition cards fit in a 2U server?
Does the RTX PRO 4500 Server Edition support MIG and vGPU?
Can the RTX PRO 4500 Server Edition run a 32B model?
Send us the models and their precision, the peak number of conversations or vGPU seats, and the server model or rack position the cards will go into. We reply within one business day with the card and the number per server, the licences the setup needs and a configuration and quote in writing.
Talk to an expertWe reply within one business day