BLOG · COMPARISON ·

L40S vs RTX PRO 4500 Server Edition: 48 GB at 350 W or 32 GB at 165 W per card

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • The L40S has 48 GB of GDDR6 at 864 GB/s on a 350 W dual-slot board; the RTX PRO 4500 Blackwell Server Edition has 32 GB of GDDR7 at 800 GB/s on a 165 W single-slot board, and both are passive
  • The Server Edition adds FP4 Tensor Cores, MIG with two 16 GB instances, PCIe 5.0 and 4:2:2 H.264 and HEVC encoding; the L40S has 73 per cent more CUDA cores and 91.6 against 51 TFLOPS of FP32
  • By our sizing rule a 32B model in FP8 fits one L40S with cache for about five 8K conversations with ECC on, eight with it off, and does not fit 32 GB, while Qwen3-14B in FP8 fits both cards, for about 35 and 13 conversations with ECC on
  • In a 2U layout of four double-wide or eight single-wide GPU slots, Lenovo’s SR650a V4 as an example, eight Server Edition cards give 256 GB of GPU memory at 1,320 W of board power, four L40S give 192 GB at 1,400 W
  • NVIDIA sets the RTX PRO 6000 Server Edition against the L40S and the RTX PRO 4500 Server Edition against the L4, and its vGPU list of 2 October 2026 shows the L40S and the RTX PRO 4500 Server Edition as fully supported

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

L40S vs RTX PRO 4500 Server Edition: the short answer

The NVIDIA L40S has 48 GB of GDDR6 at 864 GB/s on a dual-slot, 350 W board, and the RTX PRO 4500 Blackwell Server Edition has 32 GB of GDDR7 at 800 GB/s on a single-slot, 165 W board. Choose the L40S when one model with its cache needs more than about 24 GiB on a card, a 32B model in FP8 for example, or when you extend an existing L40S environment. Choose the RTX PRO 4500 Server Edition when the models fit in 32 GB and the server has single-width slots: it adds FP4 Tensor Cores, two 16 GB MIG instances, PCIe 5.0 and 4:2:2 video encoding, and a 2U server with eight single-width slots can take twice as many of them, where its maker lists the card for those slots.

L40S and RTX PRO 4500 Server Edition specifications

SPECIFICATIONL40SRTX PRO 4500 SE
ArchitectureAda LovelaceBlackwell
CUDA cores18,17610,496
GPU memory48 GB GDDR6 with ECC32 GB GDDR7 with ECC, 256-bit
Memory bandwidth864 GB/s800 GB/s
FP8 Tensor733 TFLOPS dense, 1,466 with sparsity811 TFLOPS, basis not stated
FP4 Tensornone1.6 PFLOPS, basis not stated
FP3291.6 TFLOPS51 TFLOPS
Maximum power350 W165 W
Boarddual slot, 4.4 by 10.5 in, passivesingle slot, full height, full length, passive
Host interfacePCIe 4.0 x16PCIe 5.0 x16
MIGnoup to 2 instances of 16 GB
vGPU, first release16.120.0
NVENC and NVDEC3 and 3, with AV13 and 3, with 4:2:2 H.264 and HEVC encoding
NVLinknono

NVIDIA L40S product page; NVIDIA RTX PRO 4500 Blackwell Server Edition product page, datasheet (5192801, April 2026) and product brief PB-12790-001_02 (20 April 2026); NVIDIA’s list of GPUs supported by vGPU, updated 2 October 2026.

The L40S is the larger chip, with 73 per cent more CUDA cores and 1.8 times the FP32 rate, but only 8 per cent more memory bandwidth. NVIDIA does not state whether the Server Edition’s 811 TFLOPS of FP8 include sparsity; on the L40S page, 733 is the dense figure and 1,466 the figure with sparsity. If NVIDIA’s Blackwell figure is a sparse one, the L40S has the higher FP8 rate per card. Per server the order reverses: eight Server Edition cards add up to 6,488 TFLOPS, against 2,932 dense or 5,864 sparse for four L40S, by our arithmetic.

Our RTX PRO 4500 Server Edition guide covers the card on its own, and our L40S benchmark review the published measurements of the older one.

Which models fit in 32 GB and in 48 GB

We size with the rule from our guide to how much VRAM an LLM needs: 90 per cent of the card’s memory, less about 3 GiB for activations and the runtime, holds the weights and the KV cache. Both cards have ECC memory, which on GDDR cards reduces the available memory by 6.25 per cent, according to NVIDIA’s CUDA C++ Best Practices Guide. With ECC enabled that leaves about 37.3 GiB for weights and cache on the L40S and 23.9 GiB on the RTX PRO 4500 Server Edition, with ECC off 40.0 and 25.7 GiB; nvidia-smi -q shows the ECC mode of the delivered card. The cache of one conversation follows from the model’s config.json, and with an FP8 cache at a full 8,192-token context it takes 0.5625 GiB for Qwen3-8B, 0.625 GiB for Qwen3-14B and 1 GiB for Qwen3-32B.

MODEL AND FORMATWEIGHTSL40SRTX PRO 4500 SE
Qwen3-8B, FP88.8 GiBabout 50 (55)about 26 (30)
Qwen3-14B, FP815.2 GiBabout 35 (39)about 13 (16)
Qwen3-32B, NVFP419.3 GiBno FP4 Tensor Coresabout 4 (6)
Qwen3-32B, FP832.0 GiBabout 5 (8)does not fit
Llama 3.3 70B, NVFP439.8 GiBno room for cachedoes not fit

Concurrent 8,192-token conversations per card with an FP8 KV cache, by our arithmetic, first with ECC enabled, in brackets with ECC off; card memory is our estimate of what the driver reports. Weights: NVIDIA’s FP8 and NVFP4 checkpoints and Qwen’s FP8 release of Qwen3-32B on Hugging Face, read on 10 October 2026; layers and KV heads from each config.json.

The dividing line is a 32B model in FP8, which fits the L40S with room for about five conversations (eight with ECC off) and does not fit the Server Edition. Qwen’s FP8 release scales its weights in blocks of 128 by 128, a format TensorRT-LLM’s support matrix does not list for Ada, as our L40S benchmark review notes, so confirm that your serving engine runs it on the L40S. In NVFP4 the same model fits the Blackwell card for about four conversations (six with ECC off). Ada has no FP4 Tensor Cores, so on the L40S vLLM runs NVFP4 only as weight-only compression, and other 4-bit models run as INT4 through AWQ or GPTQ. Models up to 14B in FP8 fit both, and the L40S holds more than twice the Qwen3-14B conversations per card. Speed for one user is close, because a dense model reads its weights once per token: for Qwen3-14B in FP8, with 16.3 GB of weights, the ceiling is about 53 tokens per second on the L40S and 49 on the Server Edition by our arithmetic.

Per-server totals in a 2U server with four or eight GPU slots

Lenovo Press describes the ThinkSystem SR650a V4, a 2U server, with “Support for up to 4x double-wide GPUs or 8x single-wide GPUs, installed in the front slots” (LP2128, updated 5 October 2026), and lists “4x NVIDIA L40S 350W GPUs” among its features. Lenovo’s guide to the RTX PRO 4500 Server Edition (LP2391) notes in its update of 28 July 2026 that “The SR650a V4 now supports the GPU”. Lenovo’s own eight-card example is “8x NVIDIA L4 single-wide GPUs”, and neither guide we could read states how many RTX PRO 4500 Server Edition cards the server takes. The table uses this layout as an example; the card count per server comes from the maker’s configurator.

PER 2U SERVER4 × L40S8 × RTX PRO 4500 SE
GPU memory192 GB256 GB
Board power, GPUs1,400 W1,320 W
Bandwidth, sum3,456 GB/s6,400 GB/s
FP32, sum366 TFLOPS408 TFLOPS
NVENC and NVDEC12 and 1224 and 24
MIG instancesnone16 of 16 GB
Qwen3-14B FP8, 8Kabout 140 (156)about 104 (128)
Qwen3-32B FP8, 8Kabout 20 (32)does not fit
vPC users, maximum128 with 1 GB profiles128 with 2 GB profiles

Example layout, not a configuration: slot counts from Lenovo Press LP2128; per-card figures from NVIDIA’s product pages; conversations from the table above, with ECC on and, in brackets, off; vPC users per board from NVIDIA’s vPC sizing guide, updated 19 August 2026. Sums are our arithmetic.

The eight-card server has a third more GPU memory at slightly lower board power, and nearly twice the summed bandwidth and twice the video engines. Each card holds less, so the Qwen3-14B count falls from about 140 to 104 with ECC on, and a 32B model in FP8 has to be split across two cards by tensor parallelism, over PCIe, since neither card has NVLink. With 8B models the order changes again, about 208 conversations on eight Server Edition cards against 200 on four L40S with ECC on.

We build AI servers to order with either card and check the rack, power and airflow before we quote. Tell us the models, the peak conversations and the server or rack position, and we reply with the card count per server.

MIG and vGPU on the L40S and the RTX PRO 4500 Server Edition

NVIDIA’s MIG user guide, updated 11 September 2026, lists up to two MIG 1g.16gb instances per RTX PRO 4500 Blackwell card, each with half the memory, half the SMs and one NVENC, one NVDEC and one JPEG engine, and 2g.32gb for the whole card. The L40S has no MIG and shares memory and compute in time slices. By our sizing rule each 16 GB instance holds Qwen3-8B in FP8 with cache for about four 8K conversations with ECC off, or two with it on, so one eight-card server can run 16 separate small models with hardware isolation between them.

For compute VMs, NVIDIA AI Enterprise’s vGPU reference (release 8.2, updated 2 September 2026) lists L40S types from L40S-48C down to L40S-4C, twelve 4 GB VMs per card. The smallest compute type for the RTX PRO 4500 Server Edition is 8 GB, four per card, time-sliced as DC-8C or MIG-backed as DC-2-8C. In the 2U example that makes 48 compute VMs of 4 GB on four L40S and 32 of 8 GB on eight Server Edition cards.

For virtual desktops, NVIDIA’s vPC sizing guide gives the Server Edition 16 users per board with 2 GB profiles and marks 1 GB profiles “Not supported on Blackwell”; the L40S reaches 32 users per board with 1 GB profiles. The guide adds that its table “reflects maximum supported density rather than a recommended deployment point”. Four L40S and eight Server Edition cards therefore reach the same 128 users, and each Blackwell seat has twice the frame buffer.

NVIDIA’s licensing guide states that NVIDIA AI Enterprise “is licensed on a per-GPU basis” and that a licence “is required for every GPU installed on the server or workstation” that hosts its software, so an eight-card server that runs it needs eight licences where a four-card server needs four. Virtual desktops are licensed differently: Lenovo’s guide for the card lists vPC, vApps and RTX vWS licences per concurrent user (CCU).

We supply both cards with NVIDIA vGPU and AI Enterprise licences on one EU contract and invoice. Describe your seat count, profile sizes and hypervisor in the form below, and we size the cards per host.

Video engines, power and airflow

Each card has three NVENC and three NVDEC engines, so eight cards bring 24 of each and four bring 12. NVIDIA states that the Blackwell encoders “add new support for 4:2:2 H.264 and HEVC encoding”, and the L40S page lists AV1 encode and decode for its engines.

Our guide to GPU rack power and cooling plans about 160 CFM of airflow per kW at an 11 °C rise, which by our arithmetic is about 26 CFM for each Server Edition card and 56 CFM for each L40S, or 211 and 224 CFM per server before processors, memory and drives. NVIDIA’s product brief PB-12790-001_02 gives the Server Edition an ambient operating range of 10 to 45 °C and a power limit that can be set from 165 W down to 100 W. Both cards take one 16-pin power connector, so the eight-card server needs eight GPU power cables, and Lenovo’s card guide lists an auxiliary power cable for the SR650a V4.

L40S successor and Blackwell alternatives: what NVIDIA says

The L40S product page names no successor. NVIDIA’s RTX PRO 6000 Blackwell Server Edition page states that the card “delivers a massive leap in performance versus the previous-generation NVIDIA L40S”, without test conditions on the page. The RTX PRO 4500 Server Edition page compares the card with the L4 instead, claiming “over 5x the performance of the previous-generation L4 GPU”, also without conditions. For virtual PCs its vPC sizing guide calls the RTX PRO 4500 Server Edition “a strong modernization path for customers moving from earlier generations of data center GPUs”, naming Ada Lovelace (L-series) among them. Our comparison of the L40S and the RTX PRO 6000 Server Edition covers that pairing.

The RTX PRO 4500 Server Edition becomes an L40S alternative when the server design changes with it: more 32 GB cards per server, partitioned with MIG, at less than half the power per card. NVIDIA’s vGPU support list, updated 2 October 2026, shows both cards with “Full Support” and the status “Active”.

When to choose the L40S or the RTX PRO 4500 Server Edition

Choose the L40S when a model with its cache needs more than about 24 GiB on one card, such as a 32B model in FP8, when the work leans on FP32 arithmetic, or when the servers have dual-width slots qualified for 350 W cards.

Choose the RTX PRO 4500 Server Edition for many models up to 14B in FP8 or 32B in NVFP4, for embedding and reranking models next to an assistant, for departments that need hardware isolation through MIG, for VDI with 2 GB profiles or larger and for video work that needs 4:2:2 encoding. It suits racks limited by power per rack position and servers with eight single-width slots.

What we supply

We supply the NVIDIA L40S and the RTX PRO 4500 Blackwell Server Edition as cards for servers you run, after a compatibility check of the platform, power and cooling, or in AI servers built to order, on one EU contract and invoice with manufacturer warranty. For models that outgrow 32 or 48 GB per card, we supply the RTX PRO 6000 Server Edition and the H200 NVL. NVIDIA AI Enterprise and vGPU licences come on the same invoice as the hardware. The full range of cards is on our professional GPU page.

FAQ

L40S vs RTX PRO 4500 Server Edition: which is the right card for inference?
The L40S has 48 GB at 864 GB/s and 350 W, the RTX PRO 4500 Server Edition 32 GB at 800 GB/s and 165 W, so speed for one user is close and memory decides. A 32B model in FP8 fits only the L40S, while models up to 14B in FP8 fit both and the Server Edition adds FP4 and MIG. In a 2U server with eight single-width slots, eight Server Edition cards give 256 GB against 192 GB for four L40S.
What is the successor of the NVIDIA L40S?
NVIDIA’s RTX PRO 6000 Blackwell Server Edition page calls the L40S the previous-generation card and compares the 96 GB RTX PRO 6000 Server Edition with it, while the RTX PRO 4500 Server Edition page compares that card with the L4. For virtual PCs, NVIDIA’s vPC sizing guide names the RTX PRO 4500 Server Edition as a modernisation path from Ada Lovelace (L-series) GPUs. The L40S product page names no successor, and NVIDIA’s vGPU list of 2 October 2026 shows the L40S as fully supported.
Is the RTX PRO 4500 Server Edition an L40S alternative on Blackwell?
It is when the models fit in 32 GB per card and the server takes single-width cards: Lenovo describes its 2U SR650a V4 with up to eight single-wide or four double-wide GPUs and lists the card for it, and the Server Edition draws 165 W against 350 W. It brings FP4 Tensor Cores, two 16 GB MIG instances and PCIe 5.0. It is not an alternative for a 32B model in FP8 on one card, which needs the 48 GB of the L40S or a 96 GB RTX PRO 6000 Server Edition.
How many RTX PRO 4500 Server Edition cards fit in a 2U server?
Lenovo Press describes the 2U ThinkSystem SR650a V4 with up to eight single-wide or four double-wide GPUs in its front slots, and Lenovo’s guide for the card records SR650a V4 support in its update of 28 July 2026, but neither guide we read states a maximum for this card. NVIDIA’s vPC sizing guide assumes eight boards per 2U server in its maximum-density table, the same count it gives for the dual-slot L40S, without naming a server. The count for a given server model comes from its maker’s configurator for this card.
Does the RTX PRO 4500 Server Edition support MIG and vGPU?
Yes. NVIDIA’s MIG user guide lists two 1g.16gb instances per card, and NVIDIA’s vGPU list shows the card from release 20.0 with full support. Compute types go down to 8 GB with four VMs per card, and NVIDIA’s vPC sizing guide gives 16 desktop users per card with 2 GB profiles.
Can the RTX PRO 4500 Server Edition run a 32B model?
In 4-bit it can: NVIDIA’s NVFP4 checkpoint of Qwen3-32B is 19.3 GiB and leaves room for about four 8,192-token conversations with an FP8 cache by our sizing rule with ECC on, or six with ECC off. In FP8 the model is 32.0 GiB and does not fit the 32 GB card. One L40S holds the FP8 version with cache for about five such conversations with ECC on, or eight with it off.

Send us the models and their precision, the peak number of conversations or vGPU seats, and the server model or rack position the cards will go into. We reply within one business day with the card and the number per server, the licences the setup needs and a configuration and quote in writing.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna