48 GB GPUs compared: RTX 6000 Ada, RTX 5880 Ada, RTX PRO 5000 and L40S for AI
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- All four 48 GB cards carry ECC memory; the RTX PRO 5000 Blackwell has GDDR7 at 1,344 GB/s and PCIe 5.0, the RTX 6000 Ada and RTX 5880 Ada GDDR6 at 960 GB/s and the L40S GDDR6 at 864 GB/s, all three on PCIe 4.0
- Only the RTX PRO 5000 has FP4 Tensor Cores and MIG (two 24 GB instances); the L40S is the only passive card, at 350 W, against 300 W for the RTX 6000 Ada and RTX PRO 5000 and 285 W for the RTX 5880 Ada
- NVIDIA lists vGPU support for the L40S, RTX 6000 Ada and RTX 5880 Ada; the only RTX PRO 5000 on its vGPU list is the 72 GB version
- By our sizing estimate, with ECC enabled, one 48 GB card holds a 32B model in FP8 with cache for about 5 conversations of 8K and a 14B model for about 35 (8 and 39 with ECC off); a 70B model in FP8 needs two cards
- For 4 to 8 cards in a rack server the passive L40S fits the chassis airflow; the three actively cooled cards suit workstations, where the RTX PRO 5000 brings the most bandwidth
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
48 GB GPUs compared: what differs besides the memory
Four NVIDIA cards we supply have 48 GB of memory: the RTX 6000 Ada Generation, the RTX 5880 Ada Generation, the RTX PRO 5000 Blackwell in its 48 GB version and the L40S. All four carry 48 GB with ECC, so they hold the same models, and they differ in what surrounds the memory. The RTX PRO 5000 has GDDR7 at 1,344 GB/s, PCIe 5.0, FP4 Tensor Cores and MIG with two instances. The three Ada cards have GDDR6 at 960 GB/s (RTX 6000 Ada and RTX 5880 Ada) or 864 GB/s (L40S), PCIe 4.0 and no MIG. The L40S is the only passive card of the four, built for server airflow at 350 W, while the other three have active coolers for workstations and draw 285 to 300 W.
For a 48 GB GPU running an LLM, memory bandwidth sets the ceiling on the speed of each conversation, and the memory left beside the weights sets how many conversations fit. The RTX PRO 5000 reads its memory 1.4 times as fast as the two Ada workstation cards and about 1.56 times as fast as the L40S. For virtual machines the order changes: NVIDIA lists vGPU support for the three Ada cards, while the only RTX PRO 5000 on its vGPU list is the 72 GB version.
48 GB GPU comparison: specifications side by side
| CARD | MEMORY, BANDWIDTH | TENSOR CORES | POWER, COOLING | MIG, VGPU |
|---|---|---|---|---|
| RTX 6000 Ada | 48 GB GDDR6, 960 GB/s | 4th generation, FP8 | 300 W, active | no MIG; vGPU from 15.2 |
| RTX 5880 Ada | 48 GB GDDR6, 960 GB/s | 4th generation, FP8 | 285 W, active | no MIG; vGPU from 17.0 |
| RTX PRO 5000, 48 GB | 48 GB GDDR7, 1,344 GB/s | 5th generation, adds FP4 | 300 W, active | 2 × 24 GB; no vGPU |
| L40S | 48 GB GDDR6, 864 GB/s | 4th generation, FP8 | 350 W, passive | no MIG; vGPU from 16.1 |
NVIDIA datasheets and product pages: RTX 6000 Ada (datasheet of February 2023), RTX 5880 Ada (datasheet 3082478, December 2023), RTX PRO 5000 Blackwell (datasheet of July 2026), L40S product page; NVIDIA MIG user guide (11 September 2026) and vGPU supported GPU list (2 October 2026), read on 10 October 2026. vGPU column: first supported vGPU software release.
The RTX 6000 Ada and the L40S both list 18,176 CUDA cores and 568 Tensor Cores, the RTX 5880 Ada 14,080 and 440. The RTX PRO 5000 also lists 14,080 CUDA cores, of the Blackwell generation. NVIDIA rates FP32 at 91.1 TFLOPS on the RTX 6000 Ada, 91.6 on the L40S, 69.3 on the RTX 5880 Ada and 65 on the RTX PRO 5000.
The Tensor Core peaks are published on different bases. NVIDIA rates the RTX 6000 Ada at 1,457 TFLOPS, the L40S at 1,466 and the RTX 5880 Ada at 1,108.4, each as FP8 with sparsity. For the RTX PRO 5000 it gives 2,064 AI TOPS, which the datasheet defines as FP4 with sparsity. Set beside the three FP8 figures, an FP4 peak overstates the difference between the generations.
The three Ada cards have four DisplayPort 1.4a outputs and the RTX PRO 5000 four DisplayPort 2.1b. All are dual-slot, 4.4 by 10.5 inches. The RTX 6000 Ada and RTX 5880 Ada datasheets state no NVLink, and NVIDIA lists none for the L40S and the RTX PRO 5000, so a model split over two cards exchanges its data over PCIe.
What a 48 GB GPU holds for LLMs
Our guide to how much VRAM an LLM needs explains the method; here we apply its rule to 48 GB. We use 90 per cent of the memory the driver reports, subtract 3 GiB for activations and the runtime, then the weights, and divide the rest by the KV cache of one conversation in FP8 at a full 8K context: 0.625 GiB for Qwen3-14B, 1 GiB for Qwen3-32B and 1.25 GiB for Llama 3.3 70B.
ECC changes the starting figure. NVIDIA’s CUDA C++ Best Practices Guide states that “On GPUs with GDDR memory with ECC enabled the available DRAM is reduced by 6.25% to allow for the storage of ECC bits.” With ECC off we estimate 47.8 GiB per card, in proportion to the 95.6 GiB the driver reports on a 96 GB card, as in our RTX PRO 5000 comparison; with ECC on, about 44.8 GiB. The table gives both, and nvidia-smi -q shows the ECC mode and the memory of the delivered card.
| MODEL AND PRECISION | WEIGHTS | ONE 48 GB CARD | TWO CARDS, SPLIT |
|---|---|---|---|
| Qwen3-14B, FP8 | 15.2 GiB | about 35 (39) | not needed |
| Qwen3-32B, FP8 | 32.0 GiB | about 5 (8) | about 42 (48) |
| Llama 3.3 70B, NVFP4 | 39.8 GiB | no room for one | about 27 (32) |
| Llama 3.3 70B, FP8 | 67.7 GiB | does not fit | about 5 (9) |
Our estimates. Conversations at a full 8K context in an FP8 KV cache, first with ECC enabled (44.8 GiB per card), in brackets with ECC off (47.8 GiB). Weights from Qwen’s FP8 checkpoints of Qwen3-14B and Qwen3-32B and NVIDIA’s Llama 3.3 70B FP8 and FP4 checkpoints on Hugging Face; two cards with tensor parallelism, cache pooled across both.
The weights are a fixed cost per copy of the model, so the count per card drops quickly as models grow. By the same rule a 96 GB RTX PRO 6000 holds about 51 conversations of Qwen3-32B in FP8, against 5 to 8 on one 48 GB card. The table counts memory only, not the formats each card computes; NVFP4 runs on the FP4 Tensor Cores of the RTX PRO 5000, and for the Ada cards you should check that your serving engine supports the 4-bit format of the checkpoint you plan to use.
Memory bandwidth, FP4 and measured results
For a dense model, each new token of one conversation reads all the weights once, so bandwidth caps the speed of a single stream when the model fits. At 1,344 GB/s against 960 and 864 GB/s, the RTX PRO 5000 has the higher ceiling on every model in the table above. With many conversations in one batch, the cards work closer to their compute limits, and the Tensor Core figures above matter more.
Puget Systems’ professional GPU roundup of 18 December 2025 ran the 48 GB RTX PRO 5000 and the RTX 6000 Ada in the same desktop system: in V-Ray GPU the RTX PRO 5000 “even beat the 6000 Ada by 11%”, and in Topaz Video AI it was 8 per cent ahead. The RTX 5880 Ada and the L40S were not in that test. For language models we found no published test that runs all four cards with the same engine, precision and concurrency. NVIDIA’s NIM tables for the L40S, with token rates at 1 to 250 concurrent requests, are in our article on published L40S benchmarks.
MIG and vGPU on 48 GB cards
MIG splits a card into instances with their own memory and compute. Of the four, only the RTX PRO 5000 supports it: NVIDIA’s MIG user guide lists the 48 GB card with a maximum of two instances, and the datasheet gives them as 24 GB each. Each half holds Qwen3-14B in FP8, but by our rule with cache for only about 3 to 5 conversations of 8K, so the split suits two small teams or a development and a test instance. The switch to compute mode turns off the card’s display outputs; our comparison of the RTX PRO 5000 with 48 GB and 72 GB lists the driver and tool versions it needs. The RTX 6000 Ada, RTX 5880 Ada and L40S have no MIG and share a card between virtual machines in time slices.
For vGPU, NVIDIA’s supported GPU list of 2 October 2026 shows the RTX 6000 Ada from release 15.2, the L40S from 16.1 and the RTX 5880 Ada from 17.0, each with full support. The 48 GB RTX PRO 5000 is not on the list; the 72 GB version is, from release 20.2. NVIDIA adds that “GPU support may also depend on the hypervisor software that you are using”, so check the support matrix of your vGPU release against your hypervisor. On the RTX 6000 Ada and RTX 5880 Ada the display ports are on by default, and NVIDIA’s pages say to turn them off when using vGPU software.
48 GB GPUs in an AI server: four to eight cards
A rack server cools its GPUs with the chassis fans, and passive data-centre cards such as the L40S are built for that airflow. NVIDIA rates the L40S at 350 W, so four cards draw 1.4 kW and eight draw 2.8 kW before processors and fans. An actively cooled workstation card belongs in a rack server only where the server maker lists that card for that chassis, so check the maker’s GPU support list before you plan four or more RTX 6000 Ada, RTX 5880 Ada or RTX PRO 5000 cards in one server.
Take as a worked example an internal assistant on Qwen3-14B in FP8 for 1,000 staff, with a peak of 130 conversations in flight at 8K. By the rule above one L40S with ECC enabled holds about 35, so a server with four L40S, one copy of the model per card, holds about 140. Two such servers hold about 280, and if one fails the other still covers the peak. A 70B model in FP8 on the same servers runs as one copy split over four cards with tensor parallelism over PCIe, with cache for about 65 conversations per copy by our arithmetic (73 with ECC off); for more users per card, the 96 GB and 141 GB cards in the line-up are the next step.
We build AI servers to order with four to eight L40S, and we check the rack, power and airflow before we quote. Send us the model, the peak number of conversations and the rack position through the form below.
48 GB GPUs in workstations
The three actively cooled cards are made for workstations. At 285 to 300 W per card, four cards draw up to 1.2 kW, which sets the size of the power supply and the airflow through the case.
For LLM development and local inference, the RTX PRO 5000 has the most bandwidth of the three, FP4 for NVFP4 checkpoints and MIG for two isolated users per card. The RTX 6000 Ada and RTX 5880 Ada fit a fleet that is standardised on the Ada generation, and workloads that need vGPU on a workstation-class card. Where an application vendor certifies specific cards, its list decides. The RTX 5880 Ada has the same memory and bandwidth as the RTX 6000 Ada with fewer cores and 15 W less board power; our RTX 5880 Ada and RTX 6000 Ada comparison covers the difference in detail, and the RTX 6000 Ada and L40S comparison the same core configuration as a workstation and a server card.
For a fleet of eight rendering or simulation workstations, the Puget figures above favour the RTX PRO 5000 in V-Ray GPU. For CAD seats, compare the size of your largest assemblies and scenes with 48 GB before you choose this class over a smaller card.
We supply all three workstation cards on one EU contract and invoice. Tell us how many workstations you are equipping and what runs on each.
Which 48 GB GPU for which machine
| MACHINE AND WORKLOAD | CARD | WHY |
|---|---|---|
| Rack server, LLM inference | L40S | passive cooling for chassis airflow, vGPU support |
| Rack server, VMs with vGPU | L40S | time-sliced vGPU from release 16.1 |
| Workstation, LLM development | RTX PRO 5000 | 1,344 GB/s, FP4, MIG with two 24 GB instances |
| Workstation, two users | RTX PRO 5000 | MIG, the only one of the four |
| Fleet standardised on Ada | RTX 6000 Ada or RTX 5880 Ada | same generation, vGPU on the list |
| Rendering in V-Ray GPU | RTX PRO 5000 | 11 per cent ahead of the RTX 6000 Ada in Puget’s test |
| 70B model in FP8 | two cards, or a larger card | does not fit one 48 GB card |
Our reading of NVIDIA’s datasheets, MIG guide and vGPU list (read on 10 October 2026) and of Puget Systems’ roundup of 18 December 2025.
What we supply
We supply the RTX 6000 Ada, the RTX 5880 Ada, the RTX PRO 5000 Blackwell and the L40S with manufacturer warranty, on one EU contract and invoice, as cards or in AI servers built to order. We build servers around the workload, assembled and burn-in tested, and the configuration and quote follow within one business day. For models that outgrow 48 GB, the line-up on our GPU page continues with the RTX PRO 5000 with 72 GB, the RTX PRO 6000 with 96 GB and the H200 NVL with 141 GB. NVIDIA AI Enterprise and vGPU licences come on the same invoice as the cards.
FAQ
Which 48 GB GPU should I choose for LLM inference?
RTX 6000 Ada vs RTX PRO 5000: which is faster?
What LLM fits on a GPU with 48 GB of VRAM?
Is a 48 GB GPU enough for a 70B model?
Which 48 GB GPUs support MIG and vGPU?
Can I put RTX 6000 Ada or RTX PRO 5000 cards in a GPU server?
Send us the model and its precision, the peak number of conversations, how many servers or workstations you plan and the power feed at the rack position. We reply within one business day with a configuration and a quote in writing, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day