BLOG · COMPARISON ·

DGX Station GB300 specs, memory and alternatives: RTX PRO 6000 and H200 NVL builds

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • The NVIDIA DGX Station is a deskside system with one GB300 superchip, a 72-core Grace CPU and a Blackwell Ultra GPU joined by NVLink-C2C at 900 GB/s, built by manufacturers such as Dell, HP, MSI and Supermicro
  • As of October 2026, NVIDIA lists 252 GB of HBM3e at 7.1 TB/s on the GPU and 496 GB of LPDDR5X at 396 GB/s on the CPU, 748 GB in total; its March 2025 announcement gave 784 GB
  • Weights and the KV cache of active conversations belong in the HBM3e; the LPDDR5X delivers about an eighteenth of its bandwidth, so a model that spills into it generates more slowly
  • The alternatives we supply are a tower with up to four RTX PRO 6000 Max-Q (384 GB), a rack server with up to eight Server Edition cards (768 GB) and two or four H200 NVL joined by NVLink bridges (up to 564 GB of HBM3e)
  • The DGX Station suits one large model for a few users; the PCIe builds suit many users, MIG isolation across 16 to 32 instances and models above 252 GB kept entirely in GPU memory

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

NVIDIA DGX Station GB300 specifications

The NVIDIA DGX Station is a deskside AI system built on one GB300 Grace Blackwell Ultra Desktop Superchip, which joins a 72-core Grace CPU and a Blackwell Ultra GPU over NVLink-C2C at 900 GB/s. As of October 2026, NVIDIA’s product page and datasheet list 748 GB of coherent memory in total. Of that, 252 GB is HBM3e at 7.1 TB/s on the GPU and 496 GB is LPDDR5X at 396 GB/s on the CPU. NVIDIA’s page lists Dell, Exxact, HP, MSI and Supermicro among the manufacturers that build it. We do not supply it; the alternatives below are the RTX PRO 6000 and H200 NVL builds we supply.

NVIDIA quotes 20 PFLOPS of FP4 and 10 PFLOPS of FP8, both with sparsity under the datasheet footnote that covers every Tensor Core figure not marked otherwise. The GPU splits into as many as seven MIG instances. A ConnectX-8 SuperNIC provides up to 800 Gb/s of Ethernet over two 400 Gb/s ports, and NVIDIA rates total system power at 1,600 W. NVIDIA’s product page and datasheet list Ubuntu with NVIDIA AI Developer Tools as the supported operating system; the development guide says it is based on Ubuntu 24.04 and that the Grace cores use the Arm v9.0 instruction set.

Supermicro’s version is a 40 kg tower with closed-loop liquid cooling and one 1,600 W power supply, and it takes 6U when mounted in a rack. The smaller GB10 system is covered in our comparison of the DGX Station and the DGX Spark.

DGX Station memory: 252 GB HBM3e and 496 GB LPDDR5X

NVIDIA’s announcement of 18 March 2025 gave the DGX Station “784GB of coherent memory space”. The current product page, the datasheet dated August 2026 and the DGX Station development guide, updated on 22 July 2026, give 748 GB. ServeTheHome reported on 20 March 2026 that the 2025 specification had 288 GB of HBM3e at 8 TB/s; with the 496 GB of LPDDR5X, that makes 784 GB by our arithmetic. Supermicro’s product page matches the current figures. A review or offer quoting 784 GB repeats the 2025 figure.

Only the 252 GB of HBM3e is GPU memory. The 496 GB of LPDDR5X is attached to the Grace CPU at 396 GB/s, about an eighteenth of the HBM3e bandwidth. The GPU reaches it over NVLink-C2C, which the development guide says “facilitates coherent access to CPU and GPU memory”. NVIDIA’s technical blog of 5 September 2025 describes this on Grace systems as “a single unified memory address space shared by both the CPU and the GPU”, so software can allocate more than the GPU holds without copying data by hand.

HBM3e and LPDDR5X in LLM inference

Generating each token reads every active weight once, so the speed one user sees follows memory bandwidth, as our comparison of the RTX PRO 6000 and H200 NVL for LLM inference works through. On the DGX Station, the weights and the KV cache of active conversations belong in the 252 GB of HBM3e. Weights placed in LPDDR5X are read at no more than 396 GB/s, so reading them takes about 18 times as long as from HBM3e. NVIDIA’s pages give no token rates for a model split between the two tiers, and its technical blog lists KV cache offload among the workloads that benefit from the shared memory.

Qwen3-235B-A22B in NVFP4, about 134 GB of weights, fits in the HBM3e with room for its cache, and the Blackwell Ultra GPU runs FP4 natively. The same model in FP8, at 236 GB, would leave almost no HBM3e for the cache. NVIDIA’s development guide says the 748 GB enable work with models of up to 1T parameters, and at that size most weights sit in the slower tier. Kimi K2 Thinking, a 1T-parameter model published in INT4, has 594 GB of weights, more than twice the HBM3e. Per-model needs are in our guide to LLM hardware requirements by model.

DGX Station vs RTX PRO 6000 and H200 NVL builds

The table compares the DGX Station with the three alternatives we supply, each at its largest size, which spread memory over several PCIe cards.

PER SYSTEMDGX STATION GB3004× RTX PRO MAX-Q8× RTX PRO SERVER4× H200 NVL
GPU memory252 GB HBM3e, plus 496 GB LPDDR5X on the CPU384 GB GDDR7768 GB GDDR7564 GB HBM3e
Bandwidth per GPU7.1 TB/s1,792 GB/s1,597 GB/s4.8 TB/s
FP4 and FP8both; FP4 20 PFLOPS, sparseboth; FP4 3,511 TOPS per card, sparseboth; FP4 4 PFLOPS per cardno FP4; FP8 3,341 TFLOPS per card, sparse
GPU interconnectone GPU; NVLink-C2C to the CPU, 900 GB/sPCIe Gen5 onlyPCIe Gen5 onlyNVLink bridge, 900 GB/s per GPU
MIGup to 7 instancesup to 16 × 24 GBup to 32 × 24 GBup to 28 × 16.5 GB
Power1,600 W, whole system1,200 W cards only; 1.8 to 2 kW at the wallup to 4.8 kW, cards onlyup to 2.4 kW, cards only
Form factordeskside systemtower workstationrack serverrack server

NVIDIA product pages, DGX Station datasheet (August 2026) and MIG user guide, read on 9 October 2026; sparse means with sparsity; the Server Edition page gives no sparsity note; the tower’s wall power is our estimate.

RTX PRO 6000 Max-Q workstation as a DGX Station alternative

The deskside alternative is a tower with two to four RTX PRO 6000 Blackwell Max-Q cards, each with 96 GB of GDDR7 at 1,792 GB/s and a 300 W limit. NVIDIA states that the Max-Q “enables up to four GPUs in a single system”, 384 GB in total. That is more GPU memory than the DGX Station has in HBM3e, at about a quarter of its bandwidth per GPU. No RTX PRO 6000 edition has NVLink, so a model larger than 96 GB is split across the cards over PCIe Gen5. MIG divides each card into up to four isolated 24 GB instances, sixteen in the tower; vGPU is not available on the Max-Q.

A tower with four cards draws about 1.8 to 2 kW at the wall under full load by our estimate, against the 1,600 W NVIDIA states for the DGX Station. Suitable towers, PCIe lanes, power supplies and heat are covered in our article on four RTX PRO 6000 Max-Q in one workstation.

Rack server with RTX PRO 6000 Server Edition cards

For a server room, the RTX PRO 6000 Server Edition goes into rack servers with four to eight cards. It has the same 96 GB of GDDR7, at 1,597 GB/s instead of 1,792 GB/s, and a configurable power limit of up to 600 W. Eight cards hold 768 GB, three times the HBM3e of the DGX Station, as eight separate memories joined over PCIe. MIG gives up to four 24 GB instances per card, 32 in eight cards, and the Server Edition is the only RTX PRO 6000 edition with vGPU for virtual machines.

For many users, the server keeps one copy of a model per card behind a load balancer. By our sizing rule, one card holds Llama 3.3 70B in FP8 with an FP8 cache for about twelve concurrent 8K conversations, so eight copies hold about eight times as many. Eight 600 W cards draw 4.8 kW before processors and fans, so plan three-phase feeds at the rack.

DGX Station vs H200 NVL servers with NVLink bridges

The H200 NVL is a dual-slot, air-cooled PCIe card with 141 GB of HBM3e at 4.8 TB/s and a configurable power limit of up to 600 W. Two- or four-way NVLink bridges join the cards at 900 GB/s per GPU, against 128 GB/s over PCIe Gen5. Two bridged cards hold 282 GB of HBM3e, 30 GB more than the GPU of the DGX Station, and four hold 564 GB. Hopper has FP8 but no FP4 arithmetic, and vLLM runs NVFP4 checkpoints on it as weight-only, so size it against FP8 or INT4 weights. Each card splits into up to seven MIG instances of 16.5 GB and includes a five-year NVIDIA AI Enterprise subscription.

Four bridged H200 NVL keep larger models entirely in HBM3e. GLM-4.5 in FP8, with 361 GB of weights, and Llama 4 Maverick in FP8, with 417 GB, fit in 564 GB, while the DGX Station would hold more than 100 GB of either in LPDDR5X. Where an NVFP4 checkpoint is published, it is much smaller, 134 GB against 236 GB in FP8 for Qwen3-235B, and the DGX Station runs it natively, unlike the H200 NVL. The DGX Station also has more bandwidth on its one GPU, 7.1 TB/s against 4.8 TB/s. Card counts for models of 235B to 1T parameters are in our guide to how many H200 NVL cards large models need.

We supply the RTX PRO 6000 and the H200 NVL with its NVLink bridges, as cards or in servers built to order. Tell us the models and the number of users, and we propose an RTX PRO 6000 or H200 NVL configuration.

Where the DGX Station is the stronger choice

The DGX Station is the stronger choice when one large model runs for one developer or a few users. A model too large for one 141 GB H200 NVL, but small enough with its cache for 252 GB, runs on one GPU at 7.1 TB/s without being split. Two bridged H200 NVL also hold it and, with tensor parallelism, read their halves of the weights at 9.6 TB/s combined by our arithmetic, at the cost of bridge traffic per layer. The LPDDR5X adds 496 GB of coherent memory for models or caches that do not fit, at lower speed. A PCIe card reaches host memory only over PCIe Gen5, at a seventh of the NVLink-C2C bandwidth by NVIDIA’s figures.

The ConnectX-8 provides up to 800 Gb/s of networking without an added card. NVIDIA’s datasheet says developers can work on models locally and deploy them to the cloud or data centre “using the same tools, libraries, frameworks, and pretrained models”. At 1,600 W of system power, the DGX Station also suits an office without a server room, on a circuit planned for it.

Where RTX PRO 6000 and H200 NVL builds fit better

The builds we supply are the stronger choice when many people use the system at once. Every active conversation keeps its own KV cache, so the memory a service needs grows with its users. Eight Server Edition cards provide 768 GB and four H200 NVL 564 GB, against 252 GB of HBM3e on the DGX Station. MIG isolation scales with the card count, from up to seven instances on the DGX Station to sixteen in a four-card tower and 32 in an eight-card server.

Each GPU in these builds is a standard PCIe card under manufacturer warranty, replaceable on its own. Our builds use AMD EPYC, Intel Xeon or Threadripper PRO processors, so x86 software and monitoring agents run unchanged; on the Arm-based Grace CPU of the DGX Station they need Arm builds. For fine-tuning as a shared service, H200 NVL cards combine HBM3e with NVLink at 900 GB/s between the cards.

NEEDBETTER FITWHY
One large model, few usersDGX Stationup to 252 GB in HBM3e on one GPU
Weights above 252 GB4 to 8 H200 NVL or 8× RTX PRO 6000 Serverweights stay in GPU memory
Many users on 70B to 120B8× RTX PRO 6000 Server or 2 to 4 H200 NVLa model copy per card, more room for KV cache
Fine-tuning, one developerDGX Station252 GB on one GPU, no split across cards
Fine-tuning, shared node2 or 4 H200 NVL with NVLinkNVLink between cards, AI Enterprise included
Team sharing GPUs with MIG4× Max-Q or 8× Server Edition16 or 32 instances of 24 GB, against up to 7
Office, no server roomDGX Station or 4× Max-Q tower1,600 W, or 1.8 to 2 kW at the wall

Our reading of the NVIDIA specifications in the first table.

We build RTX PRO 6000 workstations and RTX PRO 6000 or H200 NVL servers to order and check the rack, power and airflow before we quote. Write to us with the models, their precision and the number of concurrent users.

What we supply

We supply the RTX PRO 6000 Blackwell in its Workstation, Max-Q and Server editions and the H200 NVL with its two-way and four-way NVLink bridges, as professional NVIDIA GPUs for systems you already run or in AI servers built to order, which we assemble and burn-in test. All of it comes with manufacturer warranty on one EU contract and invoice, together with NVIDIA AI Enterprise and vGPU licences. The DGX Station is built by the manufacturers NVIDIA lists, and we do not supply it. Models, RAG and MLOps on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What is the NVIDIA DGX Station GB300?
It is a deskside AI system built on NVIDIA’s GB300 Grace Blackwell Ultra Desktop Superchip, a 72-core Grace CPU and a Blackwell Ultra GPU joined by NVLink-C2C at 900 GB/s. NVIDIA quotes 20 PFLOPS of FP4 with sparsity and up to 800 Gb/s of networking through a ConnectX-8 SuperNIC. Manufacturers such as Dell, HP, MSI and Supermicro build their own versions.
How much memory does the DGX Station have?
NVIDIA’s current product page and its August 2026 datasheet give 748 GB of coherent memory, made up of 252 GB of HBM3e at 7.1 TB/s on the GPU and 496 GB of LPDDR5X at 396 GB/s on the CPU. NVIDIA’s announcement of March 2025 gave 784 GB, and pages quoting that figure use the 2025 specification. Only the 252 GB of HBM3e is GPU memory at full bandwidth.
Can the DGX Station GPU use all 748 GB for an LLM?
The GPU reaches the LPDDR5X of the CPU coherently over NVLink-C2C, so software can place weights or KV cache there. That memory delivers 396 GB/s, about an eighteenth of the 7.1 TB/s of the HBM3e, so reading data from it takes about 18 times as long. Plan the weights and the cache of active conversations in the 252 GB of HBM3e and treat the LPDDR5X as overflow.
DGX Station vs RTX PRO 6000: which suits LLM inference?
The DGX Station has one GPU with 252 GB of HBM3e at 7.1 TB/s, so a model up to that size runs without being split across cards. Four RTX PRO 6000 Max-Q in a tower hold 384 GB and eight Server Edition cards in a rack server 768 GB, at 1,792 or 1,597 GB/s per card and with up to four MIG instances per card. The DGX Station suits one large model for a few users, while the RTX PRO 6000 builds suit many users and isolated instances.
DGX Station vs H200 NVL: how do they compare?
One H200 NVL has 141 GB of HBM3e at 4.8 TB/s, and NVLink bridges join two or four cards at 900 GB/s per GPU for 282 or 564 GB. The DGX Station has more bandwidth on its one GPU and FP4 arithmetic, which Hopper lacks, while four bridged H200 NVL keep models such as GLM-4.5 or Llama 4 Maverick in FP8 entirely in HBM3e. Each H200 NVL also includes a five-year NVIDIA AI Enterprise subscription.
What are the alternatives to the NVIDIA DGX Station?
The alternatives are workstations and servers with PCIe cards, such as a tower with up to four RTX PRO 6000 Max-Q (384 GB), a rack server with four to eight Server Edition cards (up to 768 GB) or a rack server with two or four H200 NVL joined by NVLink bridges (up to 564 GB of HBM3e). They spread memory over several GPUs, which suits multi-user serving and MIG isolation, while the DGX Station keeps up to 252 GB on one GPU. Eurokommerz supplies these builds; the DGX Station is built by the manufacturers NVIDIA lists.

Send us the models you plan to run, their precision, the number of concurrent users and whether the system will stand in an office or a server room. We reply within one business day with an RTX PRO 6000 or H200 NVL configuration and a written quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna