DGX Station GB300 specs, memory and alternatives: RTX PRO 6000 and H200 NVL builds
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- The NVIDIA DGX Station is a deskside system with one GB300 superchip, a 72-core Grace CPU and a Blackwell Ultra GPU joined by NVLink-C2C at 900 GB/s, built by manufacturers such as Dell, HP, MSI and Supermicro
- As of October 2026, NVIDIA lists 252 GB of HBM3e at 7.1 TB/s on the GPU and 496 GB of LPDDR5X at 396 GB/s on the CPU, 748 GB in total; its March 2025 announcement gave 784 GB
- Weights and the KV cache of active conversations belong in the HBM3e; the LPDDR5X delivers about an eighteenth of its bandwidth, so a model that spills into it generates more slowly
- The alternatives we supply are a tower with up to four RTX PRO 6000 Max-Q (384 GB), a rack server with up to eight Server Edition cards (768 GB) and two or four H200 NVL joined by NVLink bridges (up to 564 GB of HBM3e)
- The DGX Station suits one large model for a few users; the PCIe builds suit many users, MIG isolation across 16 to 32 instances and models above 252 GB kept entirely in GPU memory
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
NVIDIA DGX Station GB300 specifications
The NVIDIA DGX Station is a deskside AI system built on one GB300 Grace Blackwell Ultra Desktop Superchip, which joins a 72-core Grace CPU and a Blackwell Ultra GPU over NVLink-C2C at 900 GB/s. As of October 2026, NVIDIA’s product page and datasheet list 748 GB of coherent memory in total. Of that, 252 GB is HBM3e at 7.1 TB/s on the GPU and 496 GB is LPDDR5X at 396 GB/s on the CPU. NVIDIA’s page lists Dell, Exxact, HP, MSI and Supermicro among the manufacturers that build it. We do not supply it; the alternatives below are the RTX PRO 6000 and H200 NVL builds we supply.
NVIDIA quotes 20 PFLOPS of FP4 and 10 PFLOPS of FP8, both with sparsity under the datasheet footnote that covers every Tensor Core figure not marked otherwise. The GPU splits into as many as seven MIG instances. A ConnectX-8 SuperNIC provides up to 800 Gb/s of Ethernet over two 400 Gb/s ports, and NVIDIA rates total system power at 1,600 W. NVIDIA’s product page and datasheet list Ubuntu with NVIDIA AI Developer Tools as the supported operating system; the development guide says it is based on Ubuntu 24.04 and that the Grace cores use the Arm v9.0 instruction set.
Supermicro’s version is a 40 kg tower with closed-loop liquid cooling and one 1,600 W power supply, and it takes 6U when mounted in a rack. The smaller GB10 system is covered in our comparison of the DGX Station and the DGX Spark.
DGX Station memory: 252 GB HBM3e and 496 GB LPDDR5X
NVIDIA’s announcement of 18 March 2025 gave the DGX Station “784GB of coherent memory space”. The current product page, the datasheet dated August 2026 and the DGX Station development guide, updated on 22 July 2026, give 748 GB. ServeTheHome reported on 20 March 2026 that the 2025 specification had 288 GB of HBM3e at 8 TB/s; with the 496 GB of LPDDR5X, that makes 784 GB by our arithmetic. Supermicro’s product page matches the current figures. A review or offer quoting 784 GB repeats the 2025 figure.
Only the 252 GB of HBM3e is GPU memory. The 496 GB of LPDDR5X is attached to the Grace CPU at 396 GB/s, about an eighteenth of the HBM3e bandwidth. The GPU reaches it over NVLink-C2C, which the development guide says “facilitates coherent access to CPU and GPU memory”. NVIDIA’s technical blog of 5 September 2025 describes this on Grace systems as “a single unified memory address space shared by both the CPU and the GPU”, so software can allocate more than the GPU holds without copying data by hand.
HBM3e and LPDDR5X in LLM inference
Generating each token reads every active weight once, so the speed one user sees follows memory bandwidth, as our comparison of the RTX PRO 6000 and H200 NVL for LLM inference works through. On the DGX Station, the weights and the KV cache of active conversations belong in the 252 GB of HBM3e. Weights placed in LPDDR5X are read at no more than 396 GB/s, so reading them takes about 18 times as long as from HBM3e. NVIDIA’s pages give no token rates for a model split between the two tiers, and its technical blog lists KV cache offload among the workloads that benefit from the shared memory.
Qwen3-235B-A22B in NVFP4, about 134 GB of weights, fits in the HBM3e with room for its cache, and the Blackwell Ultra GPU runs FP4 natively. The same model in FP8, at 236 GB, would leave almost no HBM3e for the cache. NVIDIA’s development guide says the 748 GB enable work with models of up to 1T parameters, and at that size most weights sit in the slower tier. Kimi K2 Thinking, a 1T-parameter model published in INT4, has 594 GB of weights, more than twice the HBM3e. Per-model needs are in our guide to LLM hardware requirements by model.
DGX Station vs RTX PRO 6000 and H200 NVL builds
The table compares the DGX Station with the three alternatives we supply, each at its largest size, which spread memory over several PCIe cards.
| PER SYSTEM | DGX STATION GB300 | 4× RTX PRO MAX-Q | 8× RTX PRO SERVER | 4× H200 NVL |
|---|---|---|---|---|
| GPU memory | 252 GB HBM3e, plus 496 GB LPDDR5X on the CPU | 384 GB GDDR7 | 768 GB GDDR7 | 564 GB HBM3e |
| Bandwidth per GPU | 7.1 TB/s | 1,792 GB/s | 1,597 GB/s | 4.8 TB/s |
| FP4 and FP8 | both; FP4 20 PFLOPS, sparse | both; FP4 3,511 TOPS per card, sparse | both; FP4 4 PFLOPS per card | no FP4; FP8 3,341 TFLOPS per card, sparse |
| GPU interconnect | one GPU; NVLink-C2C to the CPU, 900 GB/s | PCIe Gen5 only | PCIe Gen5 only | NVLink bridge, 900 GB/s per GPU |
| MIG | up to 7 instances | up to 16 × 24 GB | up to 32 × 24 GB | up to 28 × 16.5 GB |
| Power | 1,600 W, whole system | 1,200 W cards only; 1.8 to 2 kW at the wall | up to 4.8 kW, cards only | up to 2.4 kW, cards only |
| Form factor | deskside system | tower workstation | rack server | rack server |
NVIDIA product pages, DGX Station datasheet (August 2026) and MIG user guide, read on 9 October 2026; sparse means with sparsity; the Server Edition page gives no sparsity note; the tower’s wall power is our estimate.
RTX PRO 6000 Max-Q workstation as a DGX Station alternative
The deskside alternative is a tower with two to four RTX PRO 6000 Blackwell Max-Q cards, each with 96 GB of GDDR7 at 1,792 GB/s and a 300 W limit. NVIDIA states that the Max-Q “enables up to four GPUs in a single system”, 384 GB in total. That is more GPU memory than the DGX Station has in HBM3e, at about a quarter of its bandwidth per GPU. No RTX PRO 6000 edition has NVLink, so a model larger than 96 GB is split across the cards over PCIe Gen5. MIG divides each card into up to four isolated 24 GB instances, sixteen in the tower; vGPU is not available on the Max-Q.
A tower with four cards draws about 1.8 to 2 kW at the wall under full load by our estimate, against the 1,600 W NVIDIA states for the DGX Station. Suitable towers, PCIe lanes, power supplies and heat are covered in our article on four RTX PRO 6000 Max-Q in one workstation.
Rack server with RTX PRO 6000 Server Edition cards
For a server room, the RTX PRO 6000 Server Edition goes into rack servers with four to eight cards. It has the same 96 GB of GDDR7, at 1,597 GB/s instead of 1,792 GB/s, and a configurable power limit of up to 600 W. Eight cards hold 768 GB, three times the HBM3e of the DGX Station, as eight separate memories joined over PCIe. MIG gives up to four 24 GB instances per card, 32 in eight cards, and the Server Edition is the only RTX PRO 6000 edition with vGPU for virtual machines.
For many users, the server keeps one copy of a model per card behind a load balancer. By our sizing rule, one card holds Llama 3.3 70B in FP8 with an FP8 cache for about twelve concurrent 8K conversations, so eight copies hold about eight times as many. Eight 600 W cards draw 4.8 kW before processors and fans, so plan three-phase feeds at the rack.
DGX Station vs H200 NVL servers with NVLink bridges
The H200 NVL is a dual-slot, air-cooled PCIe card with 141 GB of HBM3e at 4.8 TB/s and a configurable power limit of up to 600 W. Two- or four-way NVLink bridges join the cards at 900 GB/s per GPU, against 128 GB/s over PCIe Gen5. Two bridged cards hold 282 GB of HBM3e, 30 GB more than the GPU of the DGX Station, and four hold 564 GB. Hopper has FP8 but no FP4 arithmetic, and vLLM runs NVFP4 checkpoints on it as weight-only, so size it against FP8 or INT4 weights. Each card splits into up to seven MIG instances of 16.5 GB and includes a five-year NVIDIA AI Enterprise subscription.
Four bridged H200 NVL keep larger models entirely in HBM3e. GLM-4.5 in FP8, with 361 GB of weights, and Llama 4 Maverick in FP8, with 417 GB, fit in 564 GB, while the DGX Station would hold more than 100 GB of either in LPDDR5X. Where an NVFP4 checkpoint is published, it is much smaller, 134 GB against 236 GB in FP8 for Qwen3-235B, and the DGX Station runs it natively, unlike the H200 NVL. The DGX Station also has more bandwidth on its one GPU, 7.1 TB/s against 4.8 TB/s. Card counts for models of 235B to 1T parameters are in our guide to how many H200 NVL cards large models need.
We supply the RTX PRO 6000 and the H200 NVL with its NVLink bridges, as cards or in servers built to order. Tell us the models and the number of users, and we propose an RTX PRO 6000 or H200 NVL configuration.
Where the DGX Station is the stronger choice
The DGX Station is the stronger choice when one large model runs for one developer or a few users. A model too large for one 141 GB H200 NVL, but small enough with its cache for 252 GB, runs on one GPU at 7.1 TB/s without being split. Two bridged H200 NVL also hold it and, with tensor parallelism, read their halves of the weights at 9.6 TB/s combined by our arithmetic, at the cost of bridge traffic per layer. The LPDDR5X adds 496 GB of coherent memory for models or caches that do not fit, at lower speed. A PCIe card reaches host memory only over PCIe Gen5, at a seventh of the NVLink-C2C bandwidth by NVIDIA’s figures.
The ConnectX-8 provides up to 800 Gb/s of networking without an added card. NVIDIA’s datasheet says developers can work on models locally and deploy them to the cloud or data centre “using the same tools, libraries, frameworks, and pretrained models”. At 1,600 W of system power, the DGX Station also suits an office without a server room, on a circuit planned for it.
Where RTX PRO 6000 and H200 NVL builds fit better
The builds we supply are the stronger choice when many people use the system at once. Every active conversation keeps its own KV cache, so the memory a service needs grows with its users. Eight Server Edition cards provide 768 GB and four H200 NVL 564 GB, against 252 GB of HBM3e on the DGX Station. MIG isolation scales with the card count, from up to seven instances on the DGX Station to sixteen in a four-card tower and 32 in an eight-card server.
Each GPU in these builds is a standard PCIe card under manufacturer warranty, replaceable on its own. Our builds use AMD EPYC, Intel Xeon or Threadripper PRO processors, so x86 software and monitoring agents run unchanged; on the Arm-based Grace CPU of the DGX Station they need Arm builds. For fine-tuning as a shared service, H200 NVL cards combine HBM3e with NVLink at 900 GB/s between the cards.
| NEED | BETTER FIT | WHY |
|---|---|---|
| One large model, few users | DGX Station | up to 252 GB in HBM3e on one GPU |
| Weights above 252 GB | 4 to 8 H200 NVL or 8× RTX PRO 6000 Server | weights stay in GPU memory |
| Many users on 70B to 120B | 8× RTX PRO 6000 Server or 2 to 4 H200 NVL | a model copy per card, more room for KV cache |
| Fine-tuning, one developer | DGX Station | 252 GB on one GPU, no split across cards |
| Fine-tuning, shared node | 2 or 4 H200 NVL with NVLink | NVLink between cards, AI Enterprise included |
| Team sharing GPUs with MIG | 4× Max-Q or 8× Server Edition | 16 or 32 instances of 24 GB, against up to 7 |
| Office, no server room | DGX Station or 4× Max-Q tower | 1,600 W, or 1.8 to 2 kW at the wall |
Our reading of the NVIDIA specifications in the first table.
We build RTX PRO 6000 workstations and RTX PRO 6000 or H200 NVL servers to order and check the rack, power and airflow before we quote. Write to us with the models, their precision and the number of concurrent users.
What we supply
We supply the RTX PRO 6000 Blackwell in its Workstation, Max-Q and Server editions and the H200 NVL with its two-way and four-way NVLink bridges, as professional NVIDIA GPUs for systems you already run or in AI servers built to order, which we assemble and burn-in test. All of it comes with manufacturer warranty on one EU contract and invoice, together with NVIDIA AI Enterprise and vGPU licences. The DGX Station is built by the manufacturers NVIDIA lists, and we do not supply it. Models, RAG and MLOps on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What is the NVIDIA DGX Station GB300?
How much memory does the DGX Station have?
Can the DGX Station GPU use all 748 GB for an LLM?
DGX Station vs RTX PRO 6000: which suits LLM inference?
DGX Station vs H200 NVL: how do they compare?
What are the alternatives to the NVIDIA DGX Station?
Send us the models you plan to run, their precision, the number of concurrent users and whether the system will stand in an office or a server room. We reply within one business day with an RTX PRO 6000 or H200 NVL configuration and a written quote.
Talk to an expertWe reply within one business day