DGX Spark vs H200 NVL: a development system and a production GPU server compared
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- DGX Spark (128 GB) has unified LPDDR5x memory at 273 GB/s and FP4 Tensor Cores in a desktop system; one H200 NVL has 141 GB of HBM3e at 4.8 TB/s, NVLink bridges for 2 or 4 cards and MIG, in a server with up to 8 cards
- In a company the two usually belong to one plan, with developers building and testing on 2 to 4 DGX Spark and 500 to 2,000 employees using the service on two H200 NVL servers, so that one server can fail
- NVIDIA’s vLLM container is multi-architecture, vLLM serves the same OpenAI-compatible API on both, and NIM lists gpt-oss-20b, Llama 3.1 8B and Nemotron 3 Nano as verified on both the GB10 and the H200 NVL, and gpt-oss-120b on the H200 NVL but not the GB10
- On the server the host changes from Arm to x86 and the compute capability from 12.1 to 9.0, so images and custom CUDA kernels need builds for both, and NVFP4 checkpoints usually give way to FP8, since NVIDIA lists no FP4 Tensor Core rate for the H200 NVL
- By our memory estimate gpt-oss-120b at 32K leaves room for about 30 conversations on one DGX Spark (128 GB) and about 55 on one H200 NVL; the 17.6 times higher memory bandwidth of the H200 NVL raises the upper bound on each user’s generation speed by the same factor
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
DGX Spark vs H200 NVL: development system and production GPU
DGX Spark and the H200 NVL do different jobs. DGX Spark is a desktop system with 128 GB of unified LPDDR5x memory at 273 GB/s and FP4 Tensor Cores, built for developers who prototype, fine-tune and test models. One H200 NVL is a PCIe card with 141 GB of HBM3e at 4.8 TB/s, NVLink bridges for two or four cards and MIG, and it goes into a server that serves a model to many users at once. All figures are as NVIDIA states them on its product pages, read in October 2026.
For a company with 200 to 2,000 staff, the two usually belong to one plan. The people who build the assistant, the RAG pipeline or the agents work on 2 to 4 DGX Spark, and the service that 500 to 2,000 employees use runs on H200 NVL servers. Which hardware a first project should start on is covered in our comparison of DGX Spark, RTX PRO 6000 and H200 NVL for a first AI project. This article covers the step from the desk to the server room.
DGX Spark and H200 NVL specifications side by side
| SPECIFICATION | DGX SPARK (128 GB) | H200 NVL |
|---|---|---|
| Memory | 128 GB LPDDR5x, unified with the CPU | 141 GB HBM3e per card |
| Memory bandwidth | 273 GB/s | 4.8 TB/s |
| Tensor peak, sparse | 1 PFLOP FP4 | 3,341 TFLOPS FP8; no FP4 rate listed |
| Compute capability | 12.1 (GB10) | 9.0 (listed as H200) |
| Host CPU | 20 Arm cores in the GB10 | x86 server processors, such as AMD EPYC or Intel Xeon |
| Link to more GPUs | ConnectX-7 at 200 Gbps, up to four systems | NVLink bridge for 2 or 4 cards, 900 GB/s per GPU |
| MIG | GB10 not in NVIDIA’s MIG list | up to 7 instances of 16.5 GB |
| Power | 240 W power supply, 140 W GB10 TDP | up to 600 W per card, configurable |
| Form factor | desktop, 150 × 150 × 50.5 mm | dual-slot air-cooled PCIe card, servers with up to 8 GPUs |
| NVIDIA AI Enterprise | separate entitlement | five-year subscription included |
NVIDIA DGX Spark product page and DGX Spark hardware overview (updated 10 September 2026); NVIDIA H200 product page, H200 NVL column; NVIDIA CUDA GPUs page; MIG user guide, supported GPUs (updated 11 September 2026); host CPUs from our AI servers page.
The two Tensor Core figures are sparsity figures at different precisions and cannot be compared. NVIDIA also lists a 64 GB DGX Spark, sold only through OEM partners and shown as “Coming Soon” on its product page; this article uses the 128 GB Founders Edition, the version we supply.
Memory and bandwidth: 128 GB shared against 141 GB per card
The 128 GB of a DGX Spark serve the 20 Arm cores, DGX OS and the GPU from one pool, while the 141 GB of an H200 NVL belong to the GPU alone. Our sizing articles therefore take 102 GB, about 95 GiB, for weights and KV cache on a DGX Spark (128 GB). For a dedicated card our rule is 90 per cent of the memory the driver reports, less 3 GiB, which leaves about 123 GiB on an H200 NVL; four cards on a four-way NVLink bridge hold 564 GB between them.
For gpt-oss-120b, with its 60.8 GiB checkpoint and about 1.1 GiB of 16-bit cache per conversation of 32K tokens, that gives room for about 30 conversations on one DGX Spark (128 GB) and about 55 on one H200 NVL, by our memory estimate. At the 0.7 memory setting that NVIDIA’s vLLM release notes suggest for DGX Spark, the Spark figure falls to about 20.
Memory bandwidth sets the upper bound on generation speed, because each generated token reads the weights the model uses for it. The 4.8 TB/s of the H200 NVL is 17.6 times the 273 GB/s of DGX Spark, which raises that bound by the same factor for the same model at the same precision; measured speeds also depend on the engine, the kernels and the batch size. A test on a Spark shows that a model, a prompt template and a pipeline work, but not the response times on the server. Our DGX Spark benchmarks collect the measured figures, and the per-request speeds at 1 to 32 requests are in our article on one DGX Spark shared by a team. Load tests at production concurrency belong on the H200 NVL server before the rollout.
Develop on 2 to 4 DGX Spark, serve on H200 NVL servers
Take a company of 2,000 employees that plans a private assistant on gpt-oss-120b at a 32K context. Our private ChatGPT server sizing by company size works with example values of 80 requests in flight at the peak. Two servers with two H200 NVL each hold about 220 conversations, and about 110 if one server is down, which still covers the peak. For 500 employees and a peak of 20, one H200 NVL already holds about 55, so two servers with one card each keep the service running through the loss of one; leave slots and power for a second card per server.
On the development side, two DGX Spark suit a small platform team. One runs the candidate model behind vLLM for the application developers, and the other takes evaluation runs and fine-tuning experiments, which on one Spark would compete with serving for the same memory. NVIDIA’s product page gives fine-tuning “up to 70 billion parameters”, which NVIDIA’s fine-tuning playbook reaches with QLoRA, and inference up to 200 billion parameters on one 128 GB system, which by our arithmetic holds only for 4-bit weights. Four DGX Spark fit when several teams build applications at the same time, or when a larger model must be tested across linked systems: NVIDIA states up to 400 billion parameters on two 128 GB systems and up to 700 billion on four.
We supply the DGX Spark Founders Edition for development and AI servers built to order with H200 NVL for production, on one EU contract and invoice. Tell us your headcount, the model and how many developers build on it, and we size both ends.
What carries over from DGX Spark to the H200 NVL server
vLLM serves the OpenAI Completions and Chat Completions APIs, /v1/chat/completions among them, from its HTTP server on both systems. An application built against vLLM on a Spark points at the server’s address afterwards and sends the same requests. Prompts, evaluation sets and RAG indexes carry over too, and so do checkpoints in a precision both systems run.
NVIDIA’s vLLM container on NGC shows “Multi-Arch Support” as “Yes”; its latest tag on 28 September 2026 was 26.09-py3, which NVIDIA’s release notes give as vLLM 0.29.0 on CUDA 13.4.1. The page does not list the architectures, so check with docker manifest inspect that the tag serves both the Arm host of the Spark and the x86 host of the server.
NVIDIA NIM lists verified GPUs per model. Its support matrix, updated 6 October 2026, names both the NVIDIA-GB10 and the NVIDIA-H200-NVL for gpt-oss-20b, llama-3.1-8b-instruct and nemotron-3-nano. For gpt-oss-120b it lists the H200 NVL and not the GB10, so a team that wants that model as a NIM in production tests it on the Spark with vLLM and verifies the NIM on the server. NVIDIA’s DGX Spark quickstart gives access to eligible NIM containers “through NVIDIA Developer Program membership or through NVIDIA AI Enterprise”. NVIDIA’s H200 page states that the “H200 NVL comes with a five-year NVIDIA AI Enterprise subscription”, while on DGX Spark the entitlement exists “only if you purchased it, requested an evaluation, or received an NVIDIA Entitlement Certificate”.
What changes: Arm and x86 hosts, compute capability 12.1 and 9.0
DGX Spark runs DGX OS, based on Ubuntu 24.04, on 20 Arm cores, while our H200 NVL servers use x86 processors such as AMD EPYC or Intel Xeon. Your own images need builds for linux/arm64 and linux/amd64, and every Python package with compiled extensions needs a wheel or a build for both.
NVIDIA’s CUDA GPUs page lists the GB10 of DGX Spark at 12.1 and the H200 at 9.0. NVIDIA’s Blackwell compatibility guide states that “A cubin generated for a certain compute capability is supported to run on any GPU with the same major revision and same or higher minor revision”, and that PTX runs only on GPUs with a compute capability at or above the one it assumes. A custom kernel or extension compiled on the Spark for 12.1 therefore does not run on the H200 NVL. Build it with a 9.0 target as well, for example -gencode arch=compute_, and test both builds in the same CI pipeline.
NVIDIA’s vLLM release notes for 26.09 warn that the default memory allocation “can lead to out-of-memory errors” on unified memory such as DGX Spark and suggest --gpu-memory-utilization 0.7. They also require --max-num-seqs 4 for Nemotron Nano V3 and Super V3 in NVFP4 on Spark. A production configuration sets both values from the H200 NVL and the measured load. NVIDIA’s porting guide adds that “GPUDirect RDMA technology is not supported” on DGX Spark, so storage and network paths that rely on it can be tested only on the server.
FP4 on DGX Spark, FP8 on the H200 NVL
NVIDIA describes the Tensor Cores of DGX Spark as fifth-generation “with FP4 support”, and its vLLM playbook for the Spark serves an NVFP4 checkpoint. The H200 product page gives Tensor Core rates for FP8 and INT8 at the low end and none for FP4. A model chosen on a Spark in NVFP4 usually reaches the H200 NVL in FP8, which adds about 70 per cent to its weights: NVIDIA’s Llama 4 Scout checkpoints are 65.3 GB in NVFP4 and 111.6 GB in FP8.
Evaluate on the Spark the checkpoint you will serve, where it fits, so that answer quality is measured on the same weights. Qwen3.8-27B in its 30.9 GB FP8 version runs on one Spark and on one H200 NVL. vLLM’s quantisation page, dated 14 September 2026, marks its Marlin kernels for “GPTQ/AWQ/FP8/FP4” as supported on Hopper, and its Model Optimizer page, dated 2 October 2026, states that without a native FP4 kernel “vLLM falls back to weight-only (W4A16) execution via Marlin”. Some FP4 checkpoints therefore load on the H200 NVL and save memory there, while the arithmetic runs in 16-bit. Our article on FP4 models on H200 NVL and Ada GPUs covers which formats run and at what cost.
Serving 500 to 2,000 users: batching, MIG and NVLink
A production service batches many requests per pass over the weights, and the H200 NVL’s bandwidth sets how fast each pass reads them. It also offers two features that DGX Spark lacks. NVIDIA’s MIG user guide lists the H200 NVL with seven instances and does not list the GB10, and the product page gives “Up to 7 MIGs @16.5GB each”. One card can therefore carry the embedding model and reranker of a RAG service beside the main model on the other cards.
NVLink joins two or four H200 NVL at 900 GB/s per GPU, which lets a model larger than 141 GB, or its cache for long contexts, span cards with a fast link. DGX Spark systems connect over ConnectX-7 at 200 Gbps, about 25 GB/s.
| TASK | WHERE IT RUNS | WHY |
|---|---|---|
| Choosing a model | DGX Spark | NVIDIA gives up to 200B in 4-bit on one system |
| Fine-tuning experiments | DGX Spark | NVIDIA gives up to 70B with QLoRA on one system |
| Building the RAG pipeline | DGX Spark, multi-arch | same API and containers as the server |
| Load tests at peak | H200 NVL server | speed and batching match production |
| Serving 500 to 2,000 staff | two H200 NVL servers | 4.8 TB/s per card, one server can fail |
| Long contexts, large models | H200 NVL, NVLink | 2 or 4 cards at 900 GB/s per GPU |
| Isolated slices per team | H200 NVL with MIG | up to 7 instances of 16.5 GB |
Our reading of NVIDIA’s DGX Spark and H200 product pages and MIG user guide, October 2026; server counts from the worked example above.
We check the rack, power and airflow before we quote an H200 NVL server. Send us the rack position’s power feed and the model you tested on DGX Spark through the form below.
What we supply
We supply the DGX Spark Founders Edition (128 GB) for development and the H200 NVL with NVLink bridges, as cards or in AI servers built to order, on one EU contract and invoice with manufacturer warranty. We size the production server from the model, its precision, the context length and the peak requests in flight, and the configuration and quote follow within one business day. NVIDIA AI Enterprise licences come on the same invoice as the hardware. Our DGX Spark page and professional GPU range list the details, and running the model, RAG and MLOps on top is our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
DGX Spark vs H200: which is faster?
Can DGX Spark be used for production?
DGX Spark or GPU server for a company AI platform?
How do I migrate from DGX Spark to an H200 NVL server?
Does the H200 NVL support FP4 like DGX Spark?
How does DGX Spark memory compare with a data centre GPU?
Send us the model you test on DGX Spark, its precision, the context length, your headcount and the peak number of requests in flight. We reply within one business day with a configuration and quote for the DGX Spark systems and the H200 NVL servers; we check the rack, power and airflow before we quote.
Talk to an expertWe reply within one business day