BLOG · COMPARISON ·

DGX Spark vs H200 NVL: a development system and a production GPU server compared

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • DGX Spark (128 GB) has unified LPDDR5x memory at 273 GB/s and FP4 Tensor Cores in a desktop system; one H200 NVL has 141 GB of HBM3e at 4.8 TB/s, NVLink bridges for 2 or 4 cards and MIG, in a server with up to 8 cards
  • In a company the two usually belong to one plan, with developers building and testing on 2 to 4 DGX Spark and 500 to 2,000 employees using the service on two H200 NVL servers, so that one server can fail
  • NVIDIA’s vLLM container is multi-architecture, vLLM serves the same OpenAI-compatible API on both, and NIM lists gpt-oss-20b, Llama 3.1 8B and Nemotron 3 Nano as verified on both the GB10 and the H200 NVL, and gpt-oss-120b on the H200 NVL but not the GB10
  • On the server the host changes from Arm to x86 and the compute capability from 12.1 to 9.0, so images and custom CUDA kernels need builds for both, and NVFP4 checkpoints usually give way to FP8, since NVIDIA lists no FP4 Tensor Core rate for the H200 NVL
  • By our memory estimate gpt-oss-120b at 32K leaves room for about 30 conversations on one DGX Spark (128 GB) and about 55 on one H200 NVL; the 17.6 times higher memory bandwidth of the H200 NVL raises the upper bound on each user’s generation speed by the same factor

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

DGX Spark vs H200 NVL: development system and production GPU

DGX Spark and the H200 NVL do different jobs. DGX Spark is a desktop system with 128 GB of unified LPDDR5x memory at 273 GB/s and FP4 Tensor Cores, built for developers who prototype, fine-tune and test models. One H200 NVL is a PCIe card with 141 GB of HBM3e at 4.8 TB/s, NVLink bridges for two or four cards and MIG, and it goes into a server that serves a model to many users at once. All figures are as NVIDIA states them on its product pages, read in October 2026.

For a company with 200 to 2,000 staff, the two usually belong to one plan. The people who build the assistant, the RAG pipeline or the agents work on 2 to 4 DGX Spark, and the service that 500 to 2,000 employees use runs on H200 NVL servers. Which hardware a first project should start on is covered in our comparison of DGX Spark, RTX PRO 6000 and H200 NVL for a first AI project. This article covers the step from the desk to the server room.

DGX Spark and H200 NVL specifications side by side

SPECIFICATIONDGX SPARK (128 GB)H200 NVL
Memory128 GB LPDDR5x, unified with the CPU141 GB HBM3e per card
Memory bandwidth273 GB/s4.8 TB/s
Tensor peak, sparse1 PFLOP FP43,341 TFLOPS FP8; no FP4 rate listed
Compute capability12.1 (GB10)9.0 (listed as H200)
Host CPU20 Arm cores in the GB10x86 server processors, such as AMD EPYC or Intel Xeon
Link to more GPUsConnectX-7 at 200 Gbps, up to four systemsNVLink bridge for 2 or 4 cards, 900 GB/s per GPU
MIGGB10 not in NVIDIA’s MIG listup to 7 instances of 16.5 GB
Power240 W power supply, 140 W GB10 TDPup to 600 W per card, configurable
Form factordesktop, 150 × 150 × 50.5 mmdual-slot air-cooled PCIe card, servers with up to 8 GPUs
NVIDIA AI Enterpriseseparate entitlementfive-year subscription included

NVIDIA DGX Spark product page and DGX Spark hardware overview (updated 10 September 2026); NVIDIA H200 product page, H200 NVL column; NVIDIA CUDA GPUs page; MIG user guide, supported GPUs (updated 11 September 2026); host CPUs from our AI servers page.

The two Tensor Core figures are sparsity figures at different precisions and cannot be compared. NVIDIA also lists a 64 GB DGX Spark, sold only through OEM partners and shown as “Coming Soon” on its product page; this article uses the 128 GB Founders Edition, the version we supply.

Memory and bandwidth: 128 GB shared against 141 GB per card

The 128 GB of a DGX Spark serve the 20 Arm cores, DGX OS and the GPU from one pool, while the 141 GB of an H200 NVL belong to the GPU alone. Our sizing articles therefore take 102 GB, about 95 GiB, for weights and KV cache on a DGX Spark (128 GB). For a dedicated card our rule is 90 per cent of the memory the driver reports, less 3 GiB, which leaves about 123 GiB on an H200 NVL; four cards on a four-way NVLink bridge hold 564 GB between them.

For gpt-oss-120b, with its 60.8 GiB checkpoint and about 1.1 GiB of 16-bit cache per conversation of 32K tokens, that gives room for about 30 conversations on one DGX Spark (128 GB) and about 55 on one H200 NVL, by our memory estimate. At the 0.7 memory setting that NVIDIA’s vLLM release notes suggest for DGX Spark, the Spark figure falls to about 20.

Memory bandwidth sets the upper bound on generation speed, because each generated token reads the weights the model uses for it. The 4.8 TB/s of the H200 NVL is 17.6 times the 273 GB/s of DGX Spark, which raises that bound by the same factor for the same model at the same precision; measured speeds also depend on the engine, the kernels and the batch size. A test on a Spark shows that a model, a prompt template and a pipeline work, but not the response times on the server. Our DGX Spark benchmarks collect the measured figures, and the per-request speeds at 1 to 32 requests are in our article on one DGX Spark shared by a team. Load tests at production concurrency belong on the H200 NVL server before the rollout.

Develop on 2 to 4 DGX Spark, serve on H200 NVL servers

Take a company of 2,000 employees that plans a private assistant on gpt-oss-120b at a 32K context. Our private ChatGPT server sizing by company size works with example values of 80 requests in flight at the peak. Two servers with two H200 NVL each hold about 220 conversations, and about 110 if one server is down, which still covers the peak. For 500 employees and a peak of 20, one H200 NVL already holds about 55, so two servers with one card each keep the service running through the loss of one; leave slots and power for a second card per server.

On the development side, two DGX Spark suit a small platform team. One runs the candidate model behind vLLM for the application developers, and the other takes evaluation runs and fine-tuning experiments, which on one Spark would compete with serving for the same memory. NVIDIA’s product page gives fine-tuning “up to 70 billion parameters”, which NVIDIA’s fine-tuning playbook reaches with QLoRA, and inference up to 200 billion parameters on one 128 GB system, which by our arithmetic holds only for 4-bit weights. Four DGX Spark fit when several teams build applications at the same time, or when a larger model must be tested across linked systems: NVIDIA states up to 400 billion parameters on two 128 GB systems and up to 700 billion on four.

We supply the DGX Spark Founders Edition for development and AI servers built to order with H200 NVL for production, on one EU contract and invoice. Tell us your headcount, the model and how many developers build on it, and we size both ends.

What carries over from DGX Spark to the H200 NVL server

vLLM serves the OpenAI Completions and Chat Completions APIs, /v1/chat/completions among them, from its HTTP server on both systems. An application built against vLLM on a Spark points at the server’s address afterwards and sends the same requests. Prompts, evaluation sets and RAG indexes carry over too, and so do checkpoints in a precision both systems run.

NVIDIA’s vLLM container on NGC shows “Multi-Arch Support” as “Yes”; its latest tag on 28 September 2026 was 26.09-py3, which NVIDIA’s release notes give as vLLM 0.29.0 on CUDA 13.4.1. The page does not list the architectures, so check with docker manifest inspect that the tag serves both the Arm host of the Spark and the x86 host of the server.

NVIDIA NIM lists verified GPUs per model. Its support matrix, updated 6 October 2026, names both the NVIDIA-GB10 and the NVIDIA-H200-NVL for gpt-oss-20b, llama-3.1-8b-instruct and nemotron-3-nano. For gpt-oss-120b it lists the H200 NVL and not the GB10, so a team that wants that model as a NIM in production tests it on the Spark with vLLM and verifies the NIM on the server. NVIDIA’s DGX Spark quickstart gives access to eligible NIM containers “through NVIDIA Developer Program membership or through NVIDIA AI Enterprise”. NVIDIA’s H200 page states that the “H200 NVL comes with a five-year NVIDIA AI Enterprise subscription”, while on DGX Spark the entitlement exists “only if you purchased it, requested an evaluation, or received an NVIDIA Entitlement Certificate”.

What changes: Arm and x86 hosts, compute capability 12.1 and 9.0

DGX Spark runs DGX OS, based on Ubuntu 24.04, on 20 Arm cores, while our H200 NVL servers use x86 processors such as AMD EPYC or Intel Xeon. Your own images need builds for linux/arm64 and linux/amd64, and every Python package with compiled extensions needs a wheel or a build for both.

NVIDIA’s CUDA GPUs page lists the GB10 of DGX Spark at 12.1 and the H200 at 9.0. NVIDIA’s Blackwell compatibility guide states that “A cubin generated for a certain compute capability is supported to run on any GPU with the same major revision and same or higher minor revision”, and that PTX runs only on GPUs with a compute capability at or above the one it assumes. A custom kernel or extension compiled on the Spark for 12.1 therefore does not run on the H200 NVL. Build it with a 9.0 target as well, for example -gencode arch=compute_90,code=sm_90, and test both builds in the same CI pipeline.

NVIDIA’s vLLM release notes for 26.09 warn that the default memory allocation “can lead to out-of-memory errors” on unified memory such as DGX Spark and suggest --gpu-memory-utilization 0.7. They also require --max-num-seqs 4 for Nemotron Nano V3 and Super V3 in NVFP4 on Spark. A production configuration sets both values from the H200 NVL and the measured load. NVIDIA’s porting guide adds that “GPUDirect RDMA technology is not supported” on DGX Spark, so storage and network paths that rely on it can be tested only on the server.

FP4 on DGX Spark, FP8 on the H200 NVL

NVIDIA describes the Tensor Cores of DGX Spark as fifth-generation “with FP4 support”, and its vLLM playbook for the Spark serves an NVFP4 checkpoint. The H200 product page gives Tensor Core rates for FP8 and INT8 at the low end and none for FP4. A model chosen on a Spark in NVFP4 usually reaches the H200 NVL in FP8, which adds about 70 per cent to its weights: NVIDIA’s Llama 4 Scout checkpoints are 65.3 GB in NVFP4 and 111.6 GB in FP8.

Evaluate on the Spark the checkpoint you will serve, where it fits, so that answer quality is measured on the same weights. Qwen3.8-27B in its 30.9 GB FP8 version runs on one Spark and on one H200 NVL. vLLM’s quantisation page, dated 14 September 2026, marks its Marlin kernels for “GPTQ/AWQ/FP8/FP4” as supported on Hopper, and its Model Optimizer page, dated 2 October 2026, states that without a native FP4 kernel “vLLM falls back to weight-only (W4A16) execution via Marlin”. Some FP4 checkpoints therefore load on the H200 NVL and save memory there, while the arithmetic runs in 16-bit. Our article on FP4 models on H200 NVL and Ada GPUs covers which formats run and at what cost.

Serving 500 to 2,000 users: batching, MIG and NVLink

A production service batches many requests per pass over the weights, and the H200 NVL’s bandwidth sets how fast each pass reads them. It also offers two features that DGX Spark lacks. NVIDIA’s MIG user guide lists the H200 NVL with seven instances and does not list the GB10, and the product page gives “Up to 7 MIGs @16.5GB each”. One card can therefore carry the embedding model and reranker of a RAG service beside the main model on the other cards.

NVLink joins two or four H200 NVL at 900 GB/s per GPU, which lets a model larger than 141 GB, or its cache for long contexts, span cards with a fast link. DGX Spark systems connect over ConnectX-7 at 200 Gbps, about 25 GB/s.

TASKWHERE IT RUNSWHY
Choosing a modelDGX SparkNVIDIA gives up to 200B in 4-bit on one system
Fine-tuning experimentsDGX SparkNVIDIA gives up to 70B with QLoRA on one system
Building the RAG pipelineDGX Spark, multi-archsame API and containers as the server
Load tests at peakH200 NVL serverspeed and batching match production
Serving 500 to 2,000 stafftwo H200 NVL servers4.8 TB/s per card, one server can fail
Long contexts, large modelsH200 NVL, NVLink2 or 4 cards at 900 GB/s per GPU
Isolated slices per teamH200 NVL with MIGup to 7 instances of 16.5 GB

Our reading of NVIDIA’s DGX Spark and H200 product pages and MIG user guide, October 2026; server counts from the worked example above.

We check the rack, power and airflow before we quote an H200 NVL server. Send us the rack position’s power feed and the model you tested on DGX Spark through the form below.

What we supply

We supply the DGX Spark Founders Edition (128 GB) for development and the H200 NVL with NVLink bridges, as cards or in AI servers built to order, on one EU contract and invoice with manufacturer warranty. We size the production server from the model, its precision, the context length and the peak requests in flight, and the configuration and quote follow within one business day. NVIDIA AI Enterprise licences come on the same invoice as the hardware. Our DGX Spark page and professional GPU range list the details, and running the model, RAG and MLOps on top is our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

DGX Spark vs H200: which is faster?
The H200 NVL is faster for the same model at the same precision. Its 4.8 TB/s of memory bandwidth is 17.6 times the 273 GB/s of DGX Spark, and memory bandwidth sets the upper bound on token generation; measured speeds also depend on the engine and the batch size. DGX Spark has more memory than many cards, 128 GB unified, but one H200 NVL has 141 GB of HBM3e for the GPU alone.
Can DGX Spark be used for production?
NVIDIA positions DGX Spark for prototyping, testing and validating models, with “eventual migration” to data-centre or cloud infrastructure. One Spark can serve a pilot or a small team, but for a service used by 500 to 2,000 employees its 273 GB/s of memory bandwidth limits speed. Two H200 NVL servers serve that load and keep it running if one fails.
DGX Spark or GPU server for a company AI platform?
A company usually needs both, for different people. Developers build and test on 2 to 4 DGX Spark, and the service for employees runs on H200 NVL servers sized by the peak number of requests in flight. By our estimate, two servers with two H200 NVL each hold about 220 conversations of gpt-oss-120b at 32K.
How do I migrate from DGX Spark to an H200 NVL server?
Keep the vLLM API, the multi-architecture containers, the prompts and the evaluation sets, and change three things. Build your own images and CUDA extensions for x86 and for compute capability 9.0 as well as Arm and 12.1, and replace an NVFP4 checkpoint with an FP8 one. Set memory utilisation and concurrency from the H200 NVL instead of the Spark’s workaround values.
Does the H200 NVL support FP4 like DGX Spark?
The H200 NVL has no FP4 Tensor Core rate on NVIDIA’s product page, which lists its rates down to FP8 and INT8, while DGX Spark has fifth-generation Tensor Cores with FP4 support. vLLM loads some FP4 checkpoints on Hopper through its weight-only Marlin fallback, which saves memory while the arithmetic runs in 16-bit.
How does DGX Spark memory compare with a data centre GPU?
DGX Spark has 128 GB of LPDDR5x at 273 GB/s, shared by the Arm CPU, the operating system and the GPU. An H200 NVL has 141 GB of HBM3e at 4.8 TB/s for the GPU alone. For sizing we take 102 GB, about 95 GiB, for weights and cache on a DGX Spark (128 GB) and about 123 GiB on one H200 NVL.

Send us the model you test on DGX Spark, its precision, the context length, your headcount and the peak number of requests in flight. We reply within one business day with a configuration and quote for the DGX Spark systems and the H200 NVL servers; we check the rack, power and airflow before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna