BLOG · COMPARISON ·

DGX Spark 64 GB vs 128 GB: which configuration runs which models, and when two systems make sense

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • The 64 GB DGX Spark keeps the GB10 chip, DGX OS and NVIDIA’s software stack; NVIDIA’s specification table gives one 256-bit interface at 273 GB/s for both memory sizes and rates 64 GB for models up to 100 billion parameters, 128 GB for up to 200 billion
  • Only manufacturers sell the 64 GB configuration, six of them from 23 October 2026 according to NVIDIA’s blog; NVIDIA’s own Founders Edition comes with 128 GB and a 4 TB self-encrypting drive
  • By the same arithmetic as our 128 GB sizing, 64 GB leaves about 51 to 58 GB for weights and KV cache: Qwen3.8-27B in FP8 and Llama 3.3 70B in NVFP4 fit, while gpt-oss-120b (65.3 GB) and a 70B model in FP8 do not
  • Llama 3.3 70B in NVFP4 with an FP8 KV cache holds about 6 to 11 conversations at 8K tokens in 64 GB and about 44 to 54 in 128 GB, by our arithmetic
  • Two linked 64 GB systems pool 128 GB and NVIDIA rates them for up to 200 billion parameters, but traffic between them crosses a link NVIDIA measures at 189.85 Gb/s; NVIDIA gives no figure for more than two 64 GB systems

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

DGX Spark 64 GB vs 128 GB: what differs and what fits

The 64 GB DGX Spark is the same system as the 128 GB one with half the memory. NVIDIA states that it keeps the GB10 Grace Blackwell Superchip, DGX OS and the full NVIDIA AI software stack, and its specification table gives one memory bandwidth, 273 GB/s, for both. NVIDIA rates the 64 GB configuration for models of up to 100 billion parameters and the 128 GB configuration for up to 200 billion; by our arithmetic below, both figures hold only for 4-bit weights. Only manufacturers sell the 64 GB configuration, while NVIDIA’s own Founders Edition comes with 128 GB and a 4 TB self-encrypting drive.

The choice therefore follows from the largest model you plan to run and its precision. A dense model of about 27 billion parameters in FP8, or a 70B model at 4-bit with a short context, fits in 64 GB. gpt-oss-120b, a 70B model in FP8 and anything close to 200 billion parameters need 128 GB.

What NVIDIA announced for the 64 GB configuration

NVIDIA announced the configuration on its blog on 2 October 2026. The post says the 64 GB model keeps “the GB10 Grace Blackwell Superchip, DGX OS and full NVIDIA AI software stack”, the same as the 128 GB model, and that it “supports up to 100-billion-parameter models”. According to the blog, six manufacturers, among them Acer, Dell, HP and MSI, will sell it from 23 October 2026. Lenovo, which also builds GB10 systems, is not on that list. NVIDIA’s product page shows the configuration as “Coming Soon”, and its footnote reads “64 GB memory configuration is available exclusively through participating OEM partners.”

MSI announced a 64 GB EdgeXpert on 2 October, rated for models of up to 100 billion parameters and fitted with the ConnectX-7 adapter for clustering. Acer’s Veriton GN100 page offers “a choice of 64 GB or 128GB of unified memory”. On 9 October 2026, Dell’s GB10 shop page and HP’s ZGX Nano page listed 128 GB only, and HP’s page gave up to 405 billion parameters for two linked systems, where NVIDIA’s product page now gives 400 billion.

NVIDIA’s specification table has one storage row for both memory sizes, and its only measurement for 64 GB is the two-system test described below. For the larger system the product page states “With 128 GB of unified system memory, fine-tune models up to 70 billion parameters”, and it has no matching sentence for 64 GB.

DGX Spark 64 GB and 128 GB specifications compared

The table lists only values NVIDIA states; where its specification table has one row for both memory sizes, the value appears in both columns.

SPECIFICATION64 GB CONFIGURATION128 GB CONFIGURATION
ChipGB10 Grace BlackwellGB10 Grace Blackwell
CPU20 Arm cores: 10 Cortex-X925, 10 Cortex-A72520 Arm cores: 10 Cortex-X925, 10 Cortex-A725
Memory64 GB LPDDR5X, coherent, unified128 GB LPDDR5x, coherent, unified
Interface, bandwidth256-bit, 273 GB/s256-bit, 273 GB/s
FP4 headlineup to 1 PFLOP, with sparsityup to 1 PFLOP, with sparsity
NetworkConnectX-7 at 200 Gbps, 10 GbEConnectX-7 at 200 Gbps, 10 GbE
Storageup to 4 TB, self-encryptingup to 4 TB; Founders Edition 4 TB
Inference, one systemup to 100B parametersup to 200B parameters
Two linked systems128 GB, up to 200B parameters256 GB, up to 400B parameters
Four linked systemsno NVIDIA figure512 GB, up to 700B parameters
Fine-tuningno NVIDIA figureup to 70B parameters
Sold asmanufacturers’ systems onlyFounders Edition and manufacturers’ systems

NVIDIA DGX Spark product page and specification table (updated 7 October 2026), NVIDIA blog of 2 October 2026, DGX Spark hardware overview (updated 10 September 2026) for the sparsity condition, NVIDIA Marketplace for the Founders Edition drive; all read on 9 October 2026.

We found no separate bandwidth figure for 64 GB from NVIDIA, MSI or Acer, so the 273 GB/s rests on that single table row.

Token generation follows memory bandwidth, as our DGX Spark benchmark article shows with measured rates. Llama 3.3 70B in NVFP4 reads about 40.6 GB per generated token, which caps one user at under 7 tokens per second on either configuration. The larger configuration adds room for bigger models and more KV cache at the same generation ceiling.

How much of 64 GB a model can use

NVIDIA publishes no usable-capacity figure for either configuration. Our article on what fits in 128 GB on a DGX Spark takes the memory fractions in NVIDIA’s serving playbooks, 0.8 and 0.9, and applies them to the whole pool, which gives a working set of about 102 to 115 GB for weights and KV cache. The same arithmetic on 64 GB gives about 51 to 58 GB. We found no NVIDIA serving playbook for 64 GB, so this range is our estimate.

The memory outside the working set does not halve with the configuration. DGX OS, the desktop session, the display reserve (2 GB by default on the 128 GB system) and the processes on the Arm cores need about the same on both. On 128 GB the playbook fractions leave 13 to 26 GB for them, and on 64 GB only 6 to 13 GB. On a 64 GB system that someone also uses as a workstation, plan with the lower end, about 51 GB.

At about 55 GB, the middle of that range, BF16 weights reach about 27 billion parameters, FP8 about 55 billion and the 4-bit formats with their block scales about 100 billion. NVIDIA’s 100 billion figure is therefore a 4-bit figure. At 4.25 to 4.5 bits per weight, 100 billion parameters take 53 to 56 GB, which leaves a few GB of KV cache at the top of the range and none at the low end.

Which models fit in 64 GB and in 128 GB

The table applies the two working sets to checkpoints a buyer is likely to compare. “Left” is the room for the KV cache after the weights load.

MODEL, PRECISIONWEIGHTS64 GB128 GBLINKED SYSTEMS
gpt-oss-20b, MXFP413.8 GBfits, 37 to 44 GB leftfitsnot needed
Qwen3.8-27B, FP830.9 GBfits, 20 to 27 GB leftfitsnot needed
Qwen3.8-27B, BF1655.6 GB2 GB left at mostfits, 47 to 60 GB lefttwo 64 GB
Llama 3.3 70B, NVFP442.7 GBfits, 8.5 to 15 GB leftfits, 60 to 72 GB leftnot needed
Llama 3.3 70B, FP8about 70 GBnofits, 32 to 45 GB lefttwo 64 GB
gpt-oss-120b, MXFP465.3 GBnofits, 37 to 50 GB lefttwo 64 GB
Any 200B model, 4-bit106 to 112 GBnoat the limittwo 64 GB at the limit, or two 128 GB
Qwen3-235B-A22B, NVFP4134 GBnono, multi-node in NVIDIA’s listtwo 128 GB

Weights from the Hugging Face file lists of openai/gpt-oss, Qwen3.8-27B, nvidia/Llama-3.3-70B-Instruct-FP4 and nvidia/Qwen3-235B-A22B-FP4, read on 9 October 2026; the 70B FP8 and 200B weights are our arithmetic at 1 byte and 4.25 to 4.5 bits per weight. Room left against working sets of 51 to 58 GB and 102 to 115 GB, our estimate.

For models that fit both, the KV cache separates the two configurations. Llama 3.3 70B has 80 layers, 8 KV heads and a head dimension of 128, so one conversation at 8,192 tokens takes 1.25 GiB in an FP8 cache. With the NVFP4 weights loaded, 64 GB holds about 6 to 11 such conversations at once and 128 GB about 44 to 54, before activations and runtime overhead. That suits one developer on 64 GB; a team or agents with long prompts need 128 GB.

On 128 GB, NVIDIA’s PyTorch playbook documents a full fine-tune at 3 billion parameters, LoRA at 8 billion and QLoRA at 70 billion. NVIDIA’s NeMo fine-tuning playbook starts its container with “the 64 GB limit used for the lower-memory Spark validation” and runs LoRA recipes for 8B models and a QLoRA recipe for Llama 3.3 70B in it. It does not say that these recipes ran on a 64 GB system, and a container limit is not the same as 64 GB of shared memory. By our arithmetic a full fine-tune of a 3B model in BF16 with Adam needs about 48 GB before activations, at the edge of a 51 to 58 GB working set, while QLoRA on a 70B model starts from a 4-bit base of roughly 40 GB. Test either on a 64 GB system before you plan around it. For models beyond these, our guide to LLM hardware requirements by model gives checkpoint sizes and cache figures for each.

We supply the DGX Spark Founders Edition with 128 GB. Tell us the models, their precision and the context length in the form below, and we reply whether one Spark, two DGX Spark or four fit the work.

Two 64 GB systems or one 128 GB system

NVIDIA rates both for models of up to 200 billion parameters. Its blog says that two 64 GB systems linked over the 200 GbE fabric with the NVIDIA Sync Cluster Assistant “pool their memory to 128GB” while “delivering twice the memory bandwidth”. In NVIDIA’s test with Qwen 3.8 27B, two such systems delivered up to 1.7 times the performance “compared with a single system”. The blog does not state that system’s memory size, and we found no NVIDIA measurement of two 64 GB systems against one 128 GB system.

Twice the bandwidth is the sum of two local pools of 273 GB/s each. Data that moves between the two systems crosses the QSFP link, for which NVIDIA’s own RDMA test reports 189.85 Gb/s in total. That is about 24 GB/s, less than a tenth of the local memory bandwidth. A model split across two systems gains speed only where the parallel layout keeps most traffic local, and our article on linking two DGX Spark systems shows that a second system adds capacity first and speed per request at most in part.

Two systems also mean two installations to run. Each runs its own DGX OS, so two operating system footprints come out of the pooled 128 GB. The Founders Edition ships without the QSFP cable and NVIDIA lists no BMC for it, as our guide to buying a DGX Spark in the EU describes; check both points for any manufacturer’s 64 GB system.

From one 128 GB system, NVIDIA’s page continues to two systems with 256 GB, rated for up to 400 billion parameters, and four with 512 GB, up to 700 billion. For the 64 GB configuration the page stops at two systems. NVIDIA’s clustering guide, updated on 10 September 2026, supports up to three systems by cable and four through a switch, but does not mention the 64 GB configuration or mixing memory sizes. We found no NVIDIA or manufacturer statement that 64 GB and 128 GB systems can work in one cluster, or that memory can be added after purchase.

We supply the 128 GB Founders Edition for one developer, as two DGX Spark linked into 256 GB, or four for a team. Describe the largest model you expect to run through the form below, and we quote the number of systems it needs.

Founders Edition or a manufacturer’s GB10 system

Every Founders Edition has 128 GB of memory, and NVIDIA’s Marketplace lists its drive as “4TB NVME.M2 with self-encryption”. NVIDIA’s other pages word storage differently. The specification table on the product page (updated 7 October 2026) gives “Up to 4 TB NVME.M2 with self-encryption” for both memory sizes, and the user guide’s hardware overview (updated 10 September 2026) gives “1 TB or 4 TB NVMe M.2 with self-encryption”. Partner GB10 systems also come with 1 or 2 TB. Neither NVIDIA’s blog nor MSI’s announcement states a drive size for the 64 GB systems, and Acer gives “up to 4TB” for both memory sizes, so check the drive size and encryption of any 64 GB system before you order.

The other difference is support. NVIDIA’s release notes and recovery images cover the Founders Edition, while manufacturers’ systems receive updates and recovery media from their makers, under the maker’s warranty. The buying guide linked above covers update timing and the warranty texts.

What we supply

We supply the NVIDIA DGX Spark Founders Edition, with 128 GB of unified memory and a 4 TB self-encrypting drive, across the EU on one contract and invoice, with manufacturer warranty. We do not supply the 64 GB configuration. Whether you need one Spark, two DGX Spark linked into 256 GB or four for a team, we quote exactly what you ask for. Where a model needs more generation speed than 273 GB/s allows, we also supply professional NVIDIA GPUs such as the RTX PRO 6000 and the H200 NVL, as cards or in AI servers built to order.

FAQ

What is the DGX Spark 64 GB?
It is a configuration of DGX Spark with 64 GB of LPDDR5X unified memory instead of 128 GB, announced by NVIDIA on 2 October 2026. NVIDIA states that it keeps the GB10 Grace Blackwell Superchip, DGX OS and the NVIDIA AI software stack of the 128 GB model, and rates it for models of up to 100 billion parameters. It is sold only through manufacturers, six of them from 23 October 2026, and not as an NVIDIA Founders Edition.
What is the difference between the DGX Spark 64 GB and 128 GB?
They differ in memory size and in what follows from it. NVIDIA’s specification table gives both the same 256-bit interface at 273 GB/s, the same 20-core Arm CPU and the same ConnectX-7 networking, and rates 64 GB for models up to 100 billion parameters and 128 GB for up to 200 billion. NVIDIA’s product page gives a fine-tuning figure of up to 70 billion parameters only for the 128 GB system.
Which models fit on a DGX Spark with 64 GB?
By our arithmetic about 51 to 58 GB of the 64 GB is available for weights and KV cache. gpt-oss-20b, Qwen3.8-27B in FP8 and Llama 3.3 70B in NVFP4 fit, the last with room for about 6 to 11 conversations at 8K tokens in an FP8 cache. gpt-oss-120b at 65.3 GB and any 70B model in FP8 do not fit.
Is the 64 GB DGX Spark slower than the 128 GB version?
Not for a model that fits in both, according to NVIDIA’s specifications. Its table gives a single memory bandwidth of 273 GB/s for both configurations, and token generation on DGX Spark follows memory bandwidth. NVIDIA has published no token rates for one 64 GB system, and the 64 GB system holds fewer concurrent conversations because less memory is left for the KV cache.
Two DGX Spark 64 GB or one DGX Spark 128 GB?
NVIDIA rates both for models of up to 200 billion parameters. Two 64 GB systems pool 128 GB with twice the aggregate memory bandwidth, but data between them crosses a link NVIDIA measures at 189.85 Gb/s, and each system runs its own operating system. NVIDIA reports up to 1.7 times the performance of a single system for two of them, and gives no comparison with one 128 GB system.
Is there a DGX Spark with 256 GB?
No single DGX Spark has 256 GB; NVIDIA lists configurations with 64 GB and 128 GB of unified memory. Two linked 128 GB systems pool 256 GB, which NVIDIA rates for models of up to 400 billion parameters, and four pool 512 GB for up to 700 billion. The Founders Edition comes only with 128 GB, while the 64 GB configuration is sold exclusively through participating OEM partners.

Send us the models you plan to run or fine-tune, their precision, the context length and how many people will use them at once. We reply within one business day with whether one Spark, two DGX Spark or four fit the work, and a written quote for the Founders Edition.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna