DGX Spark 64 GB vs 128 GB: which configuration runs which models, and when two systems make sense
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- The 64 GB DGX Spark keeps the GB10 chip, DGX OS and NVIDIA’s software stack; NVIDIA’s specification table gives one 256-bit interface at 273 GB/s for both memory sizes and rates 64 GB for models up to 100 billion parameters, 128 GB for up to 200 billion
- Only manufacturers sell the 64 GB configuration, six of them from 23 October 2026 according to NVIDIA’s blog; NVIDIA’s own Founders Edition comes with 128 GB and a 4 TB self-encrypting drive
- By the same arithmetic as our 128 GB sizing, 64 GB leaves about 51 to 58 GB for weights and KV cache: Qwen3.8-27B in FP8 and Llama 3.3 70B in NVFP4 fit, while gpt-oss-120b (65.3 GB) and a 70B model in FP8 do not
- Llama 3.3 70B in NVFP4 with an FP8 KV cache holds about 6 to 11 conversations at 8K tokens in 64 GB and about 44 to 54 in 128 GB, by our arithmetic
- Two linked 64 GB systems pool 128 GB and NVIDIA rates them for up to 200 billion parameters, but traffic between them crosses a link NVIDIA measures at 189.85 Gb/s; NVIDIA gives no figure for more than two 64 GB systems
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
DGX Spark 64 GB vs 128 GB: what differs and what fits
The 64 GB DGX Spark is the same system as the 128 GB one with half the memory. NVIDIA states that it keeps the GB10 Grace Blackwell Superchip, DGX OS and the full NVIDIA AI software stack, and its specification table gives one memory bandwidth, 273 GB/s, for both. NVIDIA rates the 64 GB configuration for models of up to 100 billion parameters and the 128 GB configuration for up to 200 billion; by our arithmetic below, both figures hold only for 4-bit weights. Only manufacturers sell the 64 GB configuration, while NVIDIA’s own Founders Edition comes with 128 GB and a 4 TB self-encrypting drive.
The choice therefore follows from the largest model you plan to run and its precision. A dense model of about 27 billion parameters in FP8, or a 70B model at 4-bit with a short context, fits in 64 GB. gpt-oss-120b, a 70B model in FP8 and anything close to 200 billion parameters need 128 GB.
What NVIDIA announced for the 64 GB configuration
NVIDIA announced the configuration on its blog on 2 October 2026. The post says the 64 GB model keeps “the GB10 Grace Blackwell Superchip, DGX OS and full NVIDIA AI software stack”, the same as the 128 GB model, and that it “supports up to 100-billion-parameter models”. According to the blog, six manufacturers, among them Acer, Dell, HP and MSI, will sell it from 23 October 2026. Lenovo, which also builds GB10 systems, is not on that list. NVIDIA’s product page shows the configuration as “Coming Soon”, and its footnote reads “64 GB memory configuration is available exclusively through participating OEM partners.”
MSI announced a 64 GB EdgeXpert on 2 October, rated for models of up to 100 billion parameters and fitted with the ConnectX-7 adapter for clustering. Acer’s Veriton GN100 page offers “a choice of 64 GB or 128GB of unified memory”. On 9 October 2026, Dell’s GB10 shop page and HP’s ZGX Nano page listed 128 GB only, and HP’s page gave up to 405 billion parameters for two linked systems, where NVIDIA’s product page now gives 400 billion.
NVIDIA’s specification table has one storage row for both memory sizes, and its only measurement for 64 GB is the two-system test described below. For the larger system the product page states “With 128 GB of unified system memory, fine-tune models up to 70 billion parameters”, and it has no matching sentence for 64 GB.
DGX Spark 64 GB and 128 GB specifications compared
The table lists only values NVIDIA states; where its specification table has one row for both memory sizes, the value appears in both columns.
| SPECIFICATION | 64 GB CONFIGURATION | 128 GB CONFIGURATION |
|---|---|---|
| Chip | GB10 Grace Blackwell | GB10 Grace Blackwell |
| CPU | 20 Arm cores: 10 Cortex-X925, 10 Cortex-A725 | 20 Arm cores: 10 Cortex-X925, 10 Cortex-A725 |
| Memory | 64 GB LPDDR5X, coherent, unified | 128 GB LPDDR5x, coherent, unified |
| Interface, bandwidth | 256-bit, 273 GB/s | 256-bit, 273 GB/s |
| FP4 headline | up to 1 PFLOP, with sparsity | up to 1 PFLOP, with sparsity |
| Network | ConnectX-7 at 200 Gbps, 10 GbE | ConnectX-7 at 200 Gbps, 10 GbE |
| Storage | up to 4 TB, self-encrypting | up to 4 TB; Founders Edition 4 TB |
| Inference, one system | up to 100B parameters | up to 200B parameters |
| Two linked systems | 128 GB, up to 200B parameters | 256 GB, up to 400B parameters |
| Four linked systems | no NVIDIA figure | 512 GB, up to 700B parameters |
| Fine-tuning | no NVIDIA figure | up to 70B parameters |
| Sold as | manufacturers’ systems only | Founders Edition and manufacturers’ systems |
NVIDIA DGX Spark product page and specification table (updated 7 October 2026), NVIDIA blog of 2 October 2026, DGX Spark hardware overview (updated 10 September 2026) for the sparsity condition, NVIDIA Marketplace for the Founders Edition drive; all read on 9 October 2026.
We found no separate bandwidth figure for 64 GB from NVIDIA, MSI or Acer, so the 273 GB/s rests on that single table row.
Token generation follows memory bandwidth, as our DGX Spark benchmark article shows with measured rates. Llama 3.3 70B in NVFP4 reads about 40.6 GB per generated token, which caps one user at under 7 tokens per second on either configuration. The larger configuration adds room for bigger models and more KV cache at the same generation ceiling.
How much of 64 GB a model can use
NVIDIA publishes no usable-capacity figure for either configuration. Our article on what fits in 128 GB on a DGX Spark takes the memory fractions in NVIDIA’s serving playbooks, 0.8 and 0.9, and applies them to the whole pool, which gives a working set of about 102 to 115 GB for weights and KV cache. The same arithmetic on 64 GB gives about 51 to 58 GB. We found no NVIDIA serving playbook for 64 GB, so this range is our estimate.
The memory outside the working set does not halve with the configuration. DGX OS, the desktop session, the display reserve (2 GB by default on the 128 GB system) and the processes on the Arm cores need about the same on both. On 128 GB the playbook fractions leave 13 to 26 GB for them, and on 64 GB only 6 to 13 GB. On a 64 GB system that someone also uses as a workstation, plan with the lower end, about 51 GB.
At about 55 GB, the middle of that range, BF16 weights reach about 27 billion parameters, FP8 about 55 billion and the 4-bit formats with their block scales about 100 billion. NVIDIA’s 100 billion figure is therefore a 4-bit figure. At 4.25 to 4.5 bits per weight, 100 billion parameters take 53 to 56 GB, which leaves a few GB of KV cache at the top of the range and none at the low end.
Which models fit in 64 GB and in 128 GB
The table applies the two working sets to checkpoints a buyer is likely to compare. “Left” is the room for the KV cache after the weights load.
| MODEL, PRECISION | WEIGHTS | 64 GB | 128 GB | LINKED SYSTEMS |
|---|---|---|---|---|
| gpt-oss-20b, MXFP4 | 13.8 GB | fits, 37 to 44 GB left | fits | not needed |
| Qwen3.8-27B, FP8 | 30.9 GB | fits, 20 to 27 GB left | fits | not needed |
| Qwen3.8-27B, BF16 | 55.6 GB | 2 GB left at most | fits, 47 to 60 GB left | two 64 GB |
| Llama 3.3 70B, NVFP4 | 42.7 GB | fits, 8.5 to 15 GB left | fits, 60 to 72 GB left | not needed |
| Llama 3.3 70B, FP8 | about 70 GB | no | fits, 32 to 45 GB left | two 64 GB |
| gpt-oss-120b, MXFP4 | 65.3 GB | no | fits, 37 to 50 GB left | two 64 GB |
| Any 200B model, 4-bit | 106 to 112 GB | no | at the limit | two 64 GB at the limit, or two 128 GB |
| Qwen3-235B-A22B, NVFP4 | 134 GB | no | no, multi-node in NVIDIA’s list | two 128 GB |
Weights from the Hugging Face file lists of openai/gpt-oss, Qwen3.8-27B, nvidia/Llama-3.3-70B-Instruct-FP4 and nvidia/Qwen3-235B-A22B-FP4, read on 9 October 2026; the 70B FP8 and 200B weights are our arithmetic at 1 byte and 4.25 to 4.5 bits per weight. Room left against working sets of 51 to 58 GB and 102 to 115 GB, our estimate.
For models that fit both, the KV cache separates the two configurations. Llama 3.3 70B has 80 layers, 8 KV heads and a head dimension of 128, so one conversation at 8,192 tokens takes 1.25 GiB in an FP8 cache. With the NVFP4 weights loaded, 64 GB holds about 6 to 11 such conversations at once and 128 GB about 44 to 54, before activations and runtime overhead. That suits one developer on 64 GB; a team or agents with long prompts need 128 GB.
On 128 GB, NVIDIA’s PyTorch playbook documents a full fine-tune at 3 billion parameters, LoRA at 8 billion and QLoRA at 70 billion. NVIDIA’s NeMo fine-tuning playbook starts its container with “the 64 GB limit used for the lower-memory Spark validation” and runs LoRA recipes for 8B models and a QLoRA recipe for Llama 3.3 70B in it. It does not say that these recipes ran on a 64 GB system, and a container limit is not the same as 64 GB of shared memory. By our arithmetic a full fine-tune of a 3B model in BF16 with Adam needs about 48 GB before activations, at the edge of a 51 to 58 GB working set, while QLoRA on a 70B model starts from a 4-bit base of roughly 40 GB. Test either on a 64 GB system before you plan around it. For models beyond these, our guide to LLM hardware requirements by model gives checkpoint sizes and cache figures for each.
We supply the DGX Spark Founders Edition with 128 GB. Tell us the models, their precision and the context length in the form below, and we reply whether one Spark, two DGX Spark or four fit the work.
Two 64 GB systems or one 128 GB system
NVIDIA rates both for models of up to 200 billion parameters. Its blog says that two 64 GB systems linked over the 200 GbE fabric with the NVIDIA Sync Cluster Assistant “pool their memory to 128GB” while “delivering twice the memory bandwidth”. In NVIDIA’s test with Qwen 3.8 27B, two such systems delivered up to 1.7 times the performance “compared with a single system”. The blog does not state that system’s memory size, and we found no NVIDIA measurement of two 64 GB systems against one 128 GB system.
Twice the bandwidth is the sum of two local pools of 273 GB/s each. Data that moves between the two systems crosses the QSFP link, for which NVIDIA’s own RDMA test reports 189.85 Gb/s in total. That is about 24 GB/s, less than a tenth of the local memory bandwidth. A model split across two systems gains speed only where the parallel layout keeps most traffic local, and our article on linking two DGX Spark systems shows that a second system adds capacity first and speed per request at most in part.
Two systems also mean two installations to run. Each runs its own DGX OS, so two operating system footprints come out of the pooled 128 GB. The Founders Edition ships without the QSFP cable and NVIDIA lists no BMC for it, as our guide to buying a DGX Spark in the EU describes; check both points for any manufacturer’s 64 GB system.
From one 128 GB system, NVIDIA’s page continues to two systems with 256 GB, rated for up to 400 billion parameters, and four with 512 GB, up to 700 billion. For the 64 GB configuration the page stops at two systems. NVIDIA’s clustering guide, updated on 10 September 2026, supports up to three systems by cable and four through a switch, but does not mention the 64 GB configuration or mixing memory sizes. We found no NVIDIA or manufacturer statement that 64 GB and 128 GB systems can work in one cluster, or that memory can be added after purchase.
We supply the 128 GB Founders Edition for one developer, as two DGX Spark linked into 256 GB, or four for a team. Describe the largest model you expect to run through the form below, and we quote the number of systems it needs.
Founders Edition or a manufacturer’s GB10 system
Every Founders Edition has 128 GB of memory, and NVIDIA’s Marketplace lists its drive as “4TB NVME.M2 with self-encryption”. NVIDIA’s other pages word storage differently. The specification table on the product page (updated 7 October 2026) gives “Up to 4 TB NVME.M2 with self-encryption” for both memory sizes, and the user guide’s hardware overview (updated 10 September 2026) gives “1 TB or 4 TB NVMe M.2 with self-encryption”. Partner GB10 systems also come with 1 or 2 TB. Neither NVIDIA’s blog nor MSI’s announcement states a drive size for the 64 GB systems, and Acer gives “up to 4TB” for both memory sizes, so check the drive size and encryption of any 64 GB system before you order.
The other difference is support. NVIDIA’s release notes and recovery images cover the Founders Edition, while manufacturers’ systems receive updates and recovery media from their makers, under the maker’s warranty. The buying guide linked above covers update timing and the warranty texts.
What we supply
We supply the NVIDIA DGX Spark Founders Edition, with 128 GB of unified memory and a 4 TB self-encrypting drive, across the EU on one contract and invoice, with manufacturer warranty. We do not supply the 64 GB configuration. Whether you need one Spark, two DGX Spark linked into 256 GB or four for a team, we quote exactly what you ask for. Where a model needs more generation speed than 273 GB/s allows, we also supply professional NVIDIA GPUs such as the RTX PRO 6000 and the H200 NVL, as cards or in AI servers built to order.
FAQ
What is the DGX Spark 64 GB?
What is the difference between the DGX Spark 64 GB and 128 GB?
Which models fit on a DGX Spark with 64 GB?
Is the 64 GB DGX Spark slower than the 128 GB version?
Two DGX Spark 64 GB or one DGX Spark 128 GB?
Is there a DGX Spark with 256 GB?
Send us the models you plan to run or fine-tune, their precision, the context length and how many people will use them at once. We reply within one business day with whether one Spark, two DGX Spark or four fit the work, and a written quote for the Founders Edition.
Talk to an expertWe reply within one business day