NVIDIA DGX Spark,
supplied in Europe
A 1.2 kg desktop machine with 128 GB of memory that the CPU and the GPU share. It runs models that will not fit on a workstation card, and it is honest about speed. One unit or several, on one European contract, invoiced from Vienna.
Every figure below is from NVIDIA’s own product page, user guide or quick start guide. We do not publish prices, because configuration and licences change the number; you get a written quote instead.
| Specification | DGX Spark | What it means in practice |
|---|---|---|
| Superchip | GB10 Grace Blackwell: 20 Arm cores (10 × Cortex-X925, 10 × Cortex-A725) and a Blackwell GPU with 5th-generation Tensor Cores | One chip, not a CPU plus a card. There is no PCIe slot to fill and nothing to configure |
| Memory | 128 GB LPDDR5X, coherent and unified, 256-bit, 273 GB/s | CPU and GPU share one pool. There is no separate VRAM figure to compare with a graphics card |
| Compute | Up to 1 PFLOP at FP4 | NVIDIA’s own footnote says this figure uses the sparsity feature. It publishes no dense number |
| Storage | 1 TB or 4 TB self-encrypting NVMe M.2 | Model weights are large. The 4 TB option is the one most people end up wanting |
| Network | ConnectX-7 with two QSFP ports at 200 Gb/s each, plus 10 GbE RJ-45 and Wi-Fi 7 | The QSFP ports are how two or more units are joined, not an afterthought |
| Power | 240 W supply; NVIDIA’s own declaration measures 233.2 W maximum and 38.0 W idle at the wall | It plugs into a desk socket. No rack, no three-phase feed, no room survey |
| Size and weight | 150 × 150 × 50.5 mm, 1.2 kg | Smaller than most mini PCs. Operating range is 5 to 30 °C |
| Software | NVIDIA DGX OS, CUDA, the DGX Spark playbooks | The same CUDA you already write for. Arm, so check any binary-only dependency |
Available to order in one unit or several. Lead time is confirmed with every quote.
One Spark is a fine order
You do not need a project, a licence bundle or an engineering contract to buy from us. One unit for one developer, or a set of four for a team, any quantity: we quote exactly what you ask for, at a competitive price, invoiced in the EU. The solutions are here when you want them, not a condition of sale.
One counterparty in Vienna. No cross-border purchase to arrange, no import to clear on your side.
Where a card needs a licence to do the job you bought it for, it is in the quote, not discovered three months later.
Our partner Vixen.UNO designs, deploys and supports it. Under the same contract, or not at all – your call.
Manufacturer warranty handled through us, with support alongside the team who runs the system.
NVIDIA says inference on models up to 200 billion parameters. Do the arithmetic before you plan around it: 200 billion parameters at half a byte each is 100 GB, so that claim is a 4-bit claim. In BF16 the same box holds roughly 55 billion. Both are true; they are not the same sentence.
Read the full analysis →Every token has to read the weights through 273 GB/s. A dense 70B in FP8 was measured at 2.7 tokens per second for a single user. For a solo developer iterating on prompts that is fine. For twenty people on a chat assistant it is not.
Read the full analysis →The common belief is that Spark tops out at a pair. NVIDIA’s own user guide says up to three connected directly by cable, and up to four through a switch. Measured throughput on a direct 200 GbE link is about 190 Gb/s across both ports.
Read the full analysis →How much of the 128 GB can a model actually use?
Can it really run a 200-billion-parameter model?
What can I fine-tune on one unit?
Can I connect more than two?
Is NVIDIA AI Enterprise included?
Who invoices, and what about warranty?
The arithmetic behind the 200B claim, what each precision costs, and what you can fine-tune on one unit.
Read the article →Measured token rates, the 273 GB/s ceiling, and the gap between the slides and what owners see.
Read the article →Three ways to start, and the four questions that decide which one is yours.
Read the article →We reply within one business day