NVIDIA DGX SPARK · SUPPLY IN THE EU

NVIDIA DGX Spark,
supplied in Europe

A 1.2 kg desktop machine with 128 GB of memory that the CPU and the GPU share. It runs models that will not fit on a workstation card, and it is honest about speed. One unit or several, on one European contract, invoiced from Vienna.

AVAILABLE TO ORDER
What is actually in the box

Every figure below is from NVIDIA’s own product page, user guide or quick start guide. We do not publish prices, because configuration and licences change the number; you get a written quote instead.

SpecificationDGX SparkWhat it means in practice
SuperchipGB10 Grace Blackwell: 20 Arm cores (10 × Cortex-X925, 10 × Cortex-A725) and a Blackwell GPU with 5th-generation Tensor CoresOne chip, not a CPU plus a card. There is no PCIe slot to fill and nothing to configure
Memory128 GB LPDDR5X, coherent and unified, 256-bit, 273 GB/sCPU and GPU share one pool. There is no separate VRAM figure to compare with a graphics card
ComputeUp to 1 PFLOP at FP4NVIDIA’s own footnote says this figure uses the sparsity feature. It publishes no dense number
Storage1 TB or 4 TB self-encrypting NVMe M.2Model weights are large. The 4 TB option is the one most people end up wanting
NetworkConnectX-7 with two QSFP ports at 200 Gb/s each, plus 10 GbE RJ-45 and Wi-Fi 7The QSFP ports are how two or more units are joined, not an afterthought
Power240 W supply; NVIDIA’s own declaration measures 233.2 W maximum and 38.0 W idle at the wallIt plugs into a desk socket. No rack, no three-phase feed, no room survey
Size and weight150 × 150 × 50.5 mm, 1.2 kgSmaller than most mini PCs. Operating range is 5 to 30 °C
SoftwareNVIDIA DGX OS, CUDA, the DGX Spark playbooksThe same CUDA you already write for. Arm, so check any binary-only dependency

Available to order in one unit or several. Lead time is confirmed with every quote.

JUST THE HARDWARE?

One Spark is a fine order

You do not need a project, a licence bundle or an engineering contract to buy from us. One unit for one developer, or a set of four for a team, any quantity: we quote exactly what you ask for, at a competitive price, invoiced in the EU. The solutions are here when you want them, not a condition of sale.

WHAT COMES WITH IT
Not just a box on a pallet
EU contract and invoice

One counterparty in Vienna. No cross-border purchase to arrange, no import to clear on your side.

Licences sized with the hardware

Where a card needs a licence to do the job you bought it for, it is in the quote, not discovered three months later.

Engineering, if you want it

Our partner Vixen.UNO designs, deploys and supports it. Under the same contract, or not at all – your call.

Warranty and support

Manufacturer warranty handled through us, with support alongside the team who runs the system.

HOW TO CHOOSE
Which one is right for you
What actually fits in 128 GB

NVIDIA says inference on models up to 200 billion parameters. Do the arithmetic before you plan around it: 200 billion parameters at half a byte each is 100 GB, so that claim is a 4-bit claim. In BF16 the same box holds roughly 55 billion. Both are true; they are not the same sentence.

Read the full analysis →
Capacity is real, throughput has a ceiling

Every token has to read the weights through 273 GB/s. A dense 70B in FP8 was measured at 2.7 tokens per second for a single user. For a solo developer iterating on prompts that is fine. For twenty people on a chat assistant it is not.

Read the full analysis →
More than two units is supported

The common belief is that Spark tops out at a pair. NVIDIA’s own user guide says up to three connected directly by cable, and up to four through a switch. Measured throughput on a direct 200 GbE link is about 190 Gb/s across both ports.

Read the full analysis →
FAQ
DGX Spark, in short
How much of the 128 GB can a model actually use?
NVIDIA publishes no usable-memory figure, and on this platform nvidia-smi cannot report GPU memory use at all. The honest answer comes from NVIDIA’s own playbooks, which default to 0.8 of the pool for vLLM and 0.8 to 0.9 for TensorRT-LLM. Plan on roughly 100 to 115 GB for weights and KV cache together, and leave the rest to the operating system.
Can it really run a 200-billion-parameter model?
Yes, at 4-bit. NVIDIA states inference on models up to 200 billion parameters, and its own validated list includes GPT-OSS-120B and Llama-3.3-70B in 4-bit formats. A 235B mixture-of-experts model is marked multi-node in the same table, which tells you where the single-unit line sits.
What can I fine-tune on one unit?
NVIDIA states up to 70 billion parameters and its fine-tuning playbook shows which method gets there: a full fine-tune on a 3B model, LoRA on 8B, QLoRA on 70B. The scripts are published under those exact names, so the ladder is not our interpretation.
Can I connect more than two?
Up to three directly through QSFP cables, and up to four through a switch, per NVIDIA’s clustering guide. Cable part numbers are specified, so this is a supported configuration rather than a hack.
Is NVIDIA AI Enterprise included?
Not automatically. An entitlement exists only if you bought it, requested an evaluation or received an entitlement certificate. If you need it, we put it in the quote alongside the hardware so it is not a surprise three months later.
Who invoices, and what about warranty?
Eurokommerz in Vienna. One EU counterparty, no import for you to clear, manufacturer warranty handled through us.
Check price and availability

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry – see our privacy policy.

request@eurokommerz.at  ·  +43 1 585 1405 50  ·  Jordangasse 7, 1010 Vienna