DGX Spark vs Mac Studio vs Ryzen AI Max: which desktop system for local LLMs
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- DGX Spark has 128 GB of unified memory at 273 GB/s with NVIDIA’s CUDA stack; systems with the Ryzen AI Max+ 395 have up to 128 GB at 256 GB/s with AMD’s ROCm on Linux; the Mac Studio offers up to 128 GB at up to 614 GB/s (M5 Max) or up to 512 GB at 1.2 TB/s (M5 Ultra) on macOS
- Token generation for one user follows memory bandwidth on all three; for a dense 70B model with 8-bit weights the arithmetic ceiling is about 3.9 tokens per second on DGX Spark, 3.6 on the Ryzen AI Max+ 395 and 17 on the M5 Ultra
- In one llama.cpp comparison with the same build on both machines (22 October 2025), gpt-oss 120B generated 45.9 tokens per second on DGX Spark and 47.5 on a Ryzen AI Max+ 395 system, while prompt processing ran at 1,737 and about 1,000 tokens per second
- AMD’s ROCm documentation puts the default GPU-mappable limit under Linux at about 50 per cent of RAM and raises it through the TTM limit; AMD’s own trillion-parameter test raised it to 120 GB per 128 GB system
- DGX Spark links up to four systems over 200 Gb/s ConnectX-7 and keeps code on the CUDA stack used by NVIDIA GPU servers; the Mac Studio clusters over Thunderbolt 5 with RDMA, and AMD’s cluster test used 5 Gb/s Ethernet
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
DGX Spark vs Mac Studio vs Ryzen AI Max: what decides
For local LLM work the three systems differ on four points: how much memory a model can use, how fast that memory is read, which software stack runs on the GPU, and how the machine fits into company IT. DGX Spark has 128 GB of unified memory at 273 GB/s and runs NVIDIA’s CUDA stack on DGX OS. The Mac Studio with M5 Ultra offers up to 512 GB at 1.2 TB/s on macOS with Apple’s Metal and MLX, and the M5 Max version up to 128 GB at up to 614 GB/s. Systems with the AMD Ryzen AI Max+ 395, known by its codename Strix Halo, have up to 128 GB at 256 GB/s and run AMD’s ROCm on Linux.
Specifications side by side
| SPECIFICATION | DGX SPARK | MAC STUDIO M5 MAX | MAC STUDIO M5 ULTRA | RYZEN AI MAX+ 395 |
|---|---|---|---|---|
| Unified memory | 128 GB LPDDR5x (Founders Edition) | 36 GB; 48, 64 or 128 GB with the 40-core GPU | 96 GB; 256 or 512 GB with the 80-core GPU | up to 128 GB, 256-bit LPDDR5x-8000 |
| Memory bandwidth | 273 GB/s | 460 GB/s; 614 GB/s with the 40-core GPU | 1.2 TB/s | 256 GB/s |
| GPU | Blackwell, 6,144 CUDA cores | 32 or 40 GPU cores | 64 or 80 GPU cores | Radeon 8060S, 40 compute units |
| GPU software | CUDA | Metal, MLX | Metal, MLX | ROCm |
| Network | ConnectX-7 at 200 Gbps, 10 GbE | 10 Gb Ethernet, four Thunderbolt 5 ports | 10 Gb Ethernet, six Thunderbolt 5 ports | set by the system maker |
| Power | 240 W supply, GB10 140 W TDP | 480 W maximum continuous | 480 W maximum continuous | 55 W default TDP, 45 to 120 W configurable |
| Operating system | NVIDIA DGX OS | macOS | macOS | Ubuntu and RHEL among those AMD lists |
NVIDIA DGX Spark product page, with the CUDA core count from NVIDIA’s hardware overview; Apple Mac Studio technical specifications; AMD Ryzen AI Max+ 395 product page; AMD Ryzen AI Halo user guide for the 256 GB/s figure; all read on 9 October 2026.
The Ryzen AI Max+ 395 is a processor that several manufacturers build into their own systems, so storage, ports and chassis depend on the system. AMD also sells a box of its own, the Ryzen AI Halo, which its user guide specifies at 150 × 150 × 45.4 mm with 128 GB of LPDDR5x at 256 GB/s, a 2 TB self-encrypting SSD, one 10 Gb Ethernet port and a 120 W TDP. AMD said on 20 May 2026 that it runs models of up to 200 billion parameters. On 28 September 2026 AMD said that systems with its Ryzen AI Max and Max PRO 400 Series are available from its partners, and that the 400 Series offers “up to 192GB, with up to 160GB dedicated to the GPU”. The AMD pages we read give no bandwidth figure for the 400 Series.
How much of the memory the GPU can use
On DGX Spark the CPU, DGX OS and the GPU share one pool with no fixed GPU share apart from a display reserve. The system reports about 119 to 122 GiB of the 128 GB, and NVIDIA’s playbook settings give a working set for weights and KV cache of roughly 102 to 115 GB, as our article on what fits in 128 GB on a DGX Spark works out.
Apple’s specification page gives the unified memory as one figure per configuration and states no separate GPU share. Metal reports a value per GPU, recommendedMaxWorkingSetSize, which Apple’s developer documentation describes as “an approximation of how much memory, in bytes, this GPU device can allocate without affecting its runtime performance”. We found no Apple figure for it per Mac Studio configuration, so read it on the machine before sizing a model close to the installed memory.
On the Ryzen AI Max+ 395 the split between CPU and GPU is a setting. AMD’s Ryzen AI Halo user guide describes Variable Graphics Memory with the shared share adjustable from 10 to 90 per cent in 5 per cent steps, 75 per cent by default, and warns that higher values leave less memory for other applications. Under Linux, AMD’s ROCm documentation puts the default GPU-mappable limit (GTT) at about 50 per cent of system RAM and recommends a small dedicated reservation in the firmware, for example 0.5 GB, with the shared TTM limit raised instead. In AMD’s own cluster test of February 2026 the firmware maximum was 96 GB per system, and kernel parameters raised the GPU’s share to 120 GB of each 128 GB system.
Memory bandwidth and tokens per second
Apple’s machine learning research team, writing about its M5 chip, states the rule for token generation: “Generating subsequent tokens is bounded by memory bandwidth, rather than by compute ability.” The same holds on DGX Spark and the Ryzen AI Max+ 395. For a dense model, each new token reads every weight once, so bandwidth divided by the size of the weights gives an upper bound for one user.
| SYSTEM | BANDWIDTH | DENSE 70B, 8-BIT |
|---|---|---|
| DGX Spark | 273 GB/s | about 3.9 tokens/s |
| Ryzen AI Max+ 395 | 256 GB/s | about 3.6 tokens/s |
| Mac Studio M5 Max, 40 cores | 614 GB/s | about 8.7 tokens/s |
| Mac Studio M5 Ultra | 1.2 TB/s | about 17 tokens/s |
| RTX PRO 6000 Workstation | 1,792 GB/s | about 25 tokens/s |
Bandwidth from NVIDIA, Apple and AMD pages read on 9 October 2026; the ceilings are our arithmetic, bandwidth divided by about 70.6 GB of weights, as in our other DGX Spark articles. Measured rates stay below them: LMSYS measured 2.7 tokens/s for Llama 3.1 70B in FP8 on DGX Spark with SGLang in October 2025.
A mixture-of-experts model reads only its active experts per token and runs far faster than its total size suggests. On 22 October 2025 a contributor to the llama.cpp project’s DGX Spark discussion ran gpt-oss 120B in MXFP4 with llama-bench, the same llama.cpp build on both machines, a 2,048-token prompt, 32 generated tokens and flash attention on. DGX Spark generated 45.87 tokens per second and a Ryzen AI Max+ 395 system with ROCm 47.49, close together, as the similar bandwidth figures suggest. On 2 November 2025 the same contributor reported 60.57 tokens per second on DGX Spark with a newer llama.cpp build and NVIDIA’s 6.17.1 kernel, with no new run on the AMD system; NVIDIA’s own llama.cpp figure, in its performance blog of 24 October 2025, is 55.37. Our DGX Spark benchmark article collects the other published runs.
Prompt processing depends on compute. In the 22 October run, DGX Spark processed 1,737 prompt tokens per second against about 1,000 on the Ryzen AI Max+ 395. Apple says the M5 Ultra processes LLM prompts in LM Studio up to four times faster than the M3 Ultra, from its own testing in July 2026; Apple does not name the model, and we found no third-party measurement of the M5 Mac Studio with model, engine and date that we could cite.
Where several people need a 70B model at reading speed, the step up is an RTX PRO 6000 workstation with 96 GB per card. Tell us the model and the number of users, and we compare it with DGX Spark on your numbers.
Software stack: CUDA, MLX or ROCm
DGX Spark runs DGX OS, an Ubuntu-based Linux for Arm, with the same CUDA stack as NVIDIA’s data-centre GPUs. NVIDIA publishes playbooks for vLLM, TensorRT-LLM and NIM on it, and the same engines run on RTX PRO 6000 and H200 NVL servers. FP4 arithmetic needs a Blackwell GPU such as the RTX PRO 6000; on the H200 NVL, FP4 weights load but compute in 16-bit. Our comparison of vLLM, SGLang, TensorRT-LLM, llama.cpp and Ollama lists which formats each engine runs on DGX Spark. The Arm processor means that any binary-only x86 dependency needs an Arm build.
The Mac Studio runs macOS, with Metal for the GPU. Apple describes MLX as its open-source machine learning framework for Apple silicon, which “enables developers to run, train, and fine-tune models”. llama.cpp and LM Studio also run on the Mac. A team that uses MLX and later serves its models with vLLM or TensorRT-LLM on Linux GPU servers changes engine, and often quantisation format, at that point; GGUF files for llama.cpp run on both.
Ryzen AI Max+ 395 systems run ROCm, AMD’s GPU software stack, on Linux. The Ryzen AI Halo ships with ROCm components, PyTorch, Lemonade and ComfyUI preinstalled. AMD’s ROCm documentation requires Linux kernel 6.18.4 or later for the Ryzen AI Max series, or a distribution that carries the fixes, such as Fedora 43 or Ubuntu 26.04, so check the combination before rolling out several machines.
Linking two to four systems
DGX Spark links over its ConnectX-7 ports at 200 Gb/s per port, and NVIDIA’s clustering guide supports up to three systems by direct cable and four through a switch. NVIDIA’s product page gives up to 400 billion parameters on two 128 GB systems and up to 700 billion on four; our two-node DGX Spark article covers the cabling.
Apple says several Mac Studio systems cluster through Thunderbolt 5 with RDMA, and that a cluster of four runs AI inference up to three times faster than one system, from its July 2026 testing; Apple does not name the model. MLX documents JACCL as its backend for “RDMA over thunderbolt”, which it calls necessary “for things like tensor parallelism”.
AMD published a four-system test on 25 February 2026: four Framework Desktop systems with the Ryzen AI Max+ 395, joined by 5 Gb/s Ethernet and running llama.cpp over RPC. Kimi K2.5 in a 2-bit quantisation (UD_Q2_K_XL), a 375 GB file, generated 9.45 tokens per second for 128 tokens, and a 4,096-token prompt took 39.7 seconds to the first token, both with flash attention.
Running one in a company
NVIDIA lists no BMC for DGX Spark, which is managed through its operating system and NVIDIA Sync; when it hangs, someone next to it resets it, or a switchable power outlet power-cycles it. On the Ryzen AI Halo, SSH is on by default under Linux, and AMD’s Developer Center offers updates, a factory reset, snapshot rollback and BIOS updates. A Mac Studio is managed like any other Mac in the company.
All three run from an office socket: DGX Spark from a 240 W supply, the Mac Studio at up to 480 W of continuous power by Apple’s specification, and the Ryzen AI Halo with a 120 W TDP. Apple sells the Mac Studio through its own stores and Apple Authorized Resellers. Ryzen AI Max systems come from AMD’s partners, each under its own warranty terms, and AMD named Micro Center, a US retailer, as the exclusive retailer for the Ryzen AI Halo. We supply the DGX Spark Founders Edition on one EU contract and invoice, with manufacturer warranty.
Which system for which need
| NEED | FITS | WHY |
|---|---|---|
| Development for CUDA servers | DGX Spark | same CUDA stack, playbooks for vLLM, TensorRT-LLM and NIM |
| More speed, up to 128 GB | Mac Studio M5 Max or Ultra | 460 GB/s to 1.2 TB/s against 256 to 273 GB/s |
| Above 128 GB in one system | Mac Studio M5 Ultra | 256 or 512 GB in one system |
| Model above 128 GB, CUDA | two DGX Spark | 256 GB over one QSFP cable, up to 400B parameters per NVIDIA |
| Linux with AMD software | Ryzen AI Max+ 395 system | ROCm on Linux, 128 GB at 256 GB/s |
| Team use at reading speed | RTX PRO 6000 workstation | 96 GB at 1,792 GB/s per card, up to four cards |
Our reading of the NVIDIA, Apple and AMD pages cited above as of 9 October 2026, and of our articles on DGX Spark and the RTX PRO 6000.
A team that uses one model every day at reading speed is better served by one to four RTX PRO 6000 cards, which our four-card Max-Q workstation article covers.
We supply DGX Spark, and RTX PRO 6000 workstations built to order; we do not supply Mac Studio or Ryzen AI Max systems. Describe your models and users in the form below, and we say which of the two fits.
What we supply
We supply the NVIDIA DGX Spark Founders Edition (128 GB, 4 TB) across the EU on one contract and invoice, with manufacturer warranty, for example two DGX Spark for a model above 128 GB or four for models of up to 700 billion parameters by NVIDIA’s figure. We also supply the RTX PRO 6000 Blackwell in its Workstation, Max-Q and Server editions, as cards or in workstations and AI servers built to order, assembled and burn-in tested. With the model, its precision, the context length and the number of concurrent users, we propose a configuration and quote within one business day.
FAQ
DGX Spark vs Mac Studio: which is better for local LLMs?
DGX Spark vs Ryzen AI Max+ 395 (Strix Halo): how do they compare?
Is the Ryzen AI Max a DGX Spark alternative?
Is a Mac Studio good for running LLMs?
How much memory can the GPU use on a Ryzen AI Max+ 395?
Which local LLM desktop generates tokens faster for one user?
Send us the models you want to run, their precision, the context length, how many people will use them at once and where the models will run in production. We reply within one business day with a configuration (DGX Spark or an RTX PRO 6000 workstation) and a written quote.
Talk to an expertWe reply within one business day