Fine-tuning on DGX Spark: LoRA, QLoRA and full fine-tuning in 128 GB, and when to move to a server
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- NVIDIA states fine-tuning of models up to 70 billion parameters on DGX Spark; its PyTorch playbook reaches that size on one Spark only with QLoRA, and shows a full fine-tune at Llama 3.2 3B and LoRA at Llama 3.1 8B
- For two DGX Spark, the same playbook runs LoRA on Llama 3.1 70B with FSDP2; on a BF16 base that run holds about 144 GB of model states by our arithmetic, more than one Spark has, or about 72 GB per Spark when sharded
- NVIDIA’s NeMo playbook, updated on 4 September 2026, covers models of about 1 to 70B parameters with recipes validated for Spark: LoRA on Llama 3.1 8B and Qwen3 8B, QLoRA on Llama 3.3 70B Instruct
- Meta’s PyTorch team needed about 8 hours per epoch for a full BF16 fine-tune of Llama 3.1 8B at 16K tokens, and NVIDIA measured 759.79 tokens per second for QLoRA on a 70B model, about 22 minutes per million training tokens
- A GPU server is the next step for repeated runs on large datasets, LoRA on a BF16 base above 32B and full fine-tunes with 32-bit Adam: the H200 NVL holds the 128 GB of states of an 8B full fine-tune on one card
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
What you can fine-tune on one DGX Spark and on two
DGX Spark fine-tunes models of up to 70 billion parameters by NVIDIA’s product page as read on 9 October 2026, and the method decides how large the model can be. NVIDIA’s PyTorch fine-tuning playbook, as read on 9 October 2026 in NVIDIA’s GitHub repository, ships one script per step: a full fine-tune of Llama 3.2 3B, LoRA on Llama 3.1 8B and QLoRA on Llama 3.1 70B on one Spark. Under the heading “Run on two Sparks” it runs LoRA on the 70B model with FSDP2, sharded across two DGX Spark, and its configuration files for two systems also cover a full fine-tune of the 3B model and LoRA on the 8B model.
NVIDIA’s NeMo playbook, updated on 4 September 2026, gives the same range, models of about 1 to 70B parameters, and runs three recipes on one Spark. Meta’s PyTorch team showed a fourth case in February 2026, a full fine-tune of Llama 3.1 8B in BF16. Unsloth goes further in its own DGX Spark guide and writes that its software “enables local fine-tuning of LLMs with up to 200B parameters”.
| METHOD | ONE SPARK | TWO DGX SPARK | SOURCE |
|---|---|---|---|
| Full fine-tuning | Llama 3.2 3B; Llama 3.1 8B in BF16 | Llama 3.2 3B, configuration file | NVIDIA PyTorch playbook; PyTorch blog, February 2026 |
| LoRA | Llama 3.1 8B, Qwen3 8B | Llama 3.1 8B and 70B with FSDP2 | NVIDIA PyTorch and NeMo playbooks |
| QLoRA, 4-bit base | Llama 3.1 70B, Llama 3.3 70B Instruct; gpt-oss-120b at about 68 GB | no example published | NVIDIA PyTorch and NeMo playbooks; Unsloth guide |
NVIDIA playbooks at build.nvidia.com/spark and their READMEs in NVIDIA’s dgx-spark-playbooks repository, read on 9 October 2026; PyTorch blog of 2 February 2026; Unsloth’s DGX Spark guide (undated).
On one Spark, NVIDIA reaches 70B only with QLoRA. Before you plan a fine-tune at all, check whether the problem is missing knowledge or behaviour, since retrieval fixes the first more reliably; our comparison of fine-tuning and RAG sets out the evidence.
The fine-tuning playbooks: PyTorch, NeMo, Unsloth and LLaMA Factory
NVIDIA’s catalogue at build.nvidia.com/spark lists fine-tuning playbooks for NeMo, Unsloth, LLaMA Factory, vision-language models and FLUX.1 image models. The PyTorch playbook was not in that list on 9 October 2026, and its README remains in NVIDIA’s dgx-spark-playbooks repository on GitHub. Each playbook installs a framework and runs an example job that you then point at your own dataset.
The PyTorch playbook gives NVIDIA’s own ladder. Its scripts are named after the method, from Llama3_3B_full_finetuning.py to Llama3_70B_qLoRA_finetuning.py. The defaults are BF16 precision, a LoRA rank of 8 and a maximum sequence length of 2,048 tokens, and NVIDIA estimates 30 to 45 minutes for setup and a first run. The README carries a last-updated date of 15 January 2025, which is earlier than the product’s launch, so we cite it as read on 9 October 2026.
The NeMo playbook uses NeMo AutoModel 26.08 in the container nvcr.io/nvidia/nemo-automodel:26.08, with what NVIDIA calls “Spark-specific LoRA and QLoRA recipes”. Its three examples run LoRA on Llama 3.1 8B and Qwen3 8B and QLoRA on Llama 3.3 70B Instruct, 20 steps each. Of the 70B recipe NVIDIA writes that it “loads it in 4-bit mode, and uses the validated Spark memory settings”. The playbook asks you to keep the batch size, packed sequence size, attention, activation checkpointing and checkpoint-loading settings of these recipes unless you have validated others. NVIDIA estimates 45 to 90 minutes for setup and a first run.
The Unsloth playbook, last updated on 15 December 2025, covers LoRA and QLoRA. Its example loads unsloth/Meta-Llama-3.1-8B-bnb-4bit, a 4-bit 8B model, its custom-run example sets a batch size of 4, and the test run trains for 60 steps; NVIDIA gives 30 to 60 minutes for setup and that run. Unsloth’s own guide states that “gpt-oss-120b QLoRA 4-bit fine-tuning will use around 68GB of unified memory”.
The LLaMA Factory playbook, last updated on 31 July 2026, offers LoRA, QLoRA and full fine-tuning and uses a LoRA run on Qwen3 as its example. NVIDIA asks you to plan more than 50 GB of storage for models and checkpoints and gives 1 to 7 hours of training, “depending on model size and dataset”.
Memory per method in 128 GB
The model states follow the same arithmetic as our GPU memory guide for full fine-tuning, LoRA and QLoRA: 16 bytes per parameter for mixed-precision Adam, 2 bytes per frozen BF16 weight and about 0.516 bytes per weight in 4-bit NF4 with double quantisation. Activations come on top and grow with batch size and sequence length. The system shows about 119 to 122 GiB of the 128 GB, and the operating system and your other processes share that pool; what fits in 128 GB on a DGX Spark explains why no tool reports a VRAM figure on this machine.
Against that pool, a full fine-tune of a 3B model needs about 48 GB of model states. Llama 3.1 8B with 32-bit Adam needs 128 GB, about all of the memory the system shows, before a single activation. LoRA at rank 16 on all linear layers needs 16.7 GB for the 8B model, so most of the memory is left for the batch. LoRA on a BF16 base of a 70B Llama model needs 144.4 GB, which is consistent with NVIDIA running that example on two DGX Spark. Sharded with FSDP, it comes to about 72.2 GB per Spark, plus activations and the parameters being gathered. QLoRA cuts the same 70B model to about 42.8 GB, which leaves about 85 to 88 GB of the visible memory for activations, the operating system and everything else. NVIDIA’s PyTorch scripts default to rank 8, which lowers these figures by less than 2 GB.
Two notes in NVIDIA’s playbooks concern this platform. NVIDIA’s Unsloth README notes that DGX Spark “uses a Unified Memory Architecture (UMA)” and gives sync; echo 3 > /proc/sys/vm/drop_, run as root, for memory errors that appear below the capacity. The NeMo playbook starts its container with a 64 GB memory limit, --memory=64g, in NVIDIA’s words “the 64 GB limit used for the lower-memory Spark validation”.
We supply the 128 GB DGX Spark Founders Edition across the EU. Tell us the base model, your dataset size and the method in the form below, and we reply within one business day with whether one Spark or two fits the run.
Full fine-tuning of an 8B model on one Spark
Meta’s PyTorch team published a full fine-tune of Llama 3.1 8B Instruct on one DGX Spark on 2 February 2026, using torchtune’s single-device recipe. The run used 11k conversation pairs from the ToolACE dataset with synthetic reasoning traces, a sequence length of 16,384 tokens, a batch size of 16, 3 epochs and BF16. The team reports a peak of “roughly 80% memory usage” and “each epoch averaging just around 8 hours”.
The blog does not name the optimiser and runs its own configuration file, fft-8b.yaml. The single-device configuration that torchtune ships for Llama 3.1 8B uses the 8-bit PagedAdamW8bit optimiser from bitsandbytes, runs the optimiser step in the backward pass and turns on activation checkpointing, all marked in the file as memory savers. With FP32 master weights and Adam moments, the 16-byte count of mixed precision, the same model needs 128 GB before activations. If your training recipe expects 32-bit optimiser states, plan the 8B full fine-tune for a GPU server.
Training times from published runs
NVIDIA’s performance blog of 24 October 2025 reports peak training throughput on one Spark at 2,048-token sequences, one epoch and 64 steps. Its table gives 13,519.54 tokens per second for the 3B full fine-tune, 6,969.59 for the 8B LoRA and 759.79 for QLoRA on Llama 3.3 70B. The blog’s running text gives higher peaks for the same three runs, for example 5,079.4 tokens per second for the 70B QLoRA. We use the table, which states the configuration. NVIDIA adds that “none of these tuning workloads can run on a 32 GB consumer GPU”.
At 759.79 tokens per second, a million training tokens on the 70B QLoRA run takes about 22 minutes by our arithmetic. Other named sources give wall-clock times. Unsloth reports “1,000 steps and 4 hours of RL training” for gpt-oss-20b on DGX Spark, a reinforcement learning example that teaches the model the game 2048. Meta’s 8B full fine-tune took about 8 hours per epoch at 16K tokens. NVIDIA’s LLaMA Factory estimate is 1 to 7 hours. The two-node article lists published DGX Spark fine-tuning and multi-node figures, including Unsloth’s memory figures for a 70B LoRA across two DGX Spark.
What limits DGX Spark for training
The large matrix multiplications of training are compute-bound, as our H200 NVL and RTX PRO 6000 comparison for fine-tuning works out, so the tensor rate sets most of the pace. For DGX Spark, NVIDIA publishes up to 1 PFLOP at FP4, a figure that its datasheet footnote says uses sparsity, and we found no dense BF16 figure from NVIDIA. For the H200 NVL, NVIDIA gives 1,671 BF16 TFLOPS with sparsity, 835.5 dense when halved. Its RTX PRO whitepaper gives 503.8 dense BF16 TFLOPS for the RTX PRO 6000 Workstation Edition; for the Server Edition we found no dense figure from NVIDIA.
Memory bandwidth differs more: DGX Spark has 273 GB/s against 1,597 GB/s on the RTX PRO 6000 Server Edition and 4.8 TB/s on the H200 NVL. By the reasoning of the same comparison, the optimiser step, the adapter multiplications and, in QLoRA, the dequantisation of every 4-bit weight are limited by bandwidth. The measured DGX Spark benchmarks show the effect of the 273 GB/s on generation. A Spark also runs one heavy job at a time, since a training run and a serving engine share the same 128 GB and the same GPU.
Two DGX Spark are joined by ConnectX-7 at 200 Gb/s per port, and NVIDIA’s own RDMA test measures 189.85 Gb/s. By our arithmetic that is about 24 GB/s, against 450 GB/s each way on a two-way NVLink bridge between H200 NVL cards, and sharded runs exchange their parameters over that link.
DGX Spark or a GPU server for fine-tuning
DGX Spark suits a developer who prepares data, tests recipes and runs QLoRA up to 70B where a run of hours or a night is acceptable. NVIDIA’s product page describes it for AI development and testing at the desktop. A GPU server suits a team that retrains on schedule, runs several experiments at once or needs LoRA on a BF16 base.
| WORKLOAD | ONE SPARK | TWO DGX SPARK | GPU SERVER |
|---|---|---|---|
| LoRA, 8B, test runs | fits, 16.7 GB of states | not needed | RTX PRO 6000 with 86 GB left |
| QLoRA, 70B | fits, 42.8 GB of states | no example published | RTX PRO 6000 with 60 GB left |
| LoRA, 32B, BF16 base | 67.7 GB of states, no NVIDIA example | no example published | RTX PRO 6000 with 35 GB left |
| LoRA, 70B, BF16 base | does not fit, 144.4 GB | about 72.2 GB per Spark | two H200 NVL with NVLink |
| Full, 8B, 32-bit Adam | does not fit, 128 GB | no example published | one H200 NVL, 22 GB left |
| Scheduled retraining | slow, one job at a time | slow, one job at a time | server built to order |
Model states without activations, rank-16 LoRA on all linear layers, by the arithmetic of our GPU memory guide; RTX PRO 6000 at 95.6 GiB and H200 NVL at 140.4 GiB as the driver reports them. Methods per platform from NVIDIA’s playbooks.
PyTorch, NeMo, Unsloth and LLaMA Factory also run on GPU servers, so a recipe tested on a Spark moves to the server with its frameworks. DGX Spark is an Arm system, so containers and any binary-only dependency have to be built or pulled for the server’s x86 processors. For two or more cards, the H200 NVL bridges 2 or 4 cards over NVLink at 900 GB/s per GPU, which suits sharded runs.
We build training and fine-tuning nodes to order with the RTX PRO 6000 Server Edition or the H200 NVL with NVLink bridges, and we check the rack, power and airflow before we quote. Send us your model, method and dataset size through the form below.
What we supply
We supply the DGX Spark Founders Edition with 128 GB, for one developer or as two DGX Spark for 70B LoRA runs, on one EU contract and invoice with manufacturer warranty. When fine-tuning moves to scheduled runs, we build AI servers to order with the RTX PRO 6000 Server Edition or the H200 NVL, with fast NVMe scratch storage, assembled and burn-in tested, with configuration and quote within one business day. If you also want the platform around the training, MLOps on Kubernetes with KServe and Kubeflow and private LLMs in your infrastructure are part of our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
Can you fine-tune an LLM on DGX Spark?
What size model can I fine-tune on one DGX Spark?
Does Unsloth work on DGX Spark?
Can DGX Spark run QLoRA on a 70B model?
How long does training or fine-tuning take on DGX Spark?
Can two DGX Spark fine-tune a 70B model with LoRA?
Send us the base model, the method (full fine-tuning, LoRA or QLoRA), the size of your training data, the sequence length and how often you plan to retrain. We reply within one business day with the memory figures for your run, whether one Spark, two DGX Spark or a GPU server built to order fits it, and a written quote.
Talk to an expertWe reply within one business day