BLOG · GUIDE ·

Fine-tuning on DGX Spark: LoRA, QLoRA and full fine-tuning in 128 GB, and when to move to a server

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • NVIDIA states fine-tuning of models up to 70 billion parameters on DGX Spark; its PyTorch playbook reaches that size on one Spark only with QLoRA, and shows a full fine-tune at Llama 3.2 3B and LoRA at Llama 3.1 8B
  • For two DGX Spark, the same playbook runs LoRA on Llama 3.1 70B with FSDP2; on a BF16 base that run holds about 144 GB of model states by our arithmetic, more than one Spark has, or about 72 GB per Spark when sharded
  • NVIDIA’s NeMo playbook, updated on 4 September 2026, covers models of about 1 to 70B parameters with recipes validated for Spark: LoRA on Llama 3.1 8B and Qwen3 8B, QLoRA on Llama 3.3 70B Instruct
  • Meta’s PyTorch team needed about 8 hours per epoch for a full BF16 fine-tune of Llama 3.1 8B at 16K tokens, and NVIDIA measured 759.79 tokens per second for QLoRA on a 70B model, about 22 minutes per million training tokens
  • A GPU server is the next step for repeated runs on large datasets, LoRA on a BF16 base above 32B and full fine-tunes with 32-bit Adam: the H200 NVL holds the 128 GB of states of an 8B full fine-tune on one card

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

What you can fine-tune on one DGX Spark and on two

DGX Spark fine-tunes models of up to 70 billion parameters by NVIDIA’s product page as read on 9 October 2026, and the method decides how large the model can be. NVIDIA’s PyTorch fine-tuning playbook, as read on 9 October 2026 in NVIDIA’s GitHub repository, ships one script per step: a full fine-tune of Llama 3.2 3B, LoRA on Llama 3.1 8B and QLoRA on Llama 3.1 70B on one Spark. Under the heading “Run on two Sparks” it runs LoRA on the 70B model with FSDP2, sharded across two DGX Spark, and its configuration files for two systems also cover a full fine-tune of the 3B model and LoRA on the 8B model.

NVIDIA’s NeMo playbook, updated on 4 September 2026, gives the same range, models of about 1 to 70B parameters, and runs three recipes on one Spark. Meta’s PyTorch team showed a fourth case in February 2026, a full fine-tune of Llama 3.1 8B in BF16. Unsloth goes further in its own DGX Spark guide and writes that its software “enables local fine-tuning of LLMs with up to 200B parameters”.

METHODONE SPARKTWO DGX SPARKSOURCE
Full fine-tuningLlama 3.2 3B; Llama 3.1 8B in BF16Llama 3.2 3B, configuration fileNVIDIA PyTorch playbook; PyTorch blog, February 2026
LoRALlama 3.1 8B, Qwen3 8BLlama 3.1 8B and 70B with FSDP2NVIDIA PyTorch and NeMo playbooks
QLoRA, 4-bit baseLlama 3.1 70B, Llama 3.3 70B Instruct; gpt-oss-120b at about 68 GBno example publishedNVIDIA PyTorch and NeMo playbooks; Unsloth guide

NVIDIA playbooks at build.nvidia.com/spark and their READMEs in NVIDIA’s dgx-spark-playbooks repository, read on 9 October 2026; PyTorch blog of 2 February 2026; Unsloth’s DGX Spark guide (undated).

On one Spark, NVIDIA reaches 70B only with QLoRA. Before you plan a fine-tune at all, check whether the problem is missing knowledge or behaviour, since retrieval fixes the first more reliably; our comparison of fine-tuning and RAG sets out the evidence.

The fine-tuning playbooks: PyTorch, NeMo, Unsloth and LLaMA Factory

NVIDIA’s catalogue at build.nvidia.com/spark lists fine-tuning playbooks for NeMo, Unsloth, LLaMA Factory, vision-language models and FLUX.1 image models. The PyTorch playbook was not in that list on 9 October 2026, and its README remains in NVIDIA’s dgx-spark-playbooks repository on GitHub. Each playbook installs a framework and runs an example job that you then point at your own dataset.

The PyTorch playbook gives NVIDIA’s own ladder. Its scripts are named after the method, from Llama3_3B_full_finetuning.py to Llama3_70B_qLoRA_finetuning.py. The defaults are BF16 precision, a LoRA rank of 8 and a maximum sequence length of 2,048 tokens, and NVIDIA estimates 30 to 45 minutes for setup and a first run. The README carries a last-updated date of 15 January 2025, which is earlier than the product’s launch, so we cite it as read on 9 October 2026.

The NeMo playbook uses NeMo AutoModel 26.08 in the container nvcr.io/nvidia/nemo-automodel:26.08, with what NVIDIA calls “Spark-specific LoRA and QLoRA recipes”. Its three examples run LoRA on Llama 3.1 8B and Qwen3 8B and QLoRA on Llama 3.3 70B Instruct, 20 steps each. Of the 70B recipe NVIDIA writes that it “loads it in 4-bit mode, and uses the validated Spark memory settings”. The playbook asks you to keep the batch size, packed sequence size, attention, activation checkpointing and checkpoint-loading settings of these recipes unless you have validated others. NVIDIA estimates 45 to 90 minutes for setup and a first run.

The Unsloth playbook, last updated on 15 December 2025, covers LoRA and QLoRA. Its example loads unsloth/Meta-Llama-3.1-8B-bnb-4bit, a 4-bit 8B model, its custom-run example sets a batch size of 4, and the test run trains for 60 steps; NVIDIA gives 30 to 60 minutes for setup and that run. Unsloth’s own guide states that “gpt-oss-120b QLoRA 4-bit fine-tuning will use around 68GB of unified memory”.

The LLaMA Factory playbook, last updated on 31 July 2026, offers LoRA, QLoRA and full fine-tuning and uses a LoRA run on Qwen3 as its example. NVIDIA asks you to plan more than 50 GB of storage for models and checkpoints and gives 1 to 7 hours of training, “depending on model size and dataset”.

Memory per method in 128 GB

The model states follow the same arithmetic as our GPU memory guide for full fine-tuning, LoRA and QLoRA: 16 bytes per parameter for mixed-precision Adam, 2 bytes per frozen BF16 weight and about 0.516 bytes per weight in 4-bit NF4 with double quantisation. Activations come on top and grow with batch size and sequence length. The system shows about 119 to 122 GiB of the 128 GB, and the operating system and your other processes share that pool; what fits in 128 GB on a DGX Spark explains why no tool reports a VRAM figure on this machine.

Against that pool, a full fine-tune of a 3B model needs about 48 GB of model states. Llama 3.1 8B with 32-bit Adam needs 128 GB, about all of the memory the system shows, before a single activation. LoRA at rank 16 on all linear layers needs 16.7 GB for the 8B model, so most of the memory is left for the batch. LoRA on a BF16 base of a 70B Llama model needs 144.4 GB, which is consistent with NVIDIA running that example on two DGX Spark. Sharded with FSDP, it comes to about 72.2 GB per Spark, plus activations and the parameters being gathered. QLoRA cuts the same 70B model to about 42.8 GB, which leaves about 85 to 88 GB of the visible memory for activations, the operating system and everything else. NVIDIA’s PyTorch scripts default to rank 8, which lowers these figures by less than 2 GB.

Two notes in NVIDIA’s playbooks concern this platform. NVIDIA’s Unsloth README notes that DGX Spark “uses a Unified Memory Architecture (UMA)” and gives sync; echo 3 > /proc/sys/vm/drop_caches, run as root, for memory errors that appear below the capacity. The NeMo playbook starts its container with a 64 GB memory limit, --memory=64g, in NVIDIA’s words “the 64 GB limit used for the lower-memory Spark validation”.

We supply the 128 GB DGX Spark Founders Edition across the EU. Tell us the base model, your dataset size and the method in the form below, and we reply within one business day with whether one Spark or two fits the run.

Full fine-tuning of an 8B model on one Spark

Meta’s PyTorch team published a full fine-tune of Llama 3.1 8B Instruct on one DGX Spark on 2 February 2026, using torchtune’s single-device recipe. The run used 11k conversation pairs from the ToolACE dataset with synthetic reasoning traces, a sequence length of 16,384 tokens, a batch size of 16, 3 epochs and BF16. The team reports a peak of “roughly 80% memory usage” and “each epoch averaging just around 8 hours”.

The blog does not name the optimiser and runs its own configuration file, fft-8b.yaml. The single-device configuration that torchtune ships for Llama 3.1 8B uses the 8-bit PagedAdamW8bit optimiser from bitsandbytes, runs the optimiser step in the backward pass and turns on activation checkpointing, all marked in the file as memory savers. With FP32 master weights and Adam moments, the 16-byte count of mixed precision, the same model needs 128 GB before activations. If your training recipe expects 32-bit optimiser states, plan the 8B full fine-tune for a GPU server.

Training times from published runs

NVIDIA’s performance blog of 24 October 2025 reports peak training throughput on one Spark at 2,048-token sequences, one epoch and 64 steps. Its table gives 13,519.54 tokens per second for the 3B full fine-tune, 6,969.59 for the 8B LoRA and 759.79 for QLoRA on Llama 3.3 70B. The blog’s running text gives higher peaks for the same three runs, for example 5,079.4 tokens per second for the 70B QLoRA. We use the table, which states the configuration. NVIDIA adds that “none of these tuning workloads can run on a 32 GB consumer GPU”.

At 759.79 tokens per second, a million training tokens on the 70B QLoRA run takes about 22 minutes by our arithmetic. Other named sources give wall-clock times. Unsloth reports “1,000 steps and 4 hours of RL training” for gpt-oss-20b on DGX Spark, a reinforcement learning example that teaches the model the game 2048. Meta’s 8B full fine-tune took about 8 hours per epoch at 16K tokens. NVIDIA’s LLaMA Factory estimate is 1 to 7 hours. The two-node article lists published DGX Spark fine-tuning and multi-node figures, including Unsloth’s memory figures for a 70B LoRA across two DGX Spark.

What limits DGX Spark for training

The large matrix multiplications of training are compute-bound, as our H200 NVL and RTX PRO 6000 comparison for fine-tuning works out, so the tensor rate sets most of the pace. For DGX Spark, NVIDIA publishes up to 1 PFLOP at FP4, a figure that its datasheet footnote says uses sparsity, and we found no dense BF16 figure from NVIDIA. For the H200 NVL, NVIDIA gives 1,671 BF16 TFLOPS with sparsity, 835.5 dense when halved. Its RTX PRO whitepaper gives 503.8 dense BF16 TFLOPS for the RTX PRO 6000 Workstation Edition; for the Server Edition we found no dense figure from NVIDIA.

Memory bandwidth differs more: DGX Spark has 273 GB/s against 1,597 GB/s on the RTX PRO 6000 Server Edition and 4.8 TB/s on the H200 NVL. By the reasoning of the same comparison, the optimiser step, the adapter multiplications and, in QLoRA, the dequantisation of every 4-bit weight are limited by bandwidth. The measured DGX Spark benchmarks show the effect of the 273 GB/s on generation. A Spark also runs one heavy job at a time, since a training run and a serving engine share the same 128 GB and the same GPU.

Two DGX Spark are joined by ConnectX-7 at 200 Gb/s per port, and NVIDIA’s own RDMA test measures 189.85 Gb/s. By our arithmetic that is about 24 GB/s, against 450 GB/s each way on a two-way NVLink bridge between H200 NVL cards, and sharded runs exchange their parameters over that link.

DGX Spark or a GPU server for fine-tuning

DGX Spark suits a developer who prepares data, tests recipes and runs QLoRA up to 70B where a run of hours or a night is acceptable. NVIDIA’s product page describes it for AI development and testing at the desktop. A GPU server suits a team that retrains on schedule, runs several experiments at once or needs LoRA on a BF16 base.

WORKLOADONE SPARKTWO DGX SPARKGPU SERVER
LoRA, 8B, test runsfits, 16.7 GB of statesnot neededRTX PRO 6000 with 86 GB left
QLoRA, 70Bfits, 42.8 GB of statesno example publishedRTX PRO 6000 with 60 GB left
LoRA, 32B, BF16 base67.7 GB of states, no NVIDIA exampleno example publishedRTX PRO 6000 with 35 GB left
LoRA, 70B, BF16 basedoes not fit, 144.4 GBabout 72.2 GB per Sparktwo H200 NVL with NVLink
Full, 8B, 32-bit Adamdoes not fit, 128 GBno example publishedone H200 NVL, 22 GB left
Scheduled retrainingslow, one job at a timeslow, one job at a timeserver built to order

Model states without activations, rank-16 LoRA on all linear layers, by the arithmetic of our GPU memory guide; RTX PRO 6000 at 95.6 GiB and H200 NVL at 140.4 GiB as the driver reports them. Methods per platform from NVIDIA’s playbooks.

PyTorch, NeMo, Unsloth and LLaMA Factory also run on GPU servers, so a recipe tested on a Spark moves to the server with its frameworks. DGX Spark is an Arm system, so containers and any binary-only dependency have to be built or pulled for the server’s x86 processors. For two or more cards, the H200 NVL bridges 2 or 4 cards over NVLink at 900 GB/s per GPU, which suits sharded runs.

We build training and fine-tuning nodes to order with the RTX PRO 6000 Server Edition or the H200 NVL with NVLink bridges, and we check the rack, power and airflow before we quote. Send us your model, method and dataset size through the form below.

What we supply

We supply the DGX Spark Founders Edition with 128 GB, for one developer or as two DGX Spark for 70B LoRA runs, on one EU contract and invoice with manufacturer warranty. When fine-tuning moves to scheduled runs, we build AI servers to order with the RTX PRO 6000 Server Edition or the H200 NVL, with fast NVMe scratch storage, assembled and burn-in tested, with configuration and quote within one business day. If you also want the platform around the training, MLOps on Kubernetes with KServe and Kubeflow and private LLMs in your infrastructure are part of our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

Can you fine-tune an LLM on DGX Spark?
Yes. NVIDIA states fine-tuning of models up to 70 billion parameters and publishes playbooks for PyTorch, NeMo, Unsloth and LLaMA Factory. On one Spark, NVIDIA’s PyTorch playbook shows a full fine-tune of Llama 3.2 3B, LoRA on Llama 3.1 8B and QLoRA on Llama 3.1 70B.
What size model can I fine-tune on one DGX Spark?
Up to 70 billion parameters with QLoRA, which stores the frozen base in 4 bits at about 42.8 GB of model states for Llama 3.3 70B by our arithmetic. LoRA on a BF16 base reaches 8B in NVIDIA’s examples, and a full fine-tune 3B in NVIDIA’s playbook and 8B in BF16 in a run that Meta’s PyTorch team published in February 2026.
Does Unsloth work on DGX Spark?
Yes. NVIDIA publishes an Unsloth playbook for DGX Spark that covers LoRA and QLoRA and runs a test job on a 4-bit Llama 3.1 8B in 30 to 60 minutes including setup. Unsloth’s own guide states that QLoRA on gpt-oss-120b uses around 68 GB of unified memory.
Can DGX Spark run QLoRA on a 70B model?
Yes. Both NVIDIA’s PyTorch playbook and its NeMo playbook run QLoRA on a 70B Llama model on one Spark. NVIDIA measured 759.79 tokens per second for QLoRA on Llama 3.3 70B at 2,048-token sequences, which is about 22 minutes per million training tokens.
How long does training or fine-tuning take on DGX Spark?
It depends on the model, the method and the dataset. NVIDIA’s LLaMA Factory playbook estimates 1 to 7 hours of training, Unsloth reports 4 hours for 1,000 steps of reinforcement learning on gpt-oss-20b in a game-playing example, and Meta’s PyTorch team needed about 8 hours per epoch for a full fine-tune of Llama 3.1 8B at 16K-token sequences.
Can two DGX Spark fine-tune a 70B model with LoRA?
Yes. NVIDIA’s PyTorch playbook includes a two-Spark example that runs LoRA on Llama 3.1 70B with FSDP2. On a BF16 base the model states come to about 144 GB by our arithmetic, which exceeds one Spark, or about 72 GB per Spark when sharded across two.

Send us the base model, the method (full fine-tuning, LoRA or QLoRA), the size of your training data, the sequence length and how often you plan to retrain. We reply within one business day with the memory figures for your run, whether one Spark, two DGX Spark or a GPU server built to order fits it, and a written quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna