BLOG · GUIDE ·

GPU server for a university or research lab: Slurm, MIG and fair sharing of the cards

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • A GPU server shared by a lab needs a scheduler (Slurm with GPUs as GRES, or Kubernetes with Kueue), partitioning so small jobs do not hold a whole card, and per-group quotas with fair-share
  • Slurm supports MIG since version 21.08 but expects the instances to exist already and does not partition GPUs on demand, so the MIG layout is a node setting the administrator changes
  • NVIDIA’s MIG guide (11 September 2026) lists up to 7 instances for the H200 NVL, the smallest 1g.18gb, and up to 4 for every RTX PRO 6000 edition, 1g.24gb each; the L40S and L4 are not on the list
  • Slurm’s fair-share counts allocated CPU seconds by default; a partition’s TRESBillingWeights adds GPUs, and limits such as GrpTRES and MaxTRESPerUser cap cards per group and per user
  • For double-precision codes the H200 NVL delivers 30 TFLOPS of FP64, while the RTX PRO 6000 runs FP64 at 1/64 of its FP32 rate, about 1.9 TFLOPS on the Server Edition by our arithmetic

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

What a GPU server for a university or research lab needs

Besides the cards, a GPU server that a research lab or a course shares needs a scheduler that queues jobs and hands out GPUs, partitioning so that small jobs run without holding a whole card, and quotas with fair-share so that one group cannot occupy the machine for weeks. On bare metal the usual scheduler is Slurm, which handles GPUs as generic resources (GRES) and accounts for them per user and account. On a Kubernetes cluster, Kueue adds quotas and queueing next to the normal scheduler. NVIDIA’s Multi-Instance GPU (MIG) splits an H200 NVL into up to seven instances and an RTX PRO 6000 into up to four, each with its own memory. Training and double-precision simulation lead to the H200 NVL, inference and teaching to the RTX PRO 6000 Server Edition.

Research use comes in bursts. A doctoral student fine-tunes for two days and then pauses for a week, and a course opens thirty notebooks in the same hour. Without a scheduler, whoever logs in first takes the cards, and an idle notebook keeps a 96 GB card to itself.

Slurm and GPUs: GRES configuration, job requests and accounting

Slurm’s GRES documentation (version 26.05 as of October 2026) says “You must explicitly specify which GRES are to be managed” in slurm.conf. GresTypes=gpu names the resource type, and each node line declares its cards, for example Gres=gpu:4. The device files and the CPU cores close to each card go into gres.conf, written by hand or filled in with AutoDetect=nvml, which in Slurm’s words supplies the details “for any system-detected GPU”. The count in slurm.conf stays required, so that the controller knows how many cards to expect.

Users ask for cards with --gpus or with --gres, whose form is name, optional type and count, as in the documentation’s --gres=gpu:kepler:2. Type names are set by the administrator, which lets a lab keep H200 NVL and RTX PRO 6000 nodes in one cluster and let jobs ask for one or the other.

With AccountingStorageTRES=gres/gpu set, Slurm also gathers GPU memory and utilisation per job, as gres/gpumem and gres/gpuutil, for NVIDIA GPUs found with AutoDetect=nvml. The documentation adds that NVML does not support utilisation metrics for MIG instances, so Slurm records neither gpumem nor gpuutil for them. Card-level health, memory errors and XID events come from NVIDIA’s DCGM, which our GPU server monitoring guide covers.

MIG in Slurm: splitting an H200 NVL or an RTX PRO 6000

Slurm has supported MIG devices since version 21.08. They are detected with AutoDetect=nvml and listed in slurm.conf like ordinary GPUs, with an optional type that “must be a substring of the ‘MIG Profile’ string” the node reports; the documentation’s example uses a100_3g.20gb. The same page states that “Slurm expects MIG devices to already be partitioned, and does not support dynamic MIG partitioning.” The MIG layout is therefore a property of the node. Changing it means draining the node, rebuilding the instances and adjusting the configuration, so plan layout changes like any other maintenance. NVIDIA’s MIG guide also states that “NCCL is currently not supported with MIG”, so an instance runs single-GPU jobs, and multi-card training takes whole cards.

NVIDIA’s MIG guide describes instances with “separate and isolated paths through the entire memory system” and states that “MIG ensures one client cannot impact the work or scheduling of other clients”. A student’s notebook that crashes then stops on its own, and a colleague’s job keeps running.

CARDMEMORYMIG INSTANCESFP64FP32
H200 NVL141 GB HBM3e, 4.8 TB/s1g.18gb (up to 7), 1g.35gb (4), 2g.35gb (3), 3g.71gb (2), 4g.71gb, 7g.141gb30 TFLOPS, 60 on Tensor Cores60 TFLOPS
RTX PRO 6000 Server Edition96 GB GDDR7, 1,597 GB/s1g.24gb (up to 4), 2g.48gb (2), 4g.96gb1/64 of FP32, about 1.9 TFLOPS120 TFLOPS
RTX PRO 6000 Workstation96 GB GDDR7the same profiles, on the Max-Q too, after the display-mode switch1/64 of FP32, about 2.0 TFLOPS (Max-Q 1.7)125 TFLOPS (Max-Q 110)

NVIDIA MIG user guide (updated 11 September 2026), which lists the H200 NVL with 7 instances and gives the profiles in one table for the H200 141GB; NVIDIA’s H200 page gives the H200 NVL “Up to 7 MIGs @16.5GB each”. H200 and RTX PRO 6000 Server Edition, Workstation and Max-Q product pages, read 10 October 2026; the 1/64 FP64 rate from NVIDIA’s RTX Blackwell PRO whitepaper v1.1; RTX PRO 6000 FP64 values are our arithmetic.

On the Workstation and Max-Q editions MIG needs a switch of the display mode first, which our MIG runbook for RTX PRO Blackwell describes together with a second point that matters for Slurm. On Hopper and later GPUs, MIG mode and the instances are gone after a reboot. NVIDIA suggests a systemd service with its MIG Partition Editor (nvidia-mig-parted) to recreate them at start-up; on a Slurm node it has to run before slurmd starts, since AutoDetect reads the devices the node has. The L40S, L4, RTX PRO 4000 and RTX PRO 2000 are not in NVIDIA’s MIG guide, so they serve whole jobs or are shared without hardware isolation.

We build GPU servers to order with H200 NVL or RTX PRO 6000 Server Edition cards, with drivers and CUDA installed on request. Tell us how many people share the server and what they run.

Sharing without MIG: Slurm shards and MPS

Slurm also offers shards, a “generic mechanism where GPUs can be shared by multiple jobs”, configured as a count per node such as shard:64 beside the GPUs. The documentation states that sharding “does not fence the processes running on the GPU” and that it “works best with homogeneous workflows”. A GPU is allocated either as a gpu or as shards, not both. Shards suit a course in which every student runs the same small exercise, if the lab accepts that one job using more than its share of memory or compute can slow or block the others.

CUDA MPS is the other option, and Slurm’s documentation warns that “NVIDIA MPS has a built-in limitation regarding GPU sharing among different users”. On a server where every researcher has an account of their own, MIG is the method that separates users in hardware.

Fair-share and GPU quotas in Slurm

Slurm’s multifactor priority plugin weighs nine factors, among them age, fair-share, partition and QOS. It defines fair-share as “the difference between the portion of the computing resource that has been promised and the amount of resources that has been consumed”, and since release 19.05 the Fair Tree algorithm is the default. PriorityDecayHalfLife, 7 days by default, sets how fast past usage fades.

For a GPU server one default needs changing. “By default, the computing resource is the computing cycles delivered by a machine in the units of allocated_cpus*seconds.” A job with two CPU cores and four GPUs would then count as small. The same page continues: “Other resources can be taken into account by configuring a partition’s TRESBillingWeights option.” Giving GPUs a weight in that option makes fair-share charge for cards, and the MAX_TRES_GRES priority flag bills the largest of a node’s CPU and memory shares plus the GPUs.

Hard limits sit beside fair-share. With AccountingStorageEnforce=limits, Slurm enforces limits set on associations, that is on the cluster, an account or a user. GrpTRES caps the TRES that an association and its children use at one time, and the documentation’s example sets gres/gpu=50 for one user with sacctmgr. In a QOS, MaxTRESPerUser caps what one user holds, MaxJobsPerUser the jobs a user runs at once and MaxWallDurationPerJob the run time of each job. The documentation warns that a typed limit such as gres/gpu:tesla=1 is not enforced when a job asks for GPUs without a type, a “design limitation”, and suggests a job submit plugin that rejects requests without a type.

A typical set-up for a lab cluster runs in this order:

  1. Partition the MIG cards, then declare GPUs and MIG instances in slurm.conf and gres.conf with AutoDetect=nvml.
  2. Add gres/gpu to AccountingStorageTRES, so the accounting database records GPU use per job.
  3. Create an account per research group and the users under it, with the shares each group is promised.
  4. Set TRESBillingWeights on the GPU partitions, so fair-share counts cards and not only cores.
  5. Set AccountingStorageEnforce=limits, a GPU cap per group with GrpTRES, a per-user cap and a maximum run time in a QOS.

Kubernetes labs: Kueue quotas and MIG resources

The NVIDIA GPU Operator and device plugin expose whole cards or MIG instances as resources, and with the mixed strategy each MIG profile gets its own resource name; our Kubernetes GPU sharing guide explains the strategies. As that guide notes, we found no RTX PRO 6000 Workstation or Max-Q entry on the GPU Operator’s support list as of October 2026, so plan a Kubernetes cluster on cards from the list, such as the RTX PRO 6000 Server Edition and the H200 NVL.

Kueue, in its own words “a kubernetes-native system that manages quotas and how jobs consume them”, decides when a job waits, starts or is preempted, and “does not replace any existing Kubernetes components”. Each research group gets a ClusterQueue with a nominalQuota of GPUs. Queues in the same cohort “can borrow unused quota from each other”, up to a borrowingLimit, and a lendingLimit holds back some of a queue’s own quota. With reclaimWithinCohort set to LowerPriority or Any (the default is Never), a pending job can preempt borrowed work in other queues of the cohort that exceed their nominal quota.

Batch jobs, MPI codes and users who know sbatch from a computing centre fit Slurm. Notebooks, model services and training pipelines built as containers fit Kubernetes with Kueue.

Choosing cards for training, inference, teaching and FP64 codes

Double precision separates the two main cards. The H200 NVL delivers 30 TFLOPS of FP64 and 60 on its Tensor Cores. NVIDIA’s RTX PRO Blackwell whitepaper states that “The FP64 TFLOP rate is 1/64th the TFLOP rate of FP32 operations” and that the FP64 cores are there “to ensure any programs with FP64 code operate correctly”. Simulation, quantum chemistry and other codes that compute in FP64 belong on the H200 NVL; our comparison of GPUs for CFD and simulation covers solvers in detail. For machine learning, our H200 NVL and RTX PRO 6000 comparison for fine-tuning sets memory, tensor rates and NVLink against each other.

LAB TYPETYPICAL JOBSCONFIGURATIONSHARING
Teaching and coursesnotebooks, exercises, small modelsRTX PRO 6000 Server Edition cards; 4 cards give 16 instances of 24 GBMIG 1g.24gb as Slurm GPUs or Kubernetes resources
Machine learning groupfine-tuning, training, evaluationH200 NVL, 2 or 4 cards on an NVLink bridgewhole cards, fair-share per group
LLM and inference researchmodel services, evaluation runsRTX PRO 6000 Server Edition, 2 to 4 cardswhole cards for services, MIG for experiments
Simulation and HPC codesFP64 solvers, molecular dynamicsH200 NVLwhole cards, wall-time limits
Department-wide clusterall of the aboveone partition per node typefair-share, GPU caps per group

Our reading of NVIDIA’s MIG guide and product pages (table above) and of the reference configurations on our AI servers page, read 10 October 2026; card counts are examples, not sizing.

Our AI servers page lists two reference nodes at this scale: four RTX PRO 6000 Server Edition cards for inference, and eight H200 NVL cards on NVLink bridges with 400G networking for training. A department cluster can run both under one Slurm controller, in separate partitions.

Describe your lab in the form below, with its teams, workloads and the jobs that run at the same time, and we reply with a configuration and quote within one business day.

Storage and network for a shared GPU server

Home directories, datasets and checkpoints should have the same paths on every node, so that a job runs wherever the scheduler places it. Local NVMe drives serve as scratch space during a run, with a separate model store for what the lab keeps. For a single server, 25 GbE uplinks carry logins and data copies; several nodes training one model need a cluster fabric at 200 or 400G, which the H200 NVL training node on our AI servers page plans for.

What we supply

We build AI servers to order for shared use, with the H200 NVL and NVLink bridges, the RTX PRO 6000 Server Edition, the L40S or the L4, assembled and burn-in tested, with manufacturer warranty on every component and on one EU contract and invoice. Single cards come through our professional GPU range, including the RTX PRO 6000 Workstation and Max-Q editions for desk-side machines. We check the rack, power and airflow before we quote, and install the operating system, drivers, CUDA and a container runtime on request. If the lab wants a Kubernetes platform with Kubeflow built and supported on top, that is our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What GPU server does a university research lab need?
It needs cards that match the work plus a scheduler, partitioning and quotas, because many users run bursty jobs on the same hardware. Slurm with GPUs as generic resources, or Kubernetes with Kueue, queues the jobs, MIG splits a card into isolated instances, and fair-share with per-group limits divides the machine. H200 NVL cards suit training and double-precision codes, RTX PRO 6000 Server Edition cards suit inference and teaching.
How do I configure GPUs in Slurm?
List gpu in GresTypes and declare the cards on each node line in slurm.conf, for example Gres=gpu:4. Put the device files and CPU affinity in gres.conf, or use AutoDetect=nvml to fill them in for NVIDIA GPUs. Users then request cards with --gpus or --gres, and adding gres/gpu to AccountingStorageTRES records GPU use per job.
Does Slurm support MIG?
Yes, since version 21.08. Slurm detects MIG instances with AutoDetect=nvml and lists them like ordinary GPUs, but it expects them to be partitioned already and does not create or change instances on demand. On Hopper and later GPUs the layout disappears at a reboot, so it has to be recreated at start-up, before slurmd starts on the node.
How can researchers share GPUs fairly?
Use fair-share scheduling with per-group accounts, and give GPUs a weight with TRESBillingWeights, because Slurm counts allocated CPU seconds by default. Add hard caps such as GrpTRES per group, MaxTRESPerUser per user and a maximum wall time per job. In Kubernetes, Kueue gives each group a ClusterQueue with a GPU quota and lets unused quota be borrowed within a cohort.
Which GPU is suitable for FP64 HPC codes?
The H200 NVL, with 30 TFLOPS of FP64 and 60 on its Tensor Cores according to NVIDIA. The RTX PRO 6000 runs FP64 at 1/64 of its FP32 rate, which NVIDIA says is there so that FP64 code runs correctly; on the Server Edition that is about 1.9 TFLOPS. Double-precision solvers therefore belong on the H200 NVL.
Slurm or Kubernetes for a research GPU cluster?
Slurm fits batch jobs, MPI codes and users used to a computing centre, with fair-share and GPU limits built in. Kubernetes with Kueue fits labs that run notebooks, model services and container pipelines, with quotas per group and borrowing between them. A Kubernetes cluster should use cards on NVIDIA’s GPU Operator support list, such as the RTX PRO 6000 Server Edition and the H200 NVL.

Send us the number of researchers and students who share the server, what they run (training, inference, teaching or double-precision codes), the scheduler you use or plan, and the rack position’s power feed. We reply within one business day with a configuration and a quote, and we check the rack, power and airflow before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna