BLOG · GUIDE ·

GPU chargeback and quotas: sharing a GPU platform between departments with Kueue, Run:ai and MIG

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Departments share 8 to 32 GPUs through three controls: a fixed quota per department, borrowing of idle cards with preemption to reclaim them, and metering of GPU-hours per department for showback or chargeback
  • A Kubernetes ResourceQuota only caps a namespace: GPUs are an extended resource, so only quota items with the requests. prefix are allowed, such as requests.nvidia.com/gpu: 4, and a request over the cap is rejected with HTTP 403
  • Kueue gives each department a ClusterQueue with a nominalQuota per ResourceFlavor, one flavor per card type; queues in one cohort borrow unused quota up to a borrowingLimit, and reclaimWithinCohort, Never by default, lets the owner preempt borrowed work
  • NVIDIA Run:ai groups projects into departments with a deserved quota per node pool; only preemptible workloads may run over quota, and idle GPUs are divided by rank and then by over quota weight, 2 by default for departments when weights are enabled
  • The DCGM exporter labels GPU and MIG metrics with pod and namespace through the kubelet pod-resources API once pod mapping is on; for GPU-hours we count a 1g.24gb instance as a quarter of an RTX PRO 6000 and a 1g.18gb as a seventh of an H200 NVL

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

How departments share a GPU platform

On a shared GPU platform with quotas and chargeback, departments use one pool of 8 to 32 GPUs through three controls. Each department gets a quota of cards it can always claim, idle cards above that quota can be borrowed by others with preemption to hand them back, and a meter records GPU-hours per department for showback or chargeback. On Kubernetes, Kueue or NVIDIA Run:ai enforces the quota and the borrowing, while a plain ResourceQuota only caps each namespace. The meter is built from the per-pod metrics of NVIDIA’s DCGM exporter, and MIG sets how small the unit of allocation can be.

How one card is split between pods with MIG, time-slicing or MPS is explained in our guide to GPU sharing in Kubernetes, and Slurm with fair-share for research teams in our guide to GPU servers for university labs.

CONTROLTOOLWHAT IT ENFORCES
Hard cap per namespaceKubernetes ResourceQuotaGPUs a namespace may request at once; requests over the cap are rejected
Quota and borrowingKueue ClusterQueue in a cohortadmission against nominalQuota, borrowing up to borrowingLimit, lending up to lendingLimit
Reclaim of borrowed cardsKueue preemption, Run:ai reclaimborrowed work is preempted when the owner needs its quota back
Share of idle cardsKueue Fair Sharing weight, Run:ai over quota weightspare GPUs divided by weight
Department hierarchyRun:ai departments and projectsdepartment quota and limit above the quotas of its projects
Allocation unitMIG profilesinstances such as 1g.24gb counted as resources of their own
MeteringDCGM exporter with pod mappingGPU and MIG metrics labelled with pod and namespace

Kubernetes Resource Quotas page (modified 28 September 2026); Kueue ClusterQueue, LocalQueue and preemption pages (19 August and 30 September 2026); NVIDIA Run:ai self-hosted documentation; NVIDIA MIG user guide (11 September 2026); all read on 10 October 2026.

Kubernetes ResourceQuota: a hard cap per namespace

In the Kubernetes documentation a ResourceQuota “provides constraints that limit aggregate resource consumption per namespace”, and the page’s own scenario is “Different teams work in different namespaces.” GPUs are an extended resource, and since overcommit is not allowed for extended resources, only quota items with the prefix requests. are allowed. The documentation’s example is requests.nvidia.com/gpu: 4. With the GPU Operator’s mixed MIG strategy every profile is a resource name of its own, so a line such as requests.nvidia.com/mig-1g.24gb: 8 caps instances separately from whole cards.

When creating a pod would exceed the quota, “the control plane rejects that request with HTTP status code 403 Forbidden”. A ResourceQuota does not queue work and does not lend: a department at its cap waits even when half the cluster is idle, and a department below its cap has no claim on cards that others already hold. A quota with a PriorityClass scope applies only to pods of that priority class, which can limit how many GPUs a namespace runs at high priority. Keep it as the upper bound and add a queueing layer for the sharing.

Kueue: a ClusterQueue per department, cohorts and fair sharing

Kueue, the latest release v0.19.6 on its GitHub page on 10 October 2026, admits jobs against quotas before the Kubernetes scheduler places them. Users submit jobs to a LocalQueue, “a namespaced resource that groups closely related workloads belonging to a single tenant”, and each LocalQueue “points to one ClusterQueue from which resources are allocated to run its Workloads”. For a company that means one ClusterQueue per department, with its namespaces chosen by the .spec.namespaceSelector field, and one cohort for the whole platform.

ResourceFlavors describe the card types. Kueue’s documentation says “Flavors represent different variations of a resource (for example, different GPU models)”, and a ClusterQueue defines “the quota for each of the resources that a flavor offers”. A department can then hold a nominalQuota of six H200 NVL and four RTX PRO 6000 as two separate figures. ClusterQueues in the same cohort “can borrow unused quota from each other”. The borrowingLimit caps what a queue may take from the others, and the lendingLimit caps how much of its own idle quota it lends, so a lendingLimit of zero keeps a production service’s cards out of the pool.

Preemption has to be switched on. The field reclaimWithinCohort takes Never, the default, LowerPriority or Any, and lets a department that needs its quota back preempt workloads of other queues that are borrowing. The field withinClusterQueue, also Never by default, lets higher-priority work in a department preempt its own lower-priority work. With Fair Sharing enabled in the Kueue configuration, a queue’s share “is weighted by the .spec.fairSharing.weight defined in a ClusterQueue”, and the preemption strategies LessThanOrEqualToFinalShare and LessThanInitialShare decide when one department may preempt another.

NVIDIA Run:ai: departments, projects and over-quota

NVIDIA Run:ai builds the same controls into a two-level hierarchy. Its documentation says “Departments group multiple projects under a shared organizational scope”, and in its scheduler “project queues are bound to department queues, per node pool”. Each project has a deserved quota per node pool, and “Deserved quota means that a project is entitled to use up to a maximum number of resources defined by its quota”. A department’s quota cannot be set below the quota already assigned to its projects. Its limit, Unlimited by default, caps the GPUs the department can take from a node pool.

The workload type decides the policy. “Non-preemptible workloads can only be scheduled if their requested resources are within the deserved resource quotas”, and “Over quota resources can only be used by preemptible workloads.” Rank decides first who receives idle GPUs, and within a rank “The part each Project receives depends on its over quota weight value, and the total weights of all other projects over quota weights.” With over quota weight enabled, departments of the same rank start at a default weight of 2; with it disabled, the weight follows the assigned quota. Reclaim “returns the resources back to a project (or department) that deserves those resources as part of its deserved quota”. Inside a project, higher-priority workloads may preempt lower-priority preemptible ones, and a time-based fairshare mode calculates fairshare from historical usage.

NVIDIA’s AI Enterprise documentation of 10 August 2026 states that “both NVIDIA Run:ai self-hosted and NVIDIA Run:ai SaaS are included in your NVIDIA AI Enterprise license”. AI Enterprise is a per-GPU entitlement, as our guide to NVIDIA AI Enterprise licensing explains, and NVIDIA’s 90-day trial licence does not include Run:ai.

We supply NVIDIA AI Enterprise licences, which include Run:ai, on one EU contract and invoice with the servers. Tell us how many GPUs and departments the platform will have, and we recommend the edition and term.

MIG profiles as the unit of allocation

A quota counts whole resources, so without partitioning the smallest unit a department can hold is one card. NVIDIA’s MIG user guide of 11 September 2026 lists up to four 1g.24gb instances, two 2g.48gb or one 4g.96gb for the RTX PRO 6000 Blackwell Server Edition. For the H200 with 141 GB, the memory of the H200 NVL, it lists up to seven 1g.18gb, three 2g.35gb or two 3g.71gb instances, and Kueue quotas name each profile like any other resource.

MIG instances have their own memory and fault isolation, so a notebook in one department cannot fill the memory that another department’s instance holds on the same card. The layout is a node setting that MIG Manager changes only with no user workloads on the GPUs. Plan the profiles per node together with the quotas, and change them in agreed maintenance windows.

Metering GPU-hours per department for showback or chargeback

Showback reports each department’s use, and chargeback books it to its cost centre. Both need GPU-hours per department and per card type. NVIDIA’s blog of 4 November 2020 explains that the DCGM exporter “connects to the kubelet pod-resources server” to identify the GPU devices of each pod and “appends the GPU devices pod information to the metrics collected”. The exporter’s README shows its sample metrics with container, namespace and pod labels. NVIDIA’s DCGM Exporter page lists the pod mapping option -k with “Default: false”, so check that it is on in your deployment. For MIG, the exporter “publishes metrics for both the entire GPU as well as individual MIG devices”, labelled with GPU_I_ID and GPU_I_PROFILE.

From those series, Prometheus can count for each sampling interval how many GPUs or MIG instances a namespace held. Summed over a month and mapped from namespaces to departments, the count gives allocated GPU-hours, the time a card was held whether it was busy or not. The other basis books each department’s quota in full and reports borrowed hours on top, which follows the platform’s structure when quotas are sized to the cards bought. We weight MIG instances by the maximum count per card in NVIDIA’s guide, a 1g.24gb as a quarter of an RTX PRO 6000 and a 1g.18gb as a seventh of an H200 NVL, and report the two card types on separate lines.

A utilisation column shows which departments hold cards they do not use; our guide to DCGM monitoring explains why GPU utilisation overstates load and which fields measure it. For a time-sliced card that several pods share, NVIDIA’s DCGM command-line reference of 10 September 2026 lists the option --kubernetes-virtual-gpus, off by default, to “Attribute supported time-sharing or MPS assignments in Kubernetes mode”. Without it, meter time-sliced nodes by the replicas each namespace holds.

A shared inference service is the exception. A RAG assistant for the whole company runs in the platform team’s namespace, so its GPU-hours land on one cost centre. Split them by the tokens each department sends, taken from the per-key log of the gateway, as our guide to an LLM gateway with keys and budgets describes.

Worked example: 24 GPUs for five department workloads

As an example, take a company of 1,500 staff with three servers: two with eight RTX PRO 6000 Server Edition each and one with eight H200 NVL. Both cards are on the GPU Operator’s support list and on NVIDIA’s MIG list. Five workloads share the 24 cards.

DEPARTMENT PATTERNCARDS AND UNITQUOTAPREEMPTIONMETER
Assistant and RAG service8 RTX PRO 6000, four per server8, lends nonenever preemptedtokens per department from the gateway
Data science, fine-tuning8 H200 NVL, whole cards8, lends idle cardsreclaims borrowed cardsallocated GPU-hours
Developer notebooks4 RTX PRO 6000 as 16 × 1g.24gb16 instancesby priority inside the queueinstance-hours ÷ 4
Department model services2 RTX PRO 6000, whole cards2, lends nonenever preemptedallocated GPU-hours
Batch document jobs2 RTX PRO 6000, plus borrowed cards2, borrows up to 8preemptible above quotaown quota plus borrowed hours

Our example allocation; MIG profiles from NVIDIA’s MIG user guide (11 September 2026), quota and preemption fields from the Kueue and Run:ai documentation read on 10 October 2026.

The assistant runs four cards on each RTX server, so it keeps half of its capacity if one server fails. Its lendingLimit of zero keeps those cards available for the morning peak. Batch jobs borrow idle data-science cards at night and are preempted when fine-tuning starts. In Kueue, the set-up runs in this order.

  1. Map each department to one or more namespaces, label them and create one LocalQueue per namespace.
  2. Define one ResourceFlavor per card type and one for the nodes that carry the 1g.24gb layout.
  3. Create one ClusterQueue per department with a nominalQuota per flavor, all in one cohort, and set the lendingLimit of the two service queues to zero.
  4. List the H200 NVL flavor in the batch queue with a nominalQuota of 0 and a borrowingLimit, since “A ClusterQueue can only borrow quota for flavors that the ClusterQueue defines”, and set reclaimWithinCohort on the queues that lend.
  5. Add a ResourceQuota per namespace as the upper bound, then switch on pod mapping in the DCGM exporter and build the monthly report.

In Run:ai the same plan becomes five projects in three or four departments: the two services run non-preemptible inside their deserved quota, and the batch project runs preemptible over quota with a lower rank.

We build GPU servers to order for platforms of this size, with RTX PRO 6000 Server Edition or H200 NVL cards. Send us the departments, their workloads and the cards each should hold through the form below.

What we supply

We build AI servers to order for a shared GPU platform, with RTX PRO 6000 Server Edition and H200 NVL cards, assembled, burn-in tested and delivered anywhere in the EU with manufacturer warranty. The same cards, and the L40S and L4 for whole-GPU services, are available on their own through our professional NVIDIA GPU range. NVIDIA AI Enterprise licences, which include Run:ai, come through our software and licensing offer, on one EU contract and invoice with the hardware. Operating system, drivers, CUDA and a container runtime are installed on request, and building the Kubernetes-based platform on top is our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What is GPU chargeback?
GPU chargeback books the GPU time each department uses to its own cost centre, usually as GPU-hours per card type. The hours come from the DCGM exporter, which with pod mapping switched on labels GPU and MIG metrics with the pod and namespace that hold the device, mapped from namespaces to departments. Shared inference services are split by the tokens each department sends through the gateway.
How do I set a GPU quota per team in Kubernetes?
Give each team a namespace and a ResourceQuota with a line such as requests.nvidia.com/gpu: 4, since GPUs are an extended resource and only quota items with the requests prefix are allowed. A request over the cap is rejected with HTTP 403 Forbidden, and nothing is queued or lent. For queueing and borrowing of idle cards, add Kueue with one ClusterQueue per team in a shared cohort.
How do Run:ai quotas work?
NVIDIA Run:ai gives each project and department a deserved quota per node pool, and departments collect projects under one quota and an optional limit. Non-preemptible workloads run only within the deserved quota, while preemptible ones can use idle GPUs over quota, divided by rank and over quota weight. Reclaim returns GPUs to a project or department that needs its deserved quota back.
What is the difference between GPU showback and chargeback?
Showback reports each department’s GPU use without booking it, and chargeback books the same figures to the department’s cost centre. Both rest on GPU-hours per department and card type, either as hours held or as each quota booked in full with borrowed hours reported on top. Adding a utilisation column shows which departments hold cards they do not use.
How do I build a multi-tenant GPU cluster for several departments?
Map each department to namespaces, cap each namespace with a ResourceQuota and put the sharing into Kueue or NVIDIA Run:ai, with a quota per department and card type and borrowing of idle cards. Use MIG on the RTX PRO 6000 Server Edition or H200 NVL where departments need units smaller than a card. Turn on pod mapping in the DCGM exporter so that use can be reported per department.
Can departments borrow idle GPUs from each other?
Yes. In Kueue, ClusterQueues in one cohort can borrow unused quota from each other up to a borrowingLimit, and the owner gets cards back through preemption when reclaimWithinCohort is set to LowerPriority or Any. In NVIDIA Run:ai, preemptible workloads use idle GPUs over quota, and reclaim returns them to the project or department whose deserved quota they are.

Send us the departments that will share the platform, their workloads, the cards each should hold and whether you plan Kueue or Run:ai. We reply within one business day with a configuration and a quote for the servers and licences, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna