NVIDIA licence server: DLS or CLS, how to set it up and what happens when it is unreachable
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- The NVIDIA License System serves vGPU and NVIDIA AI Enterprise licences from a Cloud License Service (CLS) instance hosted on the NVIDIA Licensing Portal or from a Delegated License Service (DLS) instance you run on-premise; compute on bare metal or in containers needs no licence server
- A DLS is set up in order: appliance, administrator user, licence server on the portal, DLS instance token upload, binding, licence file install, client configuration token, then the token on each VM or client
- High availability needs at least two DLS instances, a primary and a secondary, with up to nine per cluster since vGPU 16.1; with two, the survivor of a failure is a single point of failure
- A compute VM that boots without a licence runs 20 minutes at full capability, then at idle-level compute; a licensed compute VM keeps working 7 days by default on CLS (configurable on DLS) after losing the service, and graphics clients up to 1 day
- One DLS instance with 4 vCPU and 8 GB served up to 32,000 clients in NVIDIA’s measurements, so for 200 to 1,000 VMs the number of HA clusters follows sites and network zones, not the client count
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
NVIDIA licence server: DLS or CLS
NVIDIA vGPU and NVIDIA AI Enterprise licences reach virtual machines from an NVIDIA licence server, a service instance of the NVIDIA License System that comes in two kinds. A Cloud License Service (CLS) instance is hosted on the NVIDIA Licensing Portal. A Delegated License Service (DLS) instance is a virtual appliance or container that you run on-premise, inside your own network. Choose CLS when the VMs may open outbound HTTPS connections to NVIDIA, and DLS when they may not, when the site is air-gapped, or when the licence service has to stay under your control with a high-availability cluster of your own.
The licence server serves VMs with compute or graphics vGPU, and GPUs in passthrough or on bare metal that run RTX vWS graphics. NVIDIA’s licensing guide for AI Enterprise, updated on 2 September 2026, states that the License System “is only required to be installed and configured when using NVIDIA vGPU for Compute drivers”. The vGPU client licensing guide applies licence enforcement to A, B and Q-series vGPUs and to passthrough and bare-metal GPUs used for professional 3D graphics. A server running models on bare metal or in containers with the data-centre driver needs no licence server. How licences are counted is in our guide to NVIDIA AI Enterprise licensing. The sources below are the NVIDIA License System guides, version 3.6.1 of 29 June 2026, and the vGPU client licensing guide of 29 September 2026.
CLS vs DLS compared
| ASPECT | CLS INSTANCE | DLS INSTANCE |
|---|---|---|
| Where it runs | NVIDIA Licensing Portal | VM or container on your network |
| Client connection | port 443 outbound to NVIDIA | port 443 to the DLS |
| Licences on the instance | allocated on the portal | downloaded, then uploaded by hand |
| Air-gapped sites | not usable | usable; node-locked file for offline clients |
| High availability | no cluster on your side | your cluster of 2 to 9 instances |
| Resources on site | none | minimum 4 vCPU, 8 GB RAM, 15 GB disk |
| Upgrades | NVIDIA and its cloud provider | you migrate the appliance |
NVIDIA License System user guide and quick start guide v3.6.1 (29 June 2026), sections 1.2, 1.4, 2.1 and 2.3.
CLS suits VMs that already reach the internet through a controlled egress; for a non-transparent proxy, the client licensing guide documents a ProxyServerAddress setting. DLS suits VMs without a route to the internet, or estates where licensing must keep working while the uplink is down. Because “a DLS instance is fully disconnected from the NVIDIA Licensing Portal”, as the user guide puts it, you “must download licenses from the NVIDIA Licensing Portal and upload them to the instance manually”. For a client with no network at all, the same guide describes node-locked licensing from a file installed locally, available since vGPU software 15.0; the rest of an offline installation is in our article on the air-gapped LLM server.
How to set up an NVIDIA DLS instance
NVIDIA’s quick start guide and user guide give the steps in this order.
- Deploy the DLS image as a VM on a supported hypervisor (the guide lists VMware vSphere, Red Hat Enterprise Linux KVM and Ubuntu among others) or as a container on Docker, Kubernetes, Podman or OpenShift, with at least 4 vCPU, 8 GB RAM and 15 GB disk, a fixed IP address and a reverse DNS entry.
- Open the appliance in a browser, choose NEW INSTALLATION, register the DLS administrator user and keep the local reset secret it displays; an HA cluster can be created now or at any time later.
- On the NVIDIA Licensing Portal, create a licence server and allocate licences from your entitlements to it.
- On the appliance, download the DLS instance token, then upload it on the portal under Register DLS Instance.
- On the portal, bind the licence server to the registered DLS instance, so that its licences are available only from that instance.
- Download the licence file (
license_*.bin) from the portal. - On the appliance, select the licence file and install it, then generate a client configuration token with at least one scope reference and download it.
- Copy the token to each client and restart the client’s licensing service.
NVIDIA explains that the appliance “requires the reverse pointer entry to determine the domain name” when the token is created. The entitlement itself, which arrives before step 3, is covered in our article on activating NVIDIA AI Enterprise licences.
The client configuration token on each VM
The path and the service to restart differ by guest OS. On Linux, the token goes into /etc/nvidia/ClientConfigToken, with file modes that let the owner read, write and execute it and others read it (744). A vGPU client also gets the FeatureType parameter in /etc/nvidia/gridd.conf; the guide gives the value 1 for vGPU, with the licence type then selected automatically, and the nvidia-gridd service is restarted. On Windows guests, the token goes into the ClientConfigToken folder under the vGPU Licensing directory of the NVIDIA installation, and the NvDisplayContainer service is restarted. After that, according to the guide, “The NVIDIA service on the client should now automatically obtain a license”.
By default, the guide says, the licence “is checked out when the VM is booted and released when the VM is shut down”, so the token belongs in the golden image of a desktop pool or the template of the compute VMs, and a new VM licenses itself on first boot. The differences between vGPU and passthrough on ESXi are in our guide to GPUs in VMware vSphere.
DLS high availability and the ports between nodes
The user guide states that “High availability requires at least two DLS instances in a failover configuration”, a primary that serves licences and one or more secondaries that act as its backup. A cluster holds at most nine instances, and clusters of more than two have been supported since vGPU software 16.1. With only two instances, “the remaining instance becomes a single point of failure” after a failure, so repair or replace the failed node promptly.
For clients to follow a failover without a new token, the guide offers Virtual IP Management, which “automatically migrates a floating IP address between cluster nodes during failover events”. Place the instances on different hosts. Between the nodes, the firewall has to pass ports 4369 (peer discovery), 5671 (AMQP over TLS) and 8080 to 8085 (HTTPS), and 443 between virtual appliances; the container version does not use 443 between nodes. Clients need only 443 to the cluster, and 80, which NVIDIA lists for licence release by Windows VMs. We found no failover time in the guide, so test a failover once after setup and note how long clients take to renew against the new primary.
What happens when the licence server is unreachable
The documented behaviour depends on whether the VM ever held a licence and on its type: compute VMs on C-series profiles follow NVIDIA AI Enterprise’s rules, graphics VMs those of the vGPU client licensing guide.
| SITUATION | DOCUMENTED BEHAVIOUR |
|---|---|
| Boot without a licence | full capability for 20 minutes |
| Compute VM after 20 minutes | “compute performance reduced to an idle level” until licensed |
| Graphics VM after 20 minutes | degraded, “further degraded” after 24 hours |
| Compute VM loses service | “7 days (CLS default; configurable for DLS)”, then degraded |
| Graphics VM loses service | “up to 1 day”, then “warned of license expiration” |
| Licence obtained again | “Full capability is restored immediately” (compute) |
| Primary DLS fails (HA) | a secondary becomes primary and serves licences |
| DLS before 3.4, vGPU 18 | licensing failures; upgrade DLS first |
| Compute on bare metal | no licence server involved |
NVIDIA AI Enterprise licensing page (infrastructure release 8, updated 2 September 2026); vGPU client licensing guide, introduction and advanced topics (updated 29 September 2026); NVIDIA License System user guide v3.6.1, sections 1.1.1 and 1.4.
The current graphics guide does not list which functions are limited. It says that a client without a licence “will periodically retry its license request to the license server”, and its introduction gives the two steps of degradation, after 20 minutes and after 24 hours. A firewall change between the VM networks and the DLS can go unnoticed while the licences last, up to a day for desktops and, for compute VMs, the period configured on the DLS (seven days is the CLS default), and then affects every VM at once. Monitor the licence state of a sample of clients and the DLS event log, and add the client ports to the change checklist of the firewall team.
We supply NVIDIA AI Enterprise and vGPU licences on the same quote as the GPUs, and our engineering partner Vixen.UNO can deploy and support the licence service. Tell us how many vGPU VMs run at each site and whether they reach the internet.
Sizing for 200 to 1,000 VMs: one DLS cluster or several
Client count rarely decides the number of DLS instances. NVIDIA measured one instance with 4 vCPU and 8 GB at 2.6 GHz, with licences borrowed for 12 hours and renewed at 15 per cent of that time, and it served up to 32,000 clients; 8 vCPU and 16 GB raised the figure to 84,000. The guide adds that “The maximum number of clients is directly proportional to the length of time for which licenses are borrowed.” In its burst test, the same 4 vCPU instance served 1,000 clients that arrived within 20 seconds in 50 seconds, results that NVIDIA calls “illustrative only”.
For an estate of 200 to 1,000 vGPU VMs, the minimum appliance therefore carries the load, including the boot storm of a whole desktop pool, with a wide margin. The number of clusters follows from sites, network zones and failure domains. Take a company with 600 virtual desktops on L40S and L4 hosts at the main site, 150 desktops at a second site and 24 compute VMs on H200 NVL for an internal assistant. One HA cluster of two instances at the main site, behind a virtual IP, serves all 774 VMs. The second site then depends on the WAN link for its licences, and its desktops run for up to a day if the link fails.
Where the second site must license its VMs on its own, as the recovery site of a DR plan for example, give it its own cluster. Licences are allocated to a licence server on the portal, and each licence server is bound to one service instance, so each site holds its own share. A recovery plan that moves desktops to the second site needs enough licences allocated there, or a change of allocation followed by a new licence file on that DLS. Sizing the desktops and GPUs themselves is covered in our guide to VDI GPU server sizing.
We quote the licences for each site together with the GPUs they run on. Describe your sites, desktop pools and compute VMs in the form below, including any site that has to license its VMs on its own during a failover.
Renewals, upgrades and the three-year licence file
From NLS release 3.4, “the license server files bound to DLS SI will have a validity of three years”, counted from the download; the portal shows the validity in the licence server’s details. Keep that date in the same calendar as the subscription renewals. A renewed subscription changes the entitlement on the portal, and a DLS learns of it only when you download and install the licence file again.
NVIDIA’s prerequisites for vGPU 18 and later say to upgrade the licence server “to a minimum version of DLS 3.4” before upgrading vGPU to 18.0 or later, then to download a fresh licence server file and install it on the upgraded appliance. A DLS appliance is upgraded by migration to a new one: the licence servers, the user registration, the IP address and the service instance move across, while event records on the old appliance are not migrated. If you use the event records for licence usage reports, save what you need from them before the migration. Plan the DLS upgrade as its own change in an agreed maintenance window with a rollback plan, ahead of the host and guest driver upgrade.
What we supply
We supply NVIDIA AI Enterprise licences for compute vGPU and the RTX vWS, vPC and vApps licences for graphics, on one EU contract and invoice with the GPUs they run on, and we size them with the hardware so that the licence lines are in the quote. We supply the L4, L40S and H200 NVL from the example above and the RTX PRO 6000 Server Edition, with manufacturer warranty, as cards or in AI servers built to order. Our NVIDIA AI Enterprise and vGPU licence page explains how each licence is counted, and the full range of cards is on our professional GPUs page. If you want it, our engineering partner Vixen.UNO designs, deploys and supports the setup under the same contract.
FAQ
What is the NVIDIA licence server?
What is the difference between NVIDIA CLS and DLS?
How do I set up an NVIDIA DLS?
Where does the NVIDIA client configuration token go?
What happens to a vGPU VM in an unlicensed state?
How many VMs can one NVIDIA DLS instance serve?
Send us the number of vGPU VMs and virtual desktops per site, the GPUs and hypervisor, and whether the VMs may reach the internet. We reply within one business day with the licences your setup needs and a written quote for them with the GPUs.
Talk to an expertWe reply within one business day