BLOG · GUIDE ·

NVIDIA licence server: DLS or CLS, how to set it up and what happens when it is unreachable

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • The NVIDIA License System serves vGPU and NVIDIA AI Enterprise licences from a Cloud License Service (CLS) instance hosted on the NVIDIA Licensing Portal or from a Delegated License Service (DLS) instance you run on-premise; compute on bare metal or in containers needs no licence server
  • A DLS is set up in order: appliance, administrator user, licence server on the portal, DLS instance token upload, binding, licence file install, client configuration token, then the token on each VM or client
  • High availability needs at least two DLS instances, a primary and a secondary, with up to nine per cluster since vGPU 16.1; with two, the survivor of a failure is a single point of failure
  • A compute VM that boots without a licence runs 20 minutes at full capability, then at idle-level compute; a licensed compute VM keeps working 7 days by default on CLS (configurable on DLS) after losing the service, and graphics clients up to 1 day
  • One DLS instance with 4 vCPU and 8 GB served up to 32,000 clients in NVIDIA’s measurements, so for 200 to 1,000 VMs the number of HA clusters follows sites and network zones, not the client count

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

NVIDIA licence server: DLS or CLS

NVIDIA vGPU and NVIDIA AI Enterprise licences reach virtual machines from an NVIDIA licence server, a service instance of the NVIDIA License System that comes in two kinds. A Cloud License Service (CLS) instance is hosted on the NVIDIA Licensing Portal. A Delegated License Service (DLS) instance is a virtual appliance or container that you run on-premise, inside your own network. Choose CLS when the VMs may open outbound HTTPS connections to NVIDIA, and DLS when they may not, when the site is air-gapped, or when the licence service has to stay under your control with a high-availability cluster of your own.

The licence server serves VMs with compute or graphics vGPU, and GPUs in passthrough or on bare metal that run RTX vWS graphics. NVIDIA’s licensing guide for AI Enterprise, updated on 2 September 2026, states that the License System “is only required to be installed and configured when using NVIDIA vGPU for Compute drivers”. The vGPU client licensing guide applies licence enforcement to A, B and Q-series vGPUs and to passthrough and bare-metal GPUs used for professional 3D graphics. A server running models on bare metal or in containers with the data-centre driver needs no licence server. How licences are counted is in our guide to NVIDIA AI Enterprise licensing. The sources below are the NVIDIA License System guides, version 3.6.1 of 29 June 2026, and the vGPU client licensing guide of 29 September 2026.

CLS vs DLS compared

ASPECTCLS INSTANCEDLS INSTANCE
Where it runsNVIDIA Licensing PortalVM or container on your network
Client connectionport 443 outbound to NVIDIAport 443 to the DLS
Licences on the instanceallocated on the portaldownloaded, then uploaded by hand
Air-gapped sitesnot usableusable; node-locked file for offline clients
High availabilityno cluster on your sideyour cluster of 2 to 9 instances
Resources on sitenoneminimum 4 vCPU, 8 GB RAM, 15 GB disk
UpgradesNVIDIA and its cloud provideryou migrate the appliance

NVIDIA License System user guide and quick start guide v3.6.1 (29 June 2026), sections 1.2, 1.4, 2.1 and 2.3.

CLS suits VMs that already reach the internet through a controlled egress; for a non-transparent proxy, the client licensing guide documents a ProxyServerAddress setting. DLS suits VMs without a route to the internet, or estates where licensing must keep working while the uplink is down. Because “a DLS instance is fully disconnected from the NVIDIA Licensing Portal”, as the user guide puts it, you “must download licenses from the NVIDIA Licensing Portal and upload them to the instance manually”. For a client with no network at all, the same guide describes node-locked licensing from a file installed locally, available since vGPU software 15.0; the rest of an offline installation is in our article on the air-gapped LLM server.

How to set up an NVIDIA DLS instance

NVIDIA’s quick start guide and user guide give the steps in this order.

  1. Deploy the DLS image as a VM on a supported hypervisor (the guide lists VMware vSphere, Red Hat Enterprise Linux KVM and Ubuntu among others) or as a container on Docker, Kubernetes, Podman or OpenShift, with at least 4 vCPU, 8 GB RAM and 15 GB disk, a fixed IP address and a reverse DNS entry.
  2. Open the appliance in a browser, choose NEW INSTALLATION, register the DLS administrator user and keep the local reset secret it displays; an HA cluster can be created now or at any time later.
  3. On the NVIDIA Licensing Portal, create a licence server and allocate licences from your entitlements to it.
  4. On the appliance, download the DLS instance token, then upload it on the portal under Register DLS Instance.
  5. On the portal, bind the licence server to the registered DLS instance, so that its licences are available only from that instance.
  6. Download the licence file (license_*.bin) from the portal.
  7. On the appliance, select the licence file and install it, then generate a client configuration token with at least one scope reference and download it.
  8. Copy the token to each client and restart the client’s licensing service.

NVIDIA explains that the appliance “requires the reverse pointer entry to determine the domain name” when the token is created. The entitlement itself, which arrives before step 3, is covered in our article on activating NVIDIA AI Enterprise licences.

The client configuration token on each VM

The path and the service to restart differ by guest OS. On Linux, the token goes into /etc/nvidia/ClientConfigToken, with file modes that let the owner read, write and execute it and others read it (744). A vGPU client also gets the FeatureType parameter in /etc/nvidia/gridd.conf; the guide gives the value 1 for vGPU, with the licence type then selected automatically, and the nvidia-gridd service is restarted. On Windows guests, the token goes into the ClientConfigToken folder under the vGPU Licensing directory of the NVIDIA installation, and the NvDisplayContainer service is restarted. After that, according to the guide, “The NVIDIA service on the client should now automatically obtain a license”.

By default, the guide says, the licence “is checked out when the VM is booted and released when the VM is shut down”, so the token belongs in the golden image of a desktop pool or the template of the compute VMs, and a new VM licenses itself on first boot. The differences between vGPU and passthrough on ESXi are in our guide to GPUs in VMware vSphere.

DLS high availability and the ports between nodes

The user guide states that “High availability requires at least two DLS instances in a failover configuration”, a primary that serves licences and one or more secondaries that act as its backup. A cluster holds at most nine instances, and clusters of more than two have been supported since vGPU software 16.1. With only two instances, “the remaining instance becomes a single point of failure” after a failure, so repair or replace the failed node promptly.

For clients to follow a failover without a new token, the guide offers Virtual IP Management, which “automatically migrates a floating IP address between cluster nodes during failover events”. Place the instances on different hosts. Between the nodes, the firewall has to pass ports 4369 (peer discovery), 5671 (AMQP over TLS) and 8080 to 8085 (HTTPS), and 443 between virtual appliances; the container version does not use 443 between nodes. Clients need only 443 to the cluster, and 80, which NVIDIA lists for licence release by Windows VMs. We found no failover time in the guide, so test a failover once after setup and note how long clients take to renew against the new primary.

What happens when the licence server is unreachable

The documented behaviour depends on whether the VM ever held a licence and on its type: compute VMs on C-series profiles follow NVIDIA AI Enterprise’s rules, graphics VMs those of the vGPU client licensing guide.

SITUATIONDOCUMENTED BEHAVIOUR
Boot without a licencefull capability for 20 minutes
Compute VM after 20 minutes“compute performance reduced to an idle level” until licensed
Graphics VM after 20 minutesdegraded, “further degraded” after 24 hours
Compute VM loses service“7 days (CLS default; configurable for DLS)”, then degraded
Graphics VM loses service“up to 1 day”, then “warned of license expiration”
Licence obtained again“Full capability is restored immediately” (compute)
Primary DLS fails (HA)a secondary becomes primary and serves licences
DLS before 3.4, vGPU 18licensing failures; upgrade DLS first
Compute on bare metalno licence server involved

NVIDIA AI Enterprise licensing page (infrastructure release 8, updated 2 September 2026); vGPU client licensing guide, introduction and advanced topics (updated 29 September 2026); NVIDIA License System user guide v3.6.1, sections 1.1.1 and 1.4.

The current graphics guide does not list which functions are limited. It says that a client without a licence “will periodically retry its license request to the license server”, and its introduction gives the two steps of degradation, after 20 minutes and after 24 hours. A firewall change between the VM networks and the DLS can go unnoticed while the licences last, up to a day for desktops and, for compute VMs, the period configured on the DLS (seven days is the CLS default), and then affects every VM at once. Monitor the licence state of a sample of clients and the DLS event log, and add the client ports to the change checklist of the firewall team.

We supply NVIDIA AI Enterprise and vGPU licences on the same quote as the GPUs, and our engineering partner Vixen.UNO can deploy and support the licence service. Tell us how many vGPU VMs run at each site and whether they reach the internet.

Sizing for 200 to 1,000 VMs: one DLS cluster or several

Client count rarely decides the number of DLS instances. NVIDIA measured one instance with 4 vCPU and 8 GB at 2.6 GHz, with licences borrowed for 12 hours and renewed at 15 per cent of that time, and it served up to 32,000 clients; 8 vCPU and 16 GB raised the figure to 84,000. The guide adds that “The maximum number of clients is directly proportional to the length of time for which licenses are borrowed.” In its burst test, the same 4 vCPU instance served 1,000 clients that arrived within 20 seconds in 50 seconds, results that NVIDIA calls “illustrative only”.

For an estate of 200 to 1,000 vGPU VMs, the minimum appliance therefore carries the load, including the boot storm of a whole desktop pool, with a wide margin. The number of clusters follows from sites, network zones and failure domains. Take a company with 600 virtual desktops on L40S and L4 hosts at the main site, 150 desktops at a second site and 24 compute VMs on H200 NVL for an internal assistant. One HA cluster of two instances at the main site, behind a virtual IP, serves all 774 VMs. The second site then depends on the WAN link for its licences, and its desktops run for up to a day if the link fails.

Where the second site must license its VMs on its own, as the recovery site of a DR plan for example, give it its own cluster. Licences are allocated to a licence server on the portal, and each licence server is bound to one service instance, so each site holds its own share. A recovery plan that moves desktops to the second site needs enough licences allocated there, or a change of allocation followed by a new licence file on that DLS. Sizing the desktops and GPUs themselves is covered in our guide to VDI GPU server sizing.

We quote the licences for each site together with the GPUs they run on. Describe your sites, desktop pools and compute VMs in the form below, including any site that has to license its VMs on its own during a failover.

Renewals, upgrades and the three-year licence file

From NLS release 3.4, “the license server files bound to DLS SI will have a validity of three years”, counted from the download; the portal shows the validity in the licence server’s details. Keep that date in the same calendar as the subscription renewals. A renewed subscription changes the entitlement on the portal, and a DLS learns of it only when you download and install the licence file again.

NVIDIA’s prerequisites for vGPU 18 and later say to upgrade the licence server “to a minimum version of DLS 3.4” before upgrading vGPU to 18.0 or later, then to download a fresh licence server file and install it on the upgraded appliance. A DLS appliance is upgraded by migration to a new one: the licence servers, the user registration, the IP address and the service instance move across, while event records on the old appliance are not migrated. If you use the event records for licence usage reports, save what you need from them before the migration. Plan the DLS upgrade as its own change in an agreed maintenance window with a rollback plan, ahead of the host and guest driver upgrade.

What we supply

We supply NVIDIA AI Enterprise licences for compute vGPU and the RTX vWS, vPC and vApps licences for graphics, on one EU contract and invoice with the GPUs they run on, and we size them with the hardware so that the licence lines are in the quote. We supply the L4, L40S and H200 NVL from the example above and the RTX PRO 6000 Server Edition, with manufacturer warranty, as cards or in AI servers built to order. Our NVIDIA AI Enterprise and vGPU licence page explains how each licence is counted, and the full range of cards is on our professional GPUs page. If you want it, our engineering partner Vixen.UNO designs, deploys and supports the setup under the same contract.

FAQ

What is the NVIDIA licence server?
It is a service instance of the NVIDIA License System that hands out vGPU and NVIDIA AI Enterprise licences to virtual machines. It runs either as a Cloud License Service (CLS) instance hosted on the NVIDIA Licensing Portal or as a Delegated License Service (DLS) instance, a virtual appliance or container on your own network. NVIDIA states that AI Enterprise needs it only for vGPU for Compute drivers, so servers running models on bare metal or in containers do not use it, while RTX vWS graphics on passthrough or bare-metal GPUs is licensed through it.
What is the difference between NVIDIA CLS and DLS?
A CLS instance is hosted by NVIDIA on its Licensing Portal, and clients reach it over outbound port 443. A DLS instance runs on-premise, needs at least 4 vCPU, 8 GB RAM and 15 GB disk, and receives its licences as a file that you download from the portal and upload by hand. DLS works on air-gapped sites and supports high-availability clusters of two to nine instances.
How do I set up an NVIDIA DLS?
Deploy the appliance, register the DLS administrator user, create a licence server on the NVIDIA Licensing Portal and register the DLS instance there with its instance token. Then bind the licence server to the instance, install the licence file on the appliance and generate a client configuration token. Copy the token to every client and restart its licensing service.
Where does the NVIDIA client configuration token go?
The location differs by guest OS: on Linux it goes into /etc/nvidia/ClientConfigToken, with the FeatureType parameter set in /etc/nvidia/gridd.conf and a restart of the nvidia-gridd service. On Windows guests it goes into the ClientConfigToken folder of the NVIDIA vGPU Licensing directory, followed by a restart of the NvDisplayContainer service. Adding it to the golden image or VM template lets new VMs license themselves on first boot.
What happens to a vGPU VM in an unlicensed state?
A VM that boots without a licence runs at full capability for 20 minutes. A compute VM on a C-series profile then drops to an idle level of compute until it obtains a licence, while a graphics VM is degraded and degraded further after 24 hours. A licensed compute VM that loses the licence server keeps working for 7 days by default on CLS, a period that is configurable on a DLS, and NVIDIA states that full capability returns immediately once a licence is acquired.
How many VMs can one NVIDIA DLS instance serve?
In NVIDIA’s measurements, one DLS instance with 4 vCPU and 8 GB RAM served up to 32,000 clients with 12-hour licence borrowing, and 8 vCPU with 16 GB up to 84,000. An estate of 200 to 1,000 vGPU VMs therefore fits on the minimum appliance. Plan the number of DLS clusters by site and failure domain, with at least two instances in each for high availability.

Send us the number of vGPU VMs and virtual desktops per site, the GPUs and hypervisor, and whether the VMs may reach the internet. We reply within one business day with the licences your setup needs and a written quote for them with the GPUs.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna