BLOG · GUIDE ·

NVIDIA L4 vGPU profiles: the profile list, max VMs per GPU and VDI users per card

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • The NVIDIA L4 has no MIG, so it is shared only through time-sliced vGPU; NVIDIA lists Q-series profiles from 1 GB to 24 GB, B-series from 1 GB to 3 GB, A-series from 24 GB downwards and C-series from 4 GB to 24 GB
  • Maximum VMs per L4 in equal-size mode: 24 on L4-1B or L4-1Q, 12 on L4-2B, 8 on L4-3B, 6 on L4-4Q or L4-4C and 1 on L4-24Q or L4-24C; in mixed-size mode the 1 GB maximum falls to 16 and the 2 GB maximum to 8
  • Licence editions per NVIDIA: Q-series needs vWS, B-series vPC or vWS, A-series vApps and C-series NVIDIA AI Enterprise; vPC and vWS are counted per concurrent user
  • Each L4 has 2 NVENC encoders, 4 NVDEC decoders and 4 JPEG decoders, so 12 desktops on L4-2B share two encoders, six desktops per encoder
  • NVIDIA’s vPC density table allows 16 L4 per 2U server as a maximum, while Lenovo lists 8 L4 for the SR650 V3 and 10 for the SR650 V4; eight L4 on L4-2B hold 96 desktops at 576 W of card power

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

NVIDIA L4 vGPU profiles at a glance

The NVIDIA L4 is shared between virtual machines only by time-sliced vGPU, because it has no Multi-Instance GPU (MIG): NVIDIA’s MIG user guide, updated on 11 September 2026, does not list it among the supported GPUs. Every L4 vGPU profile gives a VM a fixed slice of the card’s 24 GB of frame buffer, while the VMs take turns on its compute. NVIDIA’s vGPU documentation lists four series for the L4, Q for virtual workstations, B for virtual desktops, A for published applications and C for compute, with between 1 and 24 VMs per card depending on the frame buffer size.

PROFILEFRAME BUFFERMAX PER GPUMIXED-SIZE MAXLICENCE EDITION
L4-24Q24 GB11vWS
L4-12Q12 GB22vWS
L4-8Q8 GB32vWS
L4-6Q6 GB44vWS
L4-4Q4 GB64vWS
L4-3Q3 GB88vWS
L4-2Q2 GB128vWS
L4-1Q1 GB2416vWS
L4-3B3 GB88vPC or vWS
L4-2B2 GB128vPC or vWS
L4-1B1 GB2416vPC or vWS
L4-24A24 GB11vApps
L4-12A12 GB22vApps
L4-24C24 GB11NVIDIA AI Enterprise
L4-12C12 GB22NVIDIA AI Enterprise
L4-8C8 GB32NVIDIA AI Enterprise
L4-6C6 GB44NVIDIA AI Enterprise
L4-4C4 GB64NVIDIA AI Enterprise

NVIDIA Virtual GPU Types Reference in the vGPU user guide (updated 29 September 2026) for the Q and B-series; the vGPU 19.0 edition of the same reference for the A-series; NVIDIA AI Enterprise 8.2, Ada Lovelace vGPU types (updated 2 September 2026) for the C-series. NVIDIA gives Q, B and A sizes in MB, 1,024 MB per GB.

NVIDIA’s A-series table for the L4 continues below L4-12A with L4-8A at three per card; check the smaller A sizes in the reference before you plan published-app hosts on them. The C-series stops at 4 GB, so the L4 has no 1 GB, 2 GB or 3 GB compute profile. NVIDIA’s footnote to the compute types says they support “only a single display head” and provide no Quadro graphics acceleration, and each runs one display at up to 3840×2400. For how the L4 compares with the L40S as an inference card, see our comparison of the L4 and L40S for inference; the same reference for the larger card is in our guide to L40S vGPU profiles.

Time-sliced vGPU on the L4: equal and mixed sizes

NVIDIA’s RTX vWS sizing guide, updated on 19 August 2026, states: “When MIG is disabled, all vGPUs hosted on a physical GPU must have the same profile size (same frame buffer size)”. Profiles of the same size from different series may share a card, such as 2Q next to 2B in NVIDIA’s example. A card with twelve L4-2B desktops can therefore take an L4-2Q workstation in place of one of them, but not an L4-4Q.

Heterogeneous vGPU, which NVIDIA says “was introduced in vGPU 17.0”, lifts that rule and lets one card hold different sizes. The guide’s L4 example is a card hosting L4-8Q and L4-2B instances at the same time. On vSphere, NVIDIA’s validated platforms list mixed sizes for the L4 on release 9.0 and on ESXi 8.0 Update 3 and later, which every release in the vGPU 20 support matrix meets. In mixed-size mode several sizes have a lower maximum, shown in the mixed-size column of the table: L4-8Q falls from 3 to 2, L4-4Q and L4-4C from 6 to 4, the 2 GB profiles from 12 to 8 and the 1 GB profiles from 24 to 16. For a host with several L4 cards, giving each user group its own cards in equal-size mode keeps the higher counts.

Each VM keeps its own frame buffer, while compute is shared in time slices. The vWS guide says the vGPU software “uses the best effort scheduler by default”, and its example configurations add that best effort often oversubscribes the GPU compute engine two to three times, while the fixed share scheduler holds a set quality of service for each VM. Best effort suits office desktops whose load comes and goes, and a fixed share per VM suits users who need the same performance at the busiest hour.

Video encoders and displays for VDI streams

NVIDIA’s L4 product page lists the card with 24 GB and 300 GB/s of memory bandwidth at 72 W, in a “1-slot low-profile” PCIe Gen4 x16 format. It has two NVENC encoders, four NVDEC decoders and four JPEG decoders. For video it states that servers with L4 host “up to 1,040 concurrent AV1 video streams at 720p30”, measured on 8 L4 with a low-latency AV1 encode preset.

A remote display protocol that encodes on the GPU sends every desktop’s screen through these two encoders. By our arithmetic, twelve desktops on L4-2B put six on each encoder, and twenty-four on L4-1B put twelve on each. NVIDIA’s vPC sizing guide names encoder and decoder utilisation among the metrics that can be monitored and logged with nvidia-smi. We found no session limit per encoder in NVIDIA’s sizing guides, so encoder load belongs in the pilot next to frame buffer.

The profile also caps displays. The B-series profiles drive up to four screens at 2560×1600 or lower, L4-2B and L4-3B two at 3840×2160 and L4-1B one. L4-2Q and L4-3Q drive four at 3840×2400 or lower, and from L4-4Q upwards four screens at 5120×2880 or lower. The A-series profiles have a single display at up to 1280×1024, which NVIDIA’s licensing guide applies only to the console display in remote application environments.

Which L4 profile for office users and light CAD

For knowledge workers, NVIDIA’s vPC sizing guide states that modern operating systems and evolving application workloads “require larger frame buffer sizes” and that “larger profile sizes (2B or 3B) are often needed to maintain a consistent user experience”. This guide plans office desktops on L4-2B and moves a user group to L4-3B when the pilot shows its frame buffer close to the limit.

For CAD, the RTX vWS guide’s typical deployment puts light users on a 4 GB profile, DC-4Q on the RTX PRO 4500 Blackwell, at 6 to 8 users per GPU with 4 vCPUs and 8 to 12 GB of RAM each. Medium users get DC-4Q or DC-8Q at 3 to 6 per GPU, and heavy users DC-12Q to DC-24Q on the RTX PRO 6000 Blackwell at 2 to 4 per GPU. On the L4 the same frame buffer sizes are L4-4Q and L4-8Q, which allow 6 and 3 VMs per card. NVIDIA’s user counts are for Blackwell cards with more compute than the L4, so the pilot shows how many CAD users one L4 carries. Heavy 3D users at 12 or 24 GB would get one or two per L4, so they belong on larger cards.

USER TYPEL4 PROFILELICENCEVMS PER L4BASIS
Office, HD screensL4-2BvPC12vPC guide: 2B or 3B often needed
Office, 4K or heavy useL4-3BvPC8more frame buffer per desktop
Light CAD and 2D designL4-4QvWS64 GB as for light users in the vWS guide
Medium 3DL4-8QvWS38 GB as for medium users
Published app hostsL4-12A or L4-24AvApps2 or 1one VM serves many app users
Inference serviceL4-4C to L4-24CNVIDIA AI Enterprise6 to 1model and cache must fit the slice

Profiles and maximums from NVIDIA’s vGPU types references (table 1); the mapping of user types to L4 profiles is our reading of NVIDIA’s vPC sizing guide and the example configurations in its RTX vWS sizing guide (both updated 19 August 2026), to be confirmed in a pilot.

We build GPU nodes for virtual workstations with the L4, sized by seats per card. Describe your user groups, their screens and their applications in the form below, and we reply with the profile per group and the cards it needs.

Inference VMs on L4 C-series profiles

An L4 C-series profile makes sense when a small inference service needs a VM of its own, for example an embedding or reranker model for RAG, speech recognition or a small language model, kept apart from other teams’ services. L4-12C gives two such VMs 12 GB each, L4-4C gives six VMs 4 GB each, and the model weights plus their cache must fit the slice. The VMs share the card’s compute and its 300 GB/s of bandwidth in time slices, so two busy services on one L4 each run slower than one alone.

When one model needs the whole card, L4-24C gives one VM the full 24 GB. NVIDIA’s Ada vGPU types page names NVIDIA AI Enterprise as the required licence for all C-series types, and our article on NVIDIA AI Enterprise licensing explains how it is counted. NVIDIA’s vSphere notes on mixing types name only A, B and Q-series vGPUs, so we plan inference VMs on L4 cards of their own.

How many L4 cards fit in a VDI host

The L4 suits dense VDI hosts because it is a 72 W single-slot, low-profile card. NVIDIA’s vPC sizing guide describes it as a “Compact single-slot GPU for graphics, video, and inference workloads” and lists it among its recommended GPUs for vPC, with the RTX PRO 4500 Blackwell Server Edition as “the primary recommended GPU”. Our comparison of the L4 and RTX PRO 4500 Server Edition covers that choice.

NVIDIA’s maximum density table for a 2U server lists the L4 at 24 users per board with a 1 GB profile, 16 boards and 384 users per server, and says the table “reflects maximum supported density rather than a recommended deployment point”. Lenovo’s L4 product guide (LP1717, read on 10 October 2026) lists up to 8 L4 in the ThinkSystem SR650 V3, 10 in the SR650 V4 and 8 in the SR675 V3. Check the maker’s list for the exact server model and riser before you plan the card count.

Worked example: a VDI host with 8 L4

This example is our estimate from NVIDIA’s maximums, not a measured result. Take a 2U host with eight L4 for 72 office users and 12 light CAD users. Six cards run L4-2B in equal-size mode, 12 desktops each, and two cards run L4-4Q at their maximum of 6 workstations each, a CAD count per L4 the pilot has to confirm, for 84 sessions on the host. The cards draw 8 × 72 W = 576 W before processors, memory and fans. The CAD VMs need 48 vCPUs at NVIDIA’s 4 vCPUs per light user.

The same host with all eight cards on L4-2B holds 96 desktops and with L4-3B 64. The licence count follows the peak, 72 vPC and 12 vWS concurrent users in the first layout. Sizing a whole estate of such hosts with a spare host for failover is the subject of our VDI GPU sizing guide.

We check the rack, power and airflow before we quote, and the vGPU and AI Enterprise licences come on the same invoice as the cards. Send us your peak sessions per user group through the form below.

Hypervisor support and licences for L4 vGPU

NVIDIA’s support matrix for VMware vSphere, updated on 29 September 2026, lists the L4 for VMware Cloud Foundation 9.1 and 9.0 and for vSphere 8.0, where the current release “Requires VMware vSphere Hypervisor (ESXi) 8.0u3 P06 and later updates to release 8.0 unless explicitly stated otherwise”. The same matrix requires the vSphere Foundation edition of ESXi or a vSphere Enterprise Plus licence for vGPU. The R595 release notes for vSphere (vGPU 20.0 to 20.2) state: “You must use NVIDIA License System with every release in this release family”. Our article on the vGPU licence server, DLS or CLS covers that service.

vPC, vWS and vApps are licensed by concurrent user, which NVIDIA’s packaging and licensing guide (updated 24 March 2026) defines as “A method of allocating licenses based on the number of VMs that are concurrently being used”. vPC supports up to four displays at up to 5120×2880, vWS up to 7680×4320 with CUDA and OpenCL, and vApps one display. An L4 host with B, Q and C profiles therefore needs vPC and vWS licences per concurrent user and NVIDIA AI Enterprise for the inference cards.

What we supply

We supply the NVIDIA L4 as a card for your own servers, listed on our GPU page, and in AI servers built to order as VDI nodes with NVIDIA vGPU licensing, sized by seats per card. Both come with manufacturer warranty, on one EU contract and invoice. NVIDIA vPC, RTX vWS and vApps licences and NVIDIA AI Enterprise for C-series profiles come on the same invoice through our software and licensing offer. Configuration and quote follow within one business day, with the rack, power and airflow checked before we quote.

FAQ

What vGPU profiles does the NVIDIA L4 support?
NVIDIA lists Q-series profiles from L4-1Q to L4-24Q for virtual workstations, B-series L4-1B to L4-3B for virtual desktops, A-series for published applications and C-series L4-4C to L4-24C for compute. All are time-sliced, because the L4 has no MIG. Q-series needs a vWS licence, B-series vPC or vWS, A-series vApps and C-series NVIDIA AI Enterprise.
How many VDI users per L4 GPU?
NVIDIA’s maximum is 24 desktops per L4 on the 1 GB L4-1B profile, 12 on L4-2B and 8 on L4-3B in equal-size mode. NVIDIA’s vPC sizing guide says 2B or 3B profiles are often needed for a consistent user experience, so 8 to 12 office users per card is the planning range before a pilot confirms it.
What is the L4-4Q profile?
L4-4Q is the 4 GB virtual workstation profile of the NVIDIA L4, with up to 6 VMs per card in equal-size mode and 4 in mixed-size mode. It needs an NVIDIA RTX vWS licence and drives up to four displays at 5120×2880 or lower. It matches the 4 GB profile NVIDIA’s vWS sizing guide uses for light CAD users.
What is the L4-24Q profile used for?
L4-24Q gives one virtual workstation the whole 24 GB of the L4, for a single 3D or CAD user who needs all of the frame buffer. It needs a vWS licence, and NVIDIA lists up to two displays at 7680×4320 or four at 5120×2880 or lower. For a compute VM with the full card, the matching profile is L4-24C under NVIDIA AI Enterprise.
What is the maximum number of VMs on one L4 with vGPU?
The maximum is 24 VMs on the 1 GB profiles L4-1B or L4-1Q in equal-size mode, and 16 in mixed-size mode. On compute profiles the maximum is 6 VMs, on L4-4C, because the C-series has no profile below 4 GB.
Does the NVIDIA L4 support MIG?
The L4 does not support MIG, and NVIDIA’s MIG user guide, updated on 11 September 2026, does not list it among the supported GPUs. VMs share the card by time-sliced vGPU, where each VM has its own frame buffer and the compute is scheduled in turns, by default with the best effort scheduler.

Send us your user groups with their peak concurrent sessions, the screens and resolutions per user, the applications, any inference services planned for the same hosts and your hypervisor version. We reply within one business day with the L4 profile per group, the cards per host, the licence editions and a configuration and quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna