NVIDIA L4 vGPU profiles: the profile list, max VMs per GPU and VDI users per card
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- The NVIDIA L4 has no MIG, so it is shared only through time-sliced vGPU; NVIDIA lists Q-series profiles from 1 GB to 24 GB, B-series from 1 GB to 3 GB, A-series from 24 GB downwards and C-series from 4 GB to 24 GB
- Maximum VMs per L4 in equal-size mode: 24 on L4-1B or L4-1Q, 12 on L4-2B, 8 on L4-3B, 6 on L4-4Q or L4-4C and 1 on L4-24Q or L4-24C; in mixed-size mode the 1 GB maximum falls to 16 and the 2 GB maximum to 8
- Licence editions per NVIDIA: Q-series needs vWS, B-series vPC or vWS, A-series vApps and C-series NVIDIA AI Enterprise; vPC and vWS are counted per concurrent user
- Each L4 has 2 NVENC encoders, 4 NVDEC decoders and 4 JPEG decoders, so 12 desktops on L4-2B share two encoders, six desktops per encoder
- NVIDIA’s vPC density table allows 16 L4 per 2U server as a maximum, while Lenovo lists 8 L4 for the SR650 V3 and 10 for the SR650 V4; eight L4 on L4-2B hold 96 desktops at 576 W of card power
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
NVIDIA L4 vGPU profiles at a glance
The NVIDIA L4 is shared between virtual machines only by time-sliced vGPU, because it has no Multi-Instance GPU (MIG): NVIDIA’s MIG user guide, updated on 11 September 2026, does not list it among the supported GPUs. Every L4 vGPU profile gives a VM a fixed slice of the card’s 24 GB of frame buffer, while the VMs take turns on its compute. NVIDIA’s vGPU documentation lists four series for the L4, Q for virtual workstations, B for virtual desktops, A for published applications and C for compute, with between 1 and 24 VMs per card depending on the frame buffer size.
| PROFILE | FRAME BUFFER | MAX PER GPU | MIXED-SIZE MAX | LICENCE EDITION |
|---|---|---|---|---|
| L4-24Q | 24 GB | 1 | 1 | vWS |
| L4-12Q | 12 GB | 2 | 2 | vWS |
| L4-8Q | 8 GB | 3 | 2 | vWS |
| L4-6Q | 6 GB | 4 | 4 | vWS |
| L4-4Q | 4 GB | 6 | 4 | vWS |
| L4-3Q | 3 GB | 8 | 8 | vWS |
| L4-2Q | 2 GB | 12 | 8 | vWS |
| L4-1Q | 1 GB | 24 | 16 | vWS |
| L4-3B | 3 GB | 8 | 8 | vPC or vWS |
| L4-2B | 2 GB | 12 | 8 | vPC or vWS |
| L4-1B | 1 GB | 24 | 16 | vPC or vWS |
| L4-24A | 24 GB | 1 | 1 | vApps |
| L4-12A | 12 GB | 2 | 2 | vApps |
| L4-24C | 24 GB | 1 | 1 | NVIDIA AI Enterprise |
| L4-12C | 12 GB | 2 | 2 | NVIDIA AI Enterprise |
| L4-8C | 8 GB | 3 | 2 | NVIDIA AI Enterprise |
| L4-6C | 6 GB | 4 | 4 | NVIDIA AI Enterprise |
| L4-4C | 4 GB | 6 | 4 | NVIDIA AI Enterprise |
NVIDIA Virtual GPU Types Reference in the vGPU user guide (updated 29 September 2026) for the Q and B-series; the vGPU 19.0 edition of the same reference for the A-series; NVIDIA AI Enterprise 8.2, Ada Lovelace vGPU types (updated 2 September 2026) for the C-series. NVIDIA gives Q, B and A sizes in MB, 1,024 MB per GB.
NVIDIA’s A-series table for the L4 continues below L4-12A with L4-8A at three per card; check the smaller A sizes in the reference before you plan published-app hosts on them. The C-series stops at 4 GB, so the L4 has no 1 GB, 2 GB or 3 GB compute profile. NVIDIA’s footnote to the compute types says they support “only a single display head” and provide no Quadro graphics acceleration, and each runs one display at up to 3840×2400. For how the L4 compares with the L40S as an inference card, see our comparison of the L4 and L40S for inference; the same reference for the larger card is in our guide to L40S vGPU profiles.
Time-sliced vGPU on the L4: equal and mixed sizes
NVIDIA’s RTX vWS sizing guide, updated on 19 August 2026, states: “When MIG is disabled, all vGPUs hosted on a physical GPU must have the same profile size (same frame buffer size)”. Profiles of the same size from different series may share a card, such as 2Q next to 2B in NVIDIA’s example. A card with twelve L4-2B desktops can therefore take an L4-2Q workstation in place of one of them, but not an L4-4Q.
Heterogeneous vGPU, which NVIDIA says “was introduced in vGPU 17.0”, lifts that rule and lets one card hold different sizes. The guide’s L4 example is a card hosting L4-8Q and L4-2B instances at the same time. On vSphere, NVIDIA’s validated platforms list mixed sizes for the L4 on release 9.0 and on ESXi 8.0 Update 3 and later, which every release in the vGPU 20 support matrix meets. In mixed-size mode several sizes have a lower maximum, shown in the mixed-size column of the table: L4-8Q falls from 3 to 2, L4-4Q and L4-4C from 6 to 4, the 2 GB profiles from 12 to 8 and the 1 GB profiles from 24 to 16. For a host with several L4 cards, giving each user group its own cards in equal-size mode keeps the higher counts.
Each VM keeps its own frame buffer, while compute is shared in time slices. The vWS guide says the vGPU software “uses the best effort scheduler by default”, and its example configurations add that best effort often oversubscribes the GPU compute engine two to three times, while the fixed share scheduler holds a set quality of service for each VM. Best effort suits office desktops whose load comes and goes, and a fixed share per VM suits users who need the same performance at the busiest hour.
Video encoders and displays for VDI streams
NVIDIA’s L4 product page lists the card with 24 GB and 300 GB/s of memory bandwidth at 72 W, in a “1-slot low-profile” PCIe Gen4 x16 format. It has two NVENC encoders, four NVDEC decoders and four JPEG decoders. For video it states that servers with L4 host “up to 1,040 concurrent AV1 video streams at 720p30”, measured on 8 L4 with a low-latency AV1 encode preset.
A remote display protocol that encodes on the GPU sends every desktop’s screen through these two encoders. By our arithmetic, twelve desktops on L4-2B put six on each encoder, and twenty-four on L4-1B put twelve on each. NVIDIA’s vPC sizing guide names encoder and decoder utilisation among the metrics that can be monitored and logged with nvidia-smi. We found no session limit per encoder in NVIDIA’s sizing guides, so encoder load belongs in the pilot next to frame buffer.
The profile also caps displays. The B-series profiles drive up to four screens at 2560×1600 or lower, L4-2B and L4-3B two at 3840×2160 and L4-1B one. L4-2Q and L4-3Q drive four at 3840×2400 or lower, and from L4-4Q upwards four screens at 5120×2880 or lower. The A-series profiles have a single display at up to 1280×1024, which NVIDIA’s licensing guide applies only to the console display in remote application environments.
Which L4 profile for office users and light CAD
For knowledge workers, NVIDIA’s vPC sizing guide states that modern operating systems and evolving application workloads “require larger frame buffer sizes” and that “larger profile sizes (2B or 3B) are often needed to maintain a consistent user experience”. This guide plans office desktops on L4-2B and moves a user group to L4-3B when the pilot shows its frame buffer close to the limit.
For CAD, the RTX vWS guide’s typical deployment puts light users on a 4 GB profile, DC-4Q on the RTX PRO 4500 Blackwell, at 6 to 8 users per GPU with 4 vCPUs and 8 to 12 GB of RAM each. Medium users get DC-4Q or DC-8Q at 3 to 6 per GPU, and heavy users DC-12Q to DC-24Q on the RTX PRO 6000 Blackwell at 2 to 4 per GPU. On the L4 the same frame buffer sizes are L4-4Q and L4-8Q, which allow 6 and 3 VMs per card. NVIDIA’s user counts are for Blackwell cards with more compute than the L4, so the pilot shows how many CAD users one L4 carries. Heavy 3D users at 12 or 24 GB would get one or two per L4, so they belong on larger cards.
| USER TYPE | L4 PROFILE | LICENCE | VMS PER L4 | BASIS |
|---|---|---|---|---|
| Office, HD screens | L4-2B | vPC | 12 | vPC guide: 2B or 3B often needed |
| Office, 4K or heavy use | L4-3B | vPC | 8 | more frame buffer per desktop |
| Light CAD and 2D design | L4-4Q | vWS | 6 | 4 GB as for light users in the vWS guide |
| Medium 3D | L4-8Q | vWS | 3 | 8 GB as for medium users |
| Published app hosts | L4-12A or L4-24A | vApps | 2 or 1 | one VM serves many app users |
| Inference service | L4-4C to L4-24C | NVIDIA AI Enterprise | 6 to 1 | model and cache must fit the slice |
Profiles and maximums from NVIDIA’s vGPU types references (table 1); the mapping of user types to L4 profiles is our reading of NVIDIA’s vPC sizing guide and the example configurations in its RTX vWS sizing guide (both updated 19 August 2026), to be confirmed in a pilot.
We build GPU nodes for virtual workstations with the L4, sized by seats per card. Describe your user groups, their screens and their applications in the form below, and we reply with the profile per group and the cards it needs.
Inference VMs on L4 C-series profiles
An L4 C-series profile makes sense when a small inference service needs a VM of its own, for example an embedding or reranker model for RAG, speech recognition or a small language model, kept apart from other teams’ services. L4-12C gives two such VMs 12 GB each, L4-4C gives six VMs 4 GB each, and the model weights plus their cache must fit the slice. The VMs share the card’s compute and its 300 GB/s of bandwidth in time slices, so two busy services on one L4 each run slower than one alone.
When one model needs the whole card, L4-24C gives one VM the full 24 GB. NVIDIA’s Ada vGPU types page names NVIDIA AI Enterprise as the required licence for all C-series types, and our article on NVIDIA AI Enterprise licensing explains how it is counted. NVIDIA’s vSphere notes on mixing types name only A, B and Q-series vGPUs, so we plan inference VMs on L4 cards of their own.
How many L4 cards fit in a VDI host
The L4 suits dense VDI hosts because it is a 72 W single-slot, low-profile card. NVIDIA’s vPC sizing guide describes it as a “Compact single-slot GPU for graphics, video, and inference workloads” and lists it among its recommended GPUs for vPC, with the RTX PRO 4500 Blackwell Server Edition as “the primary recommended GPU”. Our comparison of the L4 and RTX PRO 4500 Server Edition covers that choice.
NVIDIA’s maximum density table for a 2U server lists the L4 at 24 users per board with a 1 GB profile, 16 boards and 384 users per server, and says the table “reflects maximum supported density rather than a recommended deployment point”. Lenovo’s L4 product guide (LP1717, read on 10 October 2026) lists up to 8 L4 in the ThinkSystem SR650 V3, 10 in the SR650 V4 and 8 in the SR675 V3. Check the maker’s list for the exact server model and riser before you plan the card count.
Worked example: a VDI host with 8 L4
This example is our estimate from NVIDIA’s maximums, not a measured result. Take a 2U host with eight L4 for 72 office users and 12 light CAD users. Six cards run L4-2B in equal-size mode, 12 desktops each, and two cards run L4-4Q at their maximum of 6 workstations each, a CAD count per L4 the pilot has to confirm, for 84 sessions on the host. The cards draw 8 × 72 W = 576 W before processors, memory and fans. The CAD VMs need 48 vCPUs at NVIDIA’s 4 vCPUs per light user.
The same host with all eight cards on L4-2B holds 96 desktops and with L4-3B 64. The licence count follows the peak, 72 vPC and 12 vWS concurrent users in the first layout. Sizing a whole estate of such hosts with a spare host for failover is the subject of our VDI GPU sizing guide.
We check the rack, power and airflow before we quote, and the vGPU and AI Enterprise licences come on the same invoice as the cards. Send us your peak sessions per user group through the form below.
Hypervisor support and licences for L4 vGPU
NVIDIA’s support matrix for VMware vSphere, updated on 29 September 2026, lists the L4 for VMware Cloud Foundation 9.1 and 9.0 and for vSphere 8.0, where the current release “Requires VMware vSphere Hypervisor (ESXi) 8.0u3 P06 and later updates to release 8.0 unless explicitly stated otherwise”. The same matrix requires the vSphere Foundation edition of ESXi or a vSphere Enterprise Plus licence for vGPU. The R595 release notes for vSphere (vGPU 20.0 to 20.2) state: “You must use NVIDIA License System with every release in this release family”. Our article on the vGPU licence server, DLS or CLS covers that service.
vPC, vWS and vApps are licensed by concurrent user, which NVIDIA’s packaging and licensing guide (updated 24 March 2026) defines as “A method of allocating licenses based on the number of VMs that are concurrently being used”. vPC supports up to four displays at up to 5120×2880, vWS up to 7680×4320 with CUDA and OpenCL, and vApps one display. An L4 host with B, Q and C profiles therefore needs vPC and vWS licences per concurrent user and NVIDIA AI Enterprise for the inference cards.
What we supply
We supply the NVIDIA L4 as a card for your own servers, listed on our GPU page, and in AI servers built to order as VDI nodes with NVIDIA vGPU licensing, sized by seats per card. Both come with manufacturer warranty, on one EU contract and invoice. NVIDIA vPC, RTX vWS and vApps licences and NVIDIA AI Enterprise for C-series profiles come on the same invoice through our software and licensing offer. Configuration and quote follow within one business day, with the rack, power and airflow checked before we quote.
FAQ
What vGPU profiles does the NVIDIA L4 support?
How many VDI users per L4 GPU?
What is the L4-4Q profile?
What is the L4-24Q profile used for?
What is the maximum number of VMs on one L4 with vGPU?
Does the NVIDIA L4 support MIG?
Send us your user groups with their peak concurrent sessions, the screens and resolutions per user, the applications, any inference services planned for the same hosts and your hypervisor version. We reply within one business day with the L4 profile per group, the cards per host, the licence editions and a configuration and quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day