BLOG · GUIDE ·

L40S vGPU profiles: every Q, B, A and C type with its frame buffer, VMs per card and licence

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • The L40S has no MIG, so it is shared only by time-sliced vGPU; NVIDIA lists ten frame buffer sizes from 1 GB (L40S-1Q) to the whole 48 GB card (L40S-48Q)
  • One L40S holds at most 32 vGPUs with a 1 GB type in equal-size mode and 16 in mixed-size mode; L40S-12Q gives 12 GB to each of four VMs in either mode
  • Q types need an RTX vWS licence, B types vPC or vWS, A types vApps, and the C types from L40S-4C to L40S-48C need NVIDIA AI Enterprise
  • On ESXi 8.0 Update 3 and on VCF 9, A, B and Q types of different sizes can share one card in mixed-size mode, which lowers the maximum for some sizes
  • Best effort is the default scheduler; equal share and fixed share give each vGPU a defined share of the card, and vGPU 20.0 added fixed share for vGPUs of different sizes on one GPU

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

L40S vGPU profile list

The NVIDIA L40S has no MIG, so a card is shared between virtual machines only by time-sliced vGPU: each VM gets a fixed part of the 48 GB frame buffer and takes turns on the GPU’s engines. NVIDIA’s vGPU types reference lists the L40S vGPU profiles in ten frame buffer sizes, from 1 GB (L40S-1Q) to the whole card (L40S-48Q), in three series: Q for virtual workstations, B for virtual desktops and A for virtual applications. The C series for compute, from 4 to 48 GB, is listed in the NVIDIA AI Enterprise documentation. One L40S holds at most 32 vGPUs, with a 1 GB type in equal-size mode, and 16 when sizes are mixed on the card.

FRAME BUFFEREQUAL-SIZE MAXMIXED-SIZE MAXTYPES
48 GB (49,152 MB)11L40S-48Q, L40S-48A, L40S-48C
24 GB (24,576 MB)22L40S-24Q, L40S-24A, L40S-24C
16 GB (16,384 MB)32L40S-16Q, L40S-16A, L40S-16C
12 GB (12,288 MB)44L40S-12Q, L40S-12A, L40S-12C
8 GB (8,192 MB)64L40S-8Q, L40S-8A, L40S-8C
6 GB (6,144 MB)88L40S-6Q, L40S-6A, L40S-6C
4 GB (4,096 MB)128L40S-4Q, L40S-4A, L40S-4C
3 GB (3,072 MB)1616L40S-3Q, L40S-3B, L40S-3A
2 GB (2,048 MB)2416L40S-2Q, L40S-2B, L40S-2A
1 GB (1,024 MB)3216L40S-1Q, L40S-1B, L40S-1A

Maximum vGPUs per GPU. Q, B and A types: NVIDIA Virtual GPU Types Reference, vGPU user guide, last updated 29 September 2026. C types: NVIDIA AI Enterprise, Ada Lovelace vGPU types, Table 112, last updated 2 September 2026.

In every row the number in the type name is the frame buffer in GB, and the maximum per GPU is the same for every series of that size. L40S-12Q has 12,288 MB, and four of them fill one card in either mode. The C series stops at 4 GB, so a compute VM on the L40S never gets less than that.

Q, B, A and C series: use case, licence and displays

NVIDIA’s licensing guide maps RTX vWS to “Q-series NVIDIA vGPUs” and “B-series NVIDIA vGPUs” as well as GPU passthrough, vPC to B-series vGPUs only, and vApps to A-series vGPUs and passthrough. The C types for the L40S are not in the vGPU types reference at all; they are listed in the NVIDIA AI Enterprise documentation under “Required license edition: NVIDIA AI Enterprise”.

SERIESNVIDIA’S USE CASELICENCE EDITIONDISPLAYS PER VGPU
Q, 1 to 48 GBvirtual workstationsRTX vWSup to 4; at 7680×4320, 2 from 8 GB, 1 from 2 to 6 GB
B, 1 to 3 GBvirtual desktopsvPC or RTX vWSup to 4 at 2560×1600 or lower
A, 1 to 48 GBvirtual applicationsvApps1 at up to 1280×1024
C, 4 to 48 GBcomputeNVIDIA AI Enterprise1 at up to 3840×2400

NVIDIA Virtual GPU Types Reference (29 September 2026); NVIDIA vGPU licensing user guide, Table 2 (29 September 2026); NVIDIA AI Enterprise Ada Lovelace vGPU types (2 September 2026). Display counts assume all displays at the same resolution, as NVIDIA states.

CUDA and OpenCL support depends on the series. NVIDIA’s vGPU user guide supports them on “All Q-series vGPU types” of a list of GPUs that includes the L40S, and names no B or A types for the card. NVIDIA AI Enterprise lists vGPUs with more than 40 GB, which on the L40S means L40S-48C only, for training workloads.

The display limits separate the small types. L40S-12Q drives two displays at 7680×4320 or four at 5120×2880 or lower. L40S-2B drives two 3840×2160 displays, while L40S-1B drives one at that resolution and four only at 2560×1600 or lower. A VM without a licence is degraded after 20 minutes and further after 24 hours, per NVIDIA’s licensing guide, which is why the NVIDIA vGPU licence server, DLS or CLS must stay reachable.

Equal-size and mixed-size mode on one L40S

NVIDIA’s release notes for vSphere state: “By default, a GPU or GPU instance supports only vGPUs with the same amount of frame buffer and, therefore, is in equal-size mode.” On ESXi 8.0 Update 3 and later 8.0 updates, and on release 9.0 and later, “Any combination of A-series, B-series, and Q-series vGPUs with any amount of frame buffer can reside on the same physical GPU simultaneously” within the card’s 48 GB. For this, the GPU “must be put into mixed-size mode”.

NVIDIA’s tables show where mixed-size mode lowers the maximum: at 1, 2, 4, 8 and 16 GB. The 3, 6, 12, 24 and 48 GB types keep their equal-size count. The vSphere API names the two modes sameSize and mixedSize, in an enum of the host graphics configuration since API release 8.0.3.0.

That statement names only A, B and Q types. For compute VMs, NVIDIA AI Enterprise documents heterogeneous vGPU, which “allows a single physical GPU to host multiple vGPU profiles”, on vSphere and on Turing and later GPUs, and its L40S table gives the C types the same mixed-size maxima as the other series. The same page states that a heterogeneous configuration “only supports the Best Effort and Equal Share schedulers”. We found no NVIDIA document that puts a C type and a Q, B or A type on one L40S. Broadcom’s blog on GPUs in vSphere 8 Update 3, of 3 July 2024, shows an L40 shared by a VM on l40-16c and one on l40-12q; until NVIDIA documents such a mix for your vGPU release, plan compute VMs on their own cards.

Broadcom’s KB 447105 states that the host policy “Spread VMs across GPUs” applies only to devices in time-sliced same-size mode and is bypassed in mixed-size and MIG mode, where placement aims to keep the GPUs from fragmenting instead of spreading VMs evenly.

vGPU scheduler on the L40S: best effort, equal share and fixed share

NVIDIA’s vGPU user guide offers three schedulers for every supported GPU. With best effort, the default, “The physical GPU’s processing cycles are shared in a way that aims to balance performance across vGPUs”, and a vGPU may use cycles the others leave idle. With equal share, “The physical GPU is shared equally amongst the running vGPUs that reside on it”, so each VM gets more when others stop. With fixed share, “Each vGPU is given a fixed share of the physical GPU’s processing cycles”, and that share stays the same as other VMs start or stop.

Under fixed share, a vGPU’s share in equal-size mode is inversely proportional to the maximum number of its type per GPU, and in mixed-size mode proportional to its part of the frame buffer. By both rules an L40S-12Q gets a quarter of the card. NVIDIA’s table sets the default time slice by the maximum number of vGPUs per GPU, 2 ms for up to eight and 1 ms for more. NVIDIA adds that “If you use the equal share or fixed share vGPU scheduler, the frame-rate limiter (FRL) is disabled.”

nvidia-smi vgpu -ss shows the current state, and the set-scheduler-state subcommand with -p 1, -p 2 or -p 3 sets best effort, equal share or fixed share at once, but only until the driver reloads or the host reboots. To keep a setting on ESXi, set the RmPVMRL registry value through esxcli on the module nvidia-gpu (Ada Lovelace and later), then reload the driver or reboot. RmPVMRL=0x01 selects equal share and RmPVMRL=0x11 fixed share, for all GPUs, or for selected cards with NVreg_RegistryDwordsPerDevice. A change fails while any vGPU is active on that GPU, so schedule it in a maintenance window. How four L40S-12Q VMs share a card under load is shown in our L40S benchmark review.

What changed between vGPU releases for the L40S

NVIDIA’s list of GPUs supported by vGPU, updated on 2 October 2026, gives release 16.1 as the first release for the L40S and marks it “Full Support”. Release 17.2 added “Support for a mixture of different types of time-sliced vGPUs on the same physical GPU starting with vSphere 8 Update 3.”

Release 20.0 added “Support for the fixed share scheduler for time-sliced vGPUs with different amounts of frame buffer on the same physical GPU” and withdrew an option: “Disabling strict round robin policy is no longer supported.” The current release, 20.2 on the R595 branch, pairs vGPU Manager 595.91.04 with the Linux guest driver 595.91.07. It supports VCF 9.1, VCF 9.0 and ESXi 8.0 Update 3 P06 or later. We found no L40S type added or removed in the 20.x release notes. NVIDIA’s display-mode table for vSphere lists the L40S as supplied in display-off mode. vMotion, DRS and passthrough rules are in our guide to GPUs in VMware vSphere, passthrough or vGPU.

Which L40S profile for CAD, office VDI and AI compute VMs

NVIDIA’s vWS sizing guide, last updated on 19 August 2026, gives its user-type recommendations for RTX PRO Blackwell cards rather than the L40S. It recommends 4 GB Q profiles for light users, 4 to 8 GB for medium users and 12 to 24 GB for heavy users, and we apply the same frame buffer steps to the L40S types in the table below.

USER TYPEL40S PROFILELICENCEPER CARD
Desktop, two 4K screensL40S-2BvPC24
Desktop, four QHD screensL40S-1BvPC32
Light CAD, 2D drawingL40S-4QRTX vWS12
Medium 3D CADL40S-8QRTX vWS6
Heavy CAD and visualisationL40S-12Q to L40S-24QRTX vWS4 to 2
Inference VM, small modelL40S-12C or L40S-24CNVIDIA AI Enterprise4 or 2
Training or fine-tuning VML40S-48CNVIDIA AI Enterprise1

Our reading. Frame buffer steps from NVIDIA’s RTX vWS sizing guide (19 August 2026); display limits, maxima per card in equal-size mode and licence editions from NVIDIA’s vGPU types reference and licensing guide; training for vGPUs over 40 GB from NVIDIA AI Enterprise.

NVIDIA’s vPC sizing guide publishes measured frame buffer use for 1B and 2B profiles, with tables for the L40 and L40S, and a pilot on your own applications settles the size per seat. How many seats and hosts an estate of several hundred desktops needs is the subject of our guide to GPU servers for virtual desktops, and the comparison with cards that also offer MIG is in MIG and vGPU: how many VMs one card holds.

We build GPU nodes for virtual workstations with L40S cards and NVIDIA vGPU licensing, sized by seats per card. Tell us your user types, seats and screens in the form below, and we reply with the profiles and card count.

Worked example: eight L40S in two vSphere hosts

This example is illustrative. A company plans 16 designers on heavy CAD, 48 office users who need GPU-accelerated desktops with two 4K screens, and four VMs for small inference services, on two ESXi hosts with four L40S each.

Each host runs two cards with four L40S-12Q (eight CAD seats), one card with 24 L40S-2B (24 desktops) and one card with two L40S-24C (two compute VMs). All cards stay in equal-size mode, so every card reaches the maxima in the first table and the “Spread VMs across GPUs” policy keeps working. The CAD cards run fixed share, set per device with RmPVMRL=0x11, so a designer’s share of the card does not change when a colleague starts a render; the desktop card keeps best effort, which lets busy desktops use cycles idle ones leave.

If one designer per host needs 24 GB, mixed-size mode could put one L40S-24Q next to two L40S-12Q on one card, a mix of Q types that NVIDIA’s notes for vSphere allow. The trade-off is placement, since the spread policy no longer applies to that card. A vGPU VM can live-migrate only to a host with the same GPU type and an unused slot for its profile, so with every slot in use, a host in maintenance takes its VMs offline. Licences follow the series, with RTX vWS for the CAD seats, vPC for the desktops and NVIDIA AI Enterprise for the GPUs that host C types.

We check the rack, power and airflow before we quote L40S hosts. Send us your seat counts per user type through the form below, and we return the profile layout and a configuration within one business day.

What we supply

We supply the NVIDIA L40S with manufacturer warranty on one EU contract and invoice, as cards or in AI servers built to order, including GPU nodes for virtual workstations sized by seats per card. NVIDIA AI Enterprise and vGPU licences come on the same invoice as the hardware, as our NVIDIA AI Enterprise and vGPU licences page describes. Where a workload outgrows 48 GB per VM or needs hardware partitioning, our professional GPU range includes the RTX PRO 6000 Server Edition and the H200 NVL. For vSphere and VCF updates on vGPU hosts, our engineering partner Vixen.UNO delivers VMware optimisation, with changes in agreed maintenance windows with a rollback plan.

FAQ

What vGPU profiles does the NVIDIA L40S support?
NVIDIA lists Q types from L40S-1Q to L40S-48Q, B types at 1, 2 and 3 GB, A types from L40S-1A to L40S-48A, and C types from L40S-4C to L40S-48C in its NVIDIA AI Enterprise documentation. All are time-sliced, because the L40S has no MIG. The frame buffer sizes are 1, 2, 3, 4, 6, 8, 12, 16, 24 and 48 GB.
What is the NVIDIA L40S-12Q profile?
L40S-12Q is a Q-series vGPU type for virtual workstations with 12,288 MB of frame buffer, and four of them fit one L40S in equal-size and in mixed-size mode. It drives two displays at 7680×4320 or four at 5120×2880 or lower, supports CUDA and OpenCL, and needs an NVIDIA RTX vWS licence.
How many VMs can one L40S run with vGPU?
Up to 32, using a 1 GB type (L40S-1Q, L40S-1B or L40S-1A) with all vGPUs on the card the same size. In mixed-size mode, where different sizes share the card, the maximum for 1 GB and 2 GB types is 16. Larger profiles hold fewer VMs, for example six with 8 GB and one with the full 48 GB.
Which licence does an L40S-24C vGPU need?
C-series types such as L40S-24C are compute profiles and need NVIDIA AI Enterprise; they are not part of the vWS, vPC or vApps editions. L40S-24C gives 24 GB of frame buffer, and two fit on one card.
Can I mix different vGPU profiles on one L40S?
On ESXi 8.0 Update 3 or later and on VCF 9, A, B and Q types of different sizes can share one L40S once the GPU is in mixed-size mode, as long as their total stays within 48 GB. Mixed-size mode lowers the maximum for several sizes, for example from six to four L40S-8Q, and the host policy that spreads VMs across GPUs applies only in same-size mode. NVIDIA’s notes for vSphere do not name C types for such a mix, so plan compute profiles on their own cards.
Which licence do L40S vWS profiles need?
The Q-series profiles of the L40S need an NVIDIA RTX vWS licence, which also covers B-series profiles and GPU passthrough. B-series desktops can instead run on vPC, A-series application profiles on vApps, and C-series compute profiles on NVIDIA AI Enterprise.

Send us the user types and seat counts, the screens each user drives, the compute VMs you plan and your vSphere or VCF version. We reply within one business day with the L40S profiles and card count that fit, the vGPU and NVIDIA AI Enterprise licences they need, and a configuration and quote for the hosts.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna