BLOG · GUIDE ·

VDI GPU sizing: how many cards and vGPU hosts 200 to 1,000 virtual desktops need

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Size a VDI estate as concurrent users per type × vGPU profile → cards → hosts, plus one spare host per GPU type for failover (N+1); a pilot on your own desktop image confirms the users per card
  • Knowledge workers run on NVIDIA vPC with B-series profiles: NVIDIA’s vPC guide says modern operating systems need larger frame buffers and that 2B or 3B profiles are often needed, and Blackwell cards offer no 1 GB profile
  • NVIDIA’s maximum per card at 2 GB is 12 users on the L4, 24 on the L40S and 32 on the RTX PRO 6000 Server Edition, or 48 with MIG-backed, time-sliced vGPU, which vGPU 20 supports on vSphere from VCF 9.1
  • For CAD and 3D, NVIDIA’s RTX vWS guide gives 3 to 6 medium users per GPU on DC-4Q or DC-8Q and 2 to 4 heavy users on the RTX PRO 6000 with DC-12Q to DC-24Q
  • vPC and RTX vWS are licensed per concurrent user (CCU), so 900 office desktops and 100 workstation sessions at peak need 900 vPC and 100 RTX vWS licences

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

How to size GPU hosts for VDI

For VDI GPU sizing, count the concurrent users of each type, give each type a vGPU profile with a fixed frame buffer, and divide each user type by the number of users one card holds at that profile. Divide the cards by the cards per host and add one host for failover (N+1). Knowledge workers run on NVIDIA Virtual PC (vPC) with B-series profiles of 2 or 3 GB, and users of CAD and 3D applications run on NVIDIA RTX Virtual Workstation (RTX vWS) with Q-series profiles. Both licences are counted per concurrent user.

The result is an estimate, and both of NVIDIA’s sizing guides tell you to confirm it with a proof of concept on the target hardware. The figures below come from NVIDIA’s vPC and RTX vWS sizing guides, both updated on 19 August 2026, from the vGPU 20 documentation updated on 29 September 2026 and from NVIDIA’s product pages, read on 10 October 2026. How many virtual machines one card holds by MIG slice and compute profile is in our guide to MIG and vGPU per card; this article sizes a whole desktop estate.

User types, vGPU profiles and licences

NVIDIA’s guides sort users into light, medium and heavy and match them to three profile series. The vPC guide describes B-profiles as “Virtual desktops for business professionals and knowledge workers”, Q-profiles as virtual workstations for creative and technical professionals and A-profiles as app streaming for users of published applications. It adds that heavy users need a large profile and may be better kept on an RTX vWS licence.

USER TYPEPROFILELICENCEUSERS PER CARD
Knowledge workerB-series, 2 GB (2B)vPC, per concurrent user12 to 48, by card
Knowledge worker, heavy useB-series, 3 GB (3B)vPC, per concurrent user8 on the L4, 16 on the L40S
Light workstation userDC-4Q (NVIDIA’s example: RTX PRO 4500 SE)RTX vWS, per concurrent user6 to 8
Medium workstation userDC-4Q or DC-8QRTX vWS, per concurrent user3 to 6
Heavy CAD and 3D userDC-12Q to DC-24Q on RTX PRO 6000 SERTX vWS, per concurrent user2 to 4

NVIDIA vPC sizing guide (density tables) and RTX vWS sizing guide (example VDI configurations), both updated 19 August 2026; vGPU user guide, virtual GPU types for the L4 and L40S; NVIDIA vGPU packaging and licensing guide, updated 24 March 2026. SE stands for Server Edition.

The workstation rows are ranges because NVIDIA counts users per GPU there, not profiles per GPU. An RTX PRO 6000 Server Edition has room for twelve DC-8Q profiles by memory, but NVIDIA’s example configuration puts 3 to 6 medium users on it, since they share its compute. The B-series rows are the frame buffer ceiling, which the next section explains.

Frame buffer per desktop: 1, 2 or 3 GB

The frame buffer is the GPU memory a desktop gets, fixed by its profile. In NVIDIA’s frame buffer tables for its knowledge worker workload, a 1 GB profile on one or two HD monitors peaked at 99 per cent on the L4 and at 97 to 98 per cent on the L40 and L40S. A 2 GB profile, tested at QHD on two or four monitors and at 4K on one or two, peaked at 96 to 98 per cent on both. NVIDIA’s vPC guide states that modern operating systems and evolving application workloads “require larger frame buffer sizes” and that “larger profile sizes (2B or 3B) are often needed to maintain a consistent user experience”. Blackwell cards drop the smallest size, since “1 GB profiles are not available on Blackwell GPUs”, and vPC on Blackwell offers 2 GB and 3 GB profiles.

This article therefore plans 2 GB per knowledge worker and moves a group to 3 GB when its pilot shows the frame buffer running near its limit. Users with several high-resolution screens, video calls inside the desktop or heavy web applications are the groups to test first.

Without MIG, all vGPUs on one physical GPU must have the same frame buffer size, unless the GPU runs in mixed-size mode, which vSphere supports from ESXi 8.0 Update 3 with vGPU 17.2 and which lowers the count: the L40S holds six L40S-8Q profiles of one size but four in mixed mode. Plan each user group on its own cards, with one profile size per card.

Which GPUs NVIDIA lists for virtual desktops

NVIDIA’s vPC guide names the RTX PRO 4500 Blackwell Server Edition “the primary recommended GPU for NVIDIA vPC”, with the L4 and the A16 beside it. Its maximum density table for a 2U server, which assumes 1 or 2 GB per user and “reflects maximum supported density rather than a recommended deployment point”, also lists the RTX PRO 6000 Server Edition and the L40S. The RTX vWS guide recommends the RTX PRO 6000 Server Edition for high-end 3D visualisation and for vPC deployments that need high user density.

Of the cards these guides name, we supply the RTX PRO 4500 Server Edition, the L4, the L40S and the RTX PRO 6000 Server Edition; the A16 is not in our range. The RTX PRO 4500 workstation card is a different product, which NVIDIA’s vSphere support matrix does not list.

CARDMEMORY, POWERNVENC ENGINESVPC USERS PER CARDCARDS PER 2U
NVIDIA L424 GB, 72 W, low profile212 at 2 GB, 8 at 3 GB16
NVIDIA L40S48 GB, 350 W324 at 2 GB, 16 at 3 GB8
RTX PRO 6000 Server Edition96 GB, up to 600 W432 at 2 GB, 48 with MIG-backed vGPU4

NVIDIA product pages for the L4, L40S and RTX PRO 6000 Blackwell Server Edition, read 10 October 2026; users per card from the vGPU user guide (L4-2B, L4-3B, L40S-2B, L40S-3B) and the vPC sizing guide; cards per 2U from the vPC guide’s maximum density table.

The remote display protocol encodes every desktop’s screen, and NVIDIA’s vWS guide lists the encoder as a metric to watch: “The Video Encoder Usage metric specifically measures how intensively the GPU’s encoder is utilized”. We found no session limit per encoder in these guides. At full density, by our arithmetic, six desktops share each encoder on the L4, eight on the L40S and eight to twelve on the RTX PRO 6000, so encoder load belongs in the pilot as much as memory. The L40S and the RTX PRO 6000 also differ in MIG, power and server fit, as our comparison of the L40S and RTX PRO 6000 Server Edition sets out.

Worked example: 1,000 concurrent desktops

The example estate has 1,000 concurrent sessions at peak: 900 knowledge workers on vPC, 80 medium and 20 heavy workstation users on RTX vWS. Cards are users divided by users per card, rounded up. Hosts are cards divided by cards per host, rounded up, plus one host. Cards per host follow NVIDIA’s 2U table; the server model you choose sets the actual figure through its slots, power and airflow.

USERS, CARD, PROFILEUSERS PER CARDCARDS NEEDEDCARDS PER HOSTHOSTS WITH N+1
900 vPC, L4-2B1275165 + 1 = 6
900 vPC, L40S-2B243885 + 1 = 6
900 vPC, L40S-3B165788 + 1 = 9
900 vPC, 4500 SE DC-2B165788 + 1 = 9
900 vPC, RTX PRO 6000 DC-2B322948 + 1 = 9
Same, MIG-backed, VCF 9.1481945 + 1 = 6
80 vWS medium, DC-8Q3 to 6, 4 used204pooled below
20 vWS heavy, DC-16Q2 to 4, 3 used7427 cards: 7 + 1 = 8

Our estimate from the users per card in NVIDIA’s vPC and RTX vWS sizing guides and the vGPU user guide; RTX PRO 6000 means the Server Edition, and 4500 SE the RTX PRO 4500 Server Edition, with 16 users and 8 cards per 2U from the vPC guide. A pilot replaces the assumed 4 and 3 workstation users per card with measured values.

For the 900 office desktops, three layouts reach 192 desktops per host: sixteen L4, eight L40S, or four RTX PRO 6000 with MIG-backed vGPU. Their card power differs, about 1.2 kW, 2.8 kW and 2.4 kW per host. The workstation pool is separate, because those users need RTX vWS licences and larger profiles, and its 27 cards fit seven hosts of four plus a spare.

The spare host has to carry the same card as its pool, because a vGPU desktop moves or restarts only on a host with the same GPU type and room for its profile. Two pools on two card types need two spare hosts, and the counts above already include them.

We supply the RTX PRO 4500 Server Edition, L4, L40S and RTX PRO 6000 Server Edition for VDI hosts, sized by seats per card. Describe your user groups and peak sessions for a configuration and quote within one business day.

Hypervisor, scheduler and licence limits

On vSphere, NVIDIA’s vGPU 20 support matrix lists VCF 9.1, VCF 9.0 and ESXi 8.0 from Update 3 P06, and the RTX PRO 6000 Server Edition needs VCF 9.0.1 or later on VCF 9.0, or ESXi 8 Update 3g or later on 8.0. Time-sliced, MIG-backed vGPUs, the mode behind 48 desktops per card, appear in the release 20.0 notes among the features supported on VCF 9.1, so on VCF 9.0 or ESXi 8 plan 32. Host preparation and vMotion of vGPU desktops are covered in our article on GPUs in vSphere, passthrough or vGPU. NVIDIA supports Nutanix AHV “as a generic Linux with KVM hypervisor”, and its AHV support page does not mention MIG-backed vGPU, so on AHV plan 32 desktops per RTX PRO 6000 until NVIDIA and Nutanix confirm more. GPU nodes and live migration on AHV are in our guide to a Nutanix GPU cluster.

The scheduler decides how desktops share a card’s compute. NVIDIA’s vGPU user guide names best effort as the default, and its vWS sizing guide notes that best effort often results in two to three times oversubscription of the GPU’s compute engine, which suits office users who are rarely busy at the same moment. The equal share and fixed share schedulers “impose a limit on GPU processing cycles used by a vGPU”, and under fixed share each vGPU’s share stays constant as others start or stop. For CAD and CAE, the same guide notes that single-threaded applications benefit from CPU clock speeds “generally above 3 GHz”, so the host processor is part of the workstation pool’s sizing.

NVIDIA’s licensing guide counts vPC and RTX vWS as concurrent user (CCU) licences: “A CCU license is required for every user who is accessing or using the software at any given time”, so the count follows the peak of concurrent sessions, not named users. vPC covers up to four displays and a maximum resolution of 5120 × 2880 (5K) and has no CUDA or OpenCL entitlement; RTX vWS covers up to four displays and up to 7680 × 4320 (8K) and includes CUDA and OpenCL. Both come as an annual subscription or as a perpetual licence bought with five years of support. Compute profiles for AI workloads are licensed per GPU through NVIDIA AI Enterprise instead, as our AI Enterprise licensing guide explains.

We supply RTX vWS, vPC and vApps licences, subscription or perpetual, on the same invoice as the hardware. Send us your peak concurrent users per user type through the form below, and we recommend the licence type and term.

Testing the sizing in a pilot

NVIDIA recommends a new proof of concept on the target GPU and software whenever the environment changes. Its own scalability tests ran the nVector knowledge worker workload, which “is designed to simulate peak usage scenarios”, on 64 and 128 VMs with two HD monitors each.

  1. Build one host with the target card, hypervisor version and vGPU release, and the production desktop image.
  2. Run the users’ own applications, or nVector, at the planned users per card for each group.
  3. Log frame buffer, GPU, encoder and decoder use with nvidia-smi, and CPU per VM with the hypervisor’s tools.
  4. Check each VM’s frame buffer: the vWS guide says it “should not frequently exceed 90% or average over 70%”, and a group above that moves to the next profile.
  5. Measure the user’s side: latency, remoted frames and image quality, where NVIDIA treats an SSIM score above 0.98 as good.
  6. Recalculate the cards and hosts with the measured users per card before the order.

What we supply

We supply the cards for VDI hosts on our GPU page: the RTX PRO 4500 Server Edition, the L4, the L40S and the RTX PRO 6000 Server Edition, with manufacturer warranty on one EU contract and invoice. We also build them into GPU nodes for virtual workstations, sized by seats per card, and we check the rack, power and airflow before we quote. The vPC and RTX vWS licences go on the same quote through our software and licensing offer.

FAQ

How many users per GPU for VDI?
NVIDIA’s maximum for office desktops at 2 GB per user is 12 on the L4, 24 on the L40S and 32 on the RTX PRO 6000 Server Edition, or 48 with MIG-backed, time-sliced vGPU, which on vSphere needs VCF 9.1. For CAD users on RTX vWS, NVIDIA’s sizing guide gives 3 to 6 medium users or 2 to 4 heavy users per GPU. A pilot with your own applications sets the final figure.
Which vGPU profile should office users get?
A B-series profile under an NVIDIA vPC licence, with 2 GB of frame buffer as the starting point. NVIDIA’s vPC guide says modern operating systems need larger frame buffers and that 2B or 3B profiles are often needed, and its tables show 1 GB profiles peaking at 97 to 99 per cent on the L4 and L40S. Blackwell cards offer no 1 GB profile at all.
Is the L40S suitable for VDI?
Yes. NVIDIA lists it in its vPC and RTX vWS sizing guides, its vGPU user guide allows 24 desktops at 2 GB or 16 at 3 GB per card, and the vPC guide’s 2U maximum-density example holds eight cards. It has three video encoders and no MIG, so desktops share it by time slicing.
How is NVIDIA vPC licensed?
Per concurrent user: NVIDIA’s licensing guide requires a CCU licence for every user accessing the software at any given time, so the count follows the peak of concurrent sessions, not named users. vPC covers up to four displays and a maximum resolution of 5K, without CUDA, and comes as an annual subscription or a perpetual licence with five years of support.
How many GPU hosts do 1,000 virtual desktops need?
For 900 office desktops at 2 GB and 100 CAD sessions, our estimate from NVIDIA’s figures is six hosts for the office desktops, with sixteen L4 or eight L40S per host, and eight hosts with four RTX PRO 6000 each for the workstation users, each pool including one spare host. The server model’s slots, power and airflow and a pilot can change these counts.
Which GPU does NVIDIA recommend for vPC?
NVIDIA’s vPC sizing guide names the RTX PRO 4500 Blackwell Server Edition as the primary recommended GPU for vPC, with the L4 and A16 as further options. Its maximum density table for 2U servers also lists the RTX PRO 6000 Server Edition and the L40S, and its RTX vWS guide recommends the RTX PRO 6000 for high-density vPC. Of these, we supply the RTX PRO 4500 Server Edition, the L4, the L40S and the RTX PRO 6000 Server Edition.

Send us your user groups with their peak concurrent sessions, the applications and screens per user, your hypervisor and its version, and the vGPU licences you hold. We reply within one business day with the cards, profiles, hosts and licence counts that fit, and a configuration and quote checked against your rack, power and airflow.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna