VDI GPU sizing: how many cards and vGPU hosts 200 to 1,000 virtual desktops need
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Size a VDI estate as concurrent users per type × vGPU profile → cards → hosts, plus one spare host per GPU type for failover (N+1); a pilot on your own desktop image confirms the users per card
- Knowledge workers run on NVIDIA vPC with B-series profiles: NVIDIA’s vPC guide says modern operating systems need larger frame buffers and that 2B or 3B profiles are often needed, and Blackwell cards offer no 1 GB profile
- NVIDIA’s maximum per card at 2 GB is 12 users on the L4, 24 on the L40S and 32 on the RTX PRO 6000 Server Edition, or 48 with MIG-backed, time-sliced vGPU, which vGPU 20 supports on vSphere from VCF 9.1
- For CAD and 3D, NVIDIA’s RTX vWS guide gives 3 to 6 medium users per GPU on DC-4Q or DC-8Q and 2 to 4 heavy users on the RTX PRO 6000 with DC-12Q to DC-24Q
- vPC and RTX vWS are licensed per concurrent user (CCU), so 900 office desktops and 100 workstation sessions at peak need 900 vPC and 100 RTX vWS licences
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
How to size GPU hosts for VDI
For VDI GPU sizing, count the concurrent users of each type, give each type a vGPU profile with a fixed frame buffer, and divide each user type by the number of users one card holds at that profile. Divide the cards by the cards per host and add one host for failover (N+1). Knowledge workers run on NVIDIA Virtual PC (vPC) with B-series profiles of 2 or 3 GB, and users of CAD and 3D applications run on NVIDIA RTX Virtual Workstation (RTX vWS) with Q-series profiles. Both licences are counted per concurrent user.
The result is an estimate, and both of NVIDIA’s sizing guides tell you to confirm it with a proof of concept on the target hardware. The figures below come from NVIDIA’s vPC and RTX vWS sizing guides, both updated on 19 August 2026, from the vGPU 20 documentation updated on 29 September 2026 and from NVIDIA’s product pages, read on 10 October 2026. How many virtual machines one card holds by MIG slice and compute profile is in our guide to MIG and vGPU per card; this article sizes a whole desktop estate.
User types, vGPU profiles and licences
NVIDIA’s guides sort users into light, medium and heavy and match them to three profile series. The vPC guide describes B-profiles as “Virtual desktops for business professionals and knowledge workers”, Q-profiles as virtual workstations for creative and technical professionals and A-profiles as app streaming for users of published applications. It adds that heavy users need a large profile and may be better kept on an RTX vWS licence.
| USER TYPE | PROFILE | LICENCE | USERS PER CARD |
|---|---|---|---|
| Knowledge worker | B-series, 2 GB (2B) | vPC, per concurrent user | 12 to 48, by card |
| Knowledge worker, heavy use | B-series, 3 GB (3B) | vPC, per concurrent user | 8 on the L4, 16 on the L40S |
| Light workstation user | DC-4Q (NVIDIA’s example: RTX PRO 4500 SE) | RTX vWS, per concurrent user | 6 to 8 |
| Medium workstation user | DC-4Q or DC-8Q | RTX vWS, per concurrent user | 3 to 6 |
| Heavy CAD and 3D user | DC-12Q to DC-24Q on RTX PRO 6000 SE | RTX vWS, per concurrent user | 2 to 4 |
NVIDIA vPC sizing guide (density tables) and RTX vWS sizing guide (example VDI configurations), both updated 19 August 2026; vGPU user guide, virtual GPU types for the L4 and L40S; NVIDIA vGPU packaging and licensing guide, updated 24 March 2026. SE stands for Server Edition.
The workstation rows are ranges because NVIDIA counts users per GPU there, not profiles per GPU. An RTX PRO 6000 Server Edition has room for twelve DC-8Q profiles by memory, but NVIDIA’s example configuration puts 3 to 6 medium users on it, since they share its compute. The B-series rows are the frame buffer ceiling, which the next section explains.
Frame buffer per desktop: 1, 2 or 3 GB
The frame buffer is the GPU memory a desktop gets, fixed by its profile. In NVIDIA’s frame buffer tables for its knowledge worker workload, a 1 GB profile on one or two HD monitors peaked at 99 per cent on the L4 and at 97 to 98 per cent on the L40 and L40S. A 2 GB profile, tested at QHD on two or four monitors and at 4K on one or two, peaked at 96 to 98 per cent on both. NVIDIA’s vPC guide states that modern operating systems and evolving application workloads “require larger frame buffer sizes” and that “larger profile sizes (2B or 3B) are often needed to maintain a consistent user experience”. Blackwell cards drop the smallest size, since “1 GB profiles are not available on Blackwell GPUs”, and vPC on Blackwell offers 2 GB and 3 GB profiles.
This article therefore plans 2 GB per knowledge worker and moves a group to 3 GB when its pilot shows the frame buffer running near its limit. Users with several high-resolution screens, video calls inside the desktop or heavy web applications are the groups to test first.
Without MIG, all vGPUs on one physical GPU must have the same frame buffer size, unless the GPU runs in mixed-size mode, which vSphere supports from ESXi 8.0 Update 3 with vGPU 17.2 and which lowers the count: the L40S holds six L40S-8Q profiles of one size but four in mixed mode. Plan each user group on its own cards, with one profile size per card.
Which GPUs NVIDIA lists for virtual desktops
NVIDIA’s vPC guide names the RTX PRO 4500 Blackwell Server Edition “the primary recommended GPU for NVIDIA vPC”, with the L4 and the A16 beside it. Its maximum density table for a 2U server, which assumes 1 or 2 GB per user and “reflects maximum supported density rather than a recommended deployment point”, also lists the RTX PRO 6000 Server Edition and the L40S. The RTX vWS guide recommends the RTX PRO 6000 Server Edition for high-end 3D visualisation and for vPC deployments that need high user density.
Of the cards these guides name, we supply the RTX PRO 4500 Server Edition, the L4, the L40S and the RTX PRO 6000 Server Edition; the A16 is not in our range. The RTX PRO 4500 workstation card is a different product, which NVIDIA’s vSphere support matrix does not list.
| CARD | MEMORY, POWER | NVENC ENGINES | VPC USERS PER CARD | CARDS PER 2U |
|---|---|---|---|---|
| NVIDIA L4 | 24 GB, 72 W, low profile | 2 | 12 at 2 GB, 8 at 3 GB | 16 |
| NVIDIA L40S | 48 GB, 350 W | 3 | 24 at 2 GB, 16 at 3 GB | 8 |
| RTX PRO 6000 Server Edition | 96 GB, up to 600 W | 4 | 32 at 2 GB, 48 with MIG-backed vGPU | 4 |
NVIDIA product pages for the L4, L40S and RTX PRO 6000 Blackwell Server Edition, read 10 October 2026; users per card from the vGPU user guide (L4-2B, L4-3B, L40S-2B, L40S-3B) and the vPC sizing guide; cards per 2U from the vPC guide’s maximum density table.
The remote display protocol encodes every desktop’s screen, and NVIDIA’s vWS guide lists the encoder as a metric to watch: “The Video Encoder Usage metric specifically measures how intensively the GPU’s encoder is utilized”. We found no session limit per encoder in these guides. At full density, by our arithmetic, six desktops share each encoder on the L4, eight on the L40S and eight to twelve on the RTX PRO 6000, so encoder load belongs in the pilot as much as memory. The L40S and the RTX PRO 6000 also differ in MIG, power and server fit, as our comparison of the L40S and RTX PRO 6000 Server Edition sets out.
Worked example: 1,000 concurrent desktops
The example estate has 1,000 concurrent sessions at peak: 900 knowledge workers on vPC, 80 medium and 20 heavy workstation users on RTX vWS. Cards are users divided by users per card, rounded up. Hosts are cards divided by cards per host, rounded up, plus one host. Cards per host follow NVIDIA’s 2U table; the server model you choose sets the actual figure through its slots, power and airflow.
| USERS, CARD, PROFILE | USERS PER CARD | CARDS NEEDED | CARDS PER HOST | HOSTS WITH N+1 |
|---|---|---|---|---|
| 900 vPC, L4-2B | 12 | 75 | 16 | 5 + 1 = 6 |
| 900 vPC, L40S-2B | 24 | 38 | 8 | 5 + 1 = 6 |
| 900 vPC, L40S-3B | 16 | 57 | 8 | 8 + 1 = 9 |
| 900 vPC, 4500 SE DC-2B | 16 | 57 | 8 | 8 + 1 = 9 |
| 900 vPC, RTX PRO 6000 DC-2B | 32 | 29 | 4 | 8 + 1 = 9 |
| Same, MIG-backed, VCF 9.1 | 48 | 19 | 4 | 5 + 1 = 6 |
| 80 vWS medium, DC-8Q | 3 to 6, 4 used | 20 | 4 | pooled below |
| 20 vWS heavy, DC-16Q | 2 to 4, 3 used | 7 | 4 | 27 cards: 7 + 1 = 8 |
Our estimate from the users per card in NVIDIA’s vPC and RTX vWS sizing guides and the vGPU user guide; RTX PRO 6000 means the Server Edition, and 4500 SE the RTX PRO 4500 Server Edition, with 16 users and 8 cards per 2U from the vPC guide. A pilot replaces the assumed 4 and 3 workstation users per card with measured values.
For the 900 office desktops, three layouts reach 192 desktops per host: sixteen L4, eight L40S, or four RTX PRO 6000 with MIG-backed vGPU. Their card power differs, about 1.2 kW, 2.8 kW and 2.4 kW per host. The workstation pool is separate, because those users need RTX vWS licences and larger profiles, and its 27 cards fit seven hosts of four plus a spare.
The spare host has to carry the same card as its pool, because a vGPU desktop moves or restarts only on a host with the same GPU type and room for its profile. Two pools on two card types need two spare hosts, and the counts above already include them.
We supply the RTX PRO 4500 Server Edition, L4, L40S and RTX PRO 6000 Server Edition for VDI hosts, sized by seats per card. Describe your user groups and peak sessions for a configuration and quote within one business day.
Hypervisor, scheduler and licence limits
On vSphere, NVIDIA’s vGPU 20 support matrix lists VCF 9.1, VCF 9.0 and ESXi 8.0 from Update 3 P06, and the RTX PRO 6000 Server Edition needs VCF 9.0.1 or later on VCF 9.0, or ESXi 8 Update 3g or later on 8.0. Time-sliced, MIG-backed vGPUs, the mode behind 48 desktops per card, appear in the release 20.0 notes among the features supported on VCF 9.1, so on VCF 9.0 or ESXi 8 plan 32. Host preparation and vMotion of vGPU desktops are covered in our article on GPUs in vSphere, passthrough or vGPU. NVIDIA supports Nutanix AHV “as a generic Linux with KVM hypervisor”, and its AHV support page does not mention MIG-backed vGPU, so on AHV plan 32 desktops per RTX PRO 6000 until NVIDIA and Nutanix confirm more. GPU nodes and live migration on AHV are in our guide to a Nutanix GPU cluster.
The scheduler decides how desktops share a card’s compute. NVIDIA’s vGPU user guide names best effort as the default, and its vWS sizing guide notes that best effort often results in two to three times oversubscription of the GPU’s compute engine, which suits office users who are rarely busy at the same moment. The equal share and fixed share schedulers “impose a limit on GPU processing cycles used by a vGPU”, and under fixed share each vGPU’s share stays constant as others start or stop. For CAD and CAE, the same guide notes that single-threaded applications benefit from CPU clock speeds “generally above 3 GHz”, so the host processor is part of the workstation pool’s sizing.
NVIDIA’s licensing guide counts vPC and RTX vWS as concurrent user (CCU) licences: “A CCU license is required for every user who is accessing or using the software at any given time”, so the count follows the peak of concurrent sessions, not named users. vPC covers up to four displays and a maximum resolution of 5120 × 2880 (5K) and has no CUDA or OpenCL entitlement; RTX vWS covers up to four displays and up to 7680 × 4320 (8K) and includes CUDA and OpenCL. Both come as an annual subscription or as a perpetual licence bought with five years of support. Compute profiles for AI workloads are licensed per GPU through NVIDIA AI Enterprise instead, as our AI Enterprise licensing guide explains.
We supply RTX vWS, vPC and vApps licences, subscription or perpetual, on the same invoice as the hardware. Send us your peak concurrent users per user type through the form below, and we recommend the licence type and term.
Testing the sizing in a pilot
NVIDIA recommends a new proof of concept on the target GPU and software whenever the environment changes. Its own scalability tests ran the nVector knowledge worker workload, which “is designed to simulate peak usage scenarios”, on 64 and 128 VMs with two HD monitors each.
- Build one host with the target card, hypervisor version and vGPU release, and the production desktop image.
- Run the users’ own applications, or nVector, at the planned users per card for each group.
- Log frame buffer, GPU, encoder and decoder use with nvidia-smi, and CPU per VM with the hypervisor’s tools.
- Check each VM’s frame buffer: the vWS guide says it “should not frequently exceed 90% or average over 70%”, and a group above that moves to the next profile.
- Measure the user’s side: latency, remoted frames and image quality, where NVIDIA treats an SSIM score above 0.98 as good.
- Recalculate the cards and hosts with the measured users per card before the order.
What we supply
We supply the cards for VDI hosts on our GPU page: the RTX PRO 4500 Server Edition, the L4, the L40S and the RTX PRO 6000 Server Edition, with manufacturer warranty on one EU contract and invoice. We also build them into GPU nodes for virtual workstations, sized by seats per card, and we check the rack, power and airflow before we quote. The vPC and RTX vWS licences go on the same quote through our software and licensing offer.
FAQ
How many users per GPU for VDI?
Which vGPU profile should office users get?
Is the L40S suitable for VDI?
How is NVIDIA vPC licensed?
How many GPU hosts do 1,000 virtual desktops need?
Which GPU does NVIDIA recommend for vPC?
Send us your user groups with their peak concurrent sessions, the applications and screens per user, your hypervisor and its version, and the vGPU licences you hold. We reply within one business day with the cards, profiles, hosts and licence counts that fit, and a configuration and quote checked against your rack, power and airflow.
Talk to an expertWe reply within one business day