BLOG · GUIDE ·

VMware Private AI Foundation with NVIDIA: what it is, what it needs and when it fits

IN BRIEF
  • VMware Private AI Foundation with NVIDIA 9.1 shipped on 12 May 2026 and 9.1.1 on 3 September 2026: deep learning VMs, GPU Kubernetes clusters on VKS and Private AI Services, all on VCF 9.1
  • Broadcom’s requirements list three licences: a VCF subscription, the Private AI Foundation licence counted in cores, and NVIDIA AI Enterprise for vGPU; since 3 November 2025 a VCF subscription may already include the second
  • With vGPU, NVIDIA AI Enterprise is required for the host driver on ESX and the guest drivers, counted per physical GPU; since 9.1 Broadcom documents DirectPath GPUs without it, though NVIDIA still requires a licence per GPU when AI Enterprise software such as NIM runs on them
  • The initial cluster of the GPU workload domain needs at least 3 GPU-enabled ESX hosts, and Broadcom sends you to its Compatibility Guide for AI/ML compute to check each GPU
  • Private AI Services 3.0, released on 3 September 2026, runs models with vLLM 0.20.0 on GPU, Infinity for embeddings and llama.cpp on CPU, and supports PostgreSQL 16.8 with pgvector 0.8.0 for RAG vectors

What it is, in September 2026

Broadcom describes VMware Private AI Foundation with NVIDIA as “a multi-component solution” that “builds on virtual infrastructure management and cloud management from VCF to run generative AI workloads on systems with NVIDIA GPU devices”. It reached initial availability in March 2024 as an add-on to VMware Cloud Foundation, and its version numbers follow VCF’s: 9.0 shipped on 17 June 2025, 9.1 on 12 May 2026 and 9.1.1 on 3 September 2026. This guide follows the 9.1 documentation, last updated by Broadcom on 21 September 2026. It does not cover VMware AI Factory, which Broadcom announced on 31 August 2026 as “the software-defined foundation of VMware Private AI Cloud”.

The documentation names two paths: deep learning VMs that administrators provision for data scientists, and Kubernetes clusters from the vSphere Kubernetes Service (VKS) for production workloads. Private AI Services adds a third layer, in which models, retrieval and agents run as a shared service.

Which parts come from Broadcom and which from NVIDIA

Each side brings its own licence and its own support path. Broadcom’s 9.1 documentation lists these components.

COMPONENTFROMWHAT IT DOES
vCenter, ESX, NSXBroadcom, in VCFthe platform; NSX Edge or Virtual Network Appliance nodes do the Supervisor’s north-south routing
Supervisor and VKSBroadcom, in VCFVMs and Kubernetes clusters on request, with vGPU or passthrough GPUs
VCF AutomationBroadcom, in VCFself-service catalogue items such as the AI workstation and the AI Kubernetes Cluster
VCF OperationsBroadcom, in VCFGPU metrics per cluster, host and GPU, with Private AI (GPU) dashboards
Deep learning VM imageBroadcomUbuntu with Docker, the NVIDIA Container Toolkit and conda; PyTorch, Triton or DCGM Exporter from NGC can be pre-installed
Private AI ServicesBroadcoma Supervisor Service: model gallery in Harbor, model endpoints, knowledge bases, agents, MCP servers
Data Services ManagerBroadcomPostgreSQL with pgvector, the vector database
vGPU host and guest driversNVIDIA AI Enterprisea VIB on each ESX host and the matching driver in each VM or node
NVIDIA License SystemNVIDIAvGPU licences from the Licensing Portal or a Delegated License Service appliance
GPU and Network OperatorsNVIDIAGPU software inside VKS clusters; RDMA and GPUDirect networking
NVIDIA NGCNVIDIAGPU-optimised containers, NIM microservices on VKS among them

Broadcom TechDocs, VMware Private AI Foundation with NVIDIA 9.1 and VCF 9.1, pages updated 15 July to 23 September 2026. A passthrough GPU runs NVIDIA’s data-centre driver instead of the vGPU drivers.

Three licence lines

Broadcom’s 9.1 requirements list three licences; a cluster that uses vGPU needs all three.

VCF subscription. VCF is licensed per physical core, with at least 16 cores per processor, as our renewal guide sets out. The requirements name a VCF subscription; we found no mention of vSphere Foundation or the standalone vSphere editions.

Private AI Foundation licence. Broadcom’s VCF 9.1 licensing overview calls it “VMware Private AI Foundation with NVIDIA (cores)”. Broadcom’s requirements assign it to the GPU-enabled workload domain; assigning it also to the management domain’s vCenter, as an add-on, activates the guided deployment UI in the vSphere Client and the quickstart wizard in VCF Automation, and “the license capacity is allocated only to the GPU-enabled workload domains and not to the management domain”. What the VCF licence covers on its own, Broadcom puts this way: “you can deploy AI workloads with and without a Supervisor enabled and use the GPU metrics in vCenter and VCF Operations under the VCF license”; the table lists what the Private AI Foundation licence adds. It may already be on the account: since 3 November 2025, “based on your VCF subscription, you might also receive a license” for it, and the VCF 9.1 FAQ of 3 September 2026 lists VCF Private AI Services as a VCF component. Check the entitlement before ordering.

NVIDIA AI Enterprise. Broadcom requires it “for running AI workloads on vGPU”, for “the host driver VIB file on ESX hosts and the guest OS drivers”, and NVIDIA licenses vGPU for Compute “only through NVIDIA AI Enterprise”, per physical GPU. Our licensing guide covers the licence server, the included subscriptions and NVIDIA’s support condition, an NVIDIA-Certified System. Broadcom’s 9.1 release notes list DirectPath enablement for GPUs as new: VMs and Kubernetes nodes get “exclusive access to GPU resources without the need for an NVIDIA AI Enterprise (NVAIE) license”. That covers the GPU and its driver; NVIDIA still requires a licence for every GPU in a server that hosts AI Enterprise software, such as NIM microservices in production.

LICENCECOUNTED INUNLOCKSNEEDED FOR
VCF subscriptionphysical cores, at least 16 per processorvSphere, vSAN, NSX, VKS, VCF Operations, VCF Automation; GPU metricsevery VCF host
Private AI Foundationcores, allocated to GPU-enabled workload domainscatalogue items, deep learning VM image, guided deployment, pgvector in Data Services Manager, Private AI Servicesthese functions; may come with VCF since 3 November 2025
NVIDIA AI Enterprisephysical GPUs, every GPU in a server that runs itvGPU host and guest drivers; AI containers from NGC, per BroadcomvGPU, and AI Enterprise software such as NIM in production; not for DirectPath GPU access in 9.1

Broadcom TechDocs: Private AI Foundation 9.1 requirements, licence assignment and release notes, VCF 9.1 licensing pages, September 2026; NVIDIA AI Enterprise documentation, 2 September 2026.

Matching the two Broadcom lines to what the estate will run is the licence part of our VMware optimisation service, in which our engineering partner Vixen.UNO audits the estate and its licences and matches editions and subscriptions to the real workloads, within Broadcom’s current licensing logic; the NVIDIA AI Enterprise count goes into the quote with the hardware, as our NVIDIA AI Enterprise page explains.

GPUs and servers the documents name

Broadcom sets the host count: “at least 3 GPU-enabled ESX hosts to include in the initial cluster of a workload domain”. Its 9.1 requirements list no GPU models; they ask you to “verify in the Broadcom Compatibility Guide for GPUs and Accelerators for AI/ML Compute that the GPUs on your ESX hosts are supported”, an online search tool. vGPU hosts need SR-IOV enabled in the BIOS and the NVIDIA vGPU host driver. For Advanced NIC Passthrough the requirements name the NVIDIA ConnectX-6 Dx, for Enhanced DirectPath I/O only, the ConnectX-7 and the BlueField-3 in NIC mode. The 9.1 release notes also add NVIDIA HGX platforms with Blackwell GPUs and NVSwitch; on the HGX B200 and B300 with vSphere, NVIDIA AI Enterprise supports full-GPU and multi-vGPU VMs only, not fractional vGPUs.

NVIDIA’s vGPU documentation for vSphere, releases 20.0 to 20.2, names host driver packages for VCF 9.1, VCF 9.0 and vSphere 8.0, and its GPU list includes the L4, L40S and RTX PRO 6000 Blackwell Server Edition. NVIDIA AI Enterprise 8.2, whose support matrix covers VMware ESXi “8.0 and later, 9.0 and later”, also lists the H200 NVL, with compute vGPU profiles.

CARDMEMORY AND MIGNVIDIA ON VSPHEREAI ENTERPRISE
RTX PRO 6000 Server Edition96 GB, up to 4 MIG instanceson the vGPU list, from VCF 9.0.1 on the 9.0 line; in AI Enterprise 8.2none included, licensed per GPU
H200 NVL141 GB, up to 7 MIG instancesin AI Enterprise 8.2, with compute vGPU profilesfive-year subscription included
L40S48 GB, no MIGon the vGPU list; in AI Enterprise 8.2none included, licensed per GPU
L424 GB, no MIGon the vGPU list; in AI Enterprise 8.2none included, licensed per GPU
Workstation editionsRTX PRO 6000 Workstation and Max-Q 96 GB, 5000 48 or 72 GB, 4500 32 GB, 4000 24 GBnot on NVIDIA’s vGPU list for vSpherenot in the 8.2 support matrix

NVIDIA vGPU documentation for VMware vSphere, 22 September 2026; NVIDIA AI Enterprise 8.2 support matrix and vGPU types, 2 September 2026; MIG counts from NVIDIA’s MIG user guide. The H200 NVL subscription is activated with the GPU serial number.

We plan such clusters on the cards NVIDIA lists for vGPU on vSphere, such as the RTX PRO 6000 Server Edition, the L40S and the L4, or on the H200 NVL, which NVIDIA AI Enterprise 8.2 lists with compute vGPU profiles. The H200 NVL also changes the licence line: each card includes a five-year NVIDIA AI Enterprise subscription, activated with its serial number, and the five years start 90 days after the card ships to the server maker, not at activation, per NVIDIA’s licensing guide; the other three cards need one licence per GPU for vGPU.

What an administrator builds

The management domain runs vCenter, NSX Manager, VCF Operations, VCF Automation and Data Services Manager. The GPUs go into a workload domain where “the vSphere Supervisor option is mandatory”, with an NSX Edge or Virtual Network Appliance cluster for the Supervisor’s north-south routing. Each host’s GPUs are set to Shared Direct for vGPU, with time-sliced or MIG-backed profiles, or to Fixed or Dynamic DirectPath for passthrough, and the host is then restarted in maintenance mode. A passthrough VM gets the whole card, which Broadcom says “leads to improved performance”, but deep learning VMs with passthrough GPUs cannot be moved between hosts with vSphere vMotion. What each mode allows and blocks, from DRS to HA and snapshots, is in our guide to passthrough or vGPU.

Users reach the GPUs through VM classes, in which “as a VI administrator, you set the compute, networking, and GPU requirements according to the GPU configuration on the ESX hosts”. Each class carries a vGPU or passthrough profile, and reserved classes are allocated “to an organization in VCF Automation”, so GPU capacity becomes a catalogue the VI team controls. The Private AI Foundation Quickstart then adds the catalogue items, among them the AI workstation, which deploys a deep learning VM, and the AI Kubernetes Cluster, a VKS cluster with the NVIDIA GPU Operator. Broadcom notes that the item’s GPU Operator v25.3.1 “has reached End-of-Life”; knowledge base article 439984 describes moving the blueprint to v26.3.1.

Private AI Services is “installed as a Supervisor Service”, activated per VCF Automation namespace, and needs a Harbor registry for the model gallery; version 3.0 needs Data Services Manager, for the pgvector database, when knowledge bases and agents are used. Model Runtime serves models behind an OpenAI-compatible API with open-source engines: vLLM for models on GPU, Infinity for embeddings, llama.cpp on CPU. Version 3.0 of 3 September 2026, for VCF 9.1.x, lists vLLM 0.20.0, llama.cpp b9309 and Infinity 0.0.76 as its inference engines and supports PostgreSQL 16.8 with pgvector 0.8.0.

RAG and the vector database

Broadcom’s 9.1.x release notes list the VCF Automation catalogue items for NVIDIA RAG as removed, because “NVIDIA is no longer providing support for the blueprints that back these catalog items”. Retrieval now sits in Private AI Services: a knowledge base “consists of metadata, text chunks, and vector embeddings created for each processed document in the scope of a data source”, embedded by a Model Runtime endpoint and indexed in pgvector. Documents are uploaded as PDF, DOCX, PPTX, TXT, HTML, Markdown or CSV files, or read from linked sources such as Google Drive folders, Confluence spaces and S3-compatible stores, each with one set of credentials.

We found nothing in the 9.1 documentation about carrying each user’s document permissions into the answers. Where different people may see different documents, that separation has to be designed before the first knowledge base is filled; RAG assistants that respect each user’s access rights are part of our AI/ML integration work.

When it fits, and when a leaner stack is enough

It fits when the estate already runs VCF 9.1, or its modernisation plan leads there, and the Private AI Foundation licence is in the subscription; when several teams need GPU VMs and Kubernetes clusters as self-service, under the same operations as the rest of vSphere; when models should run as a shared service with endpoints, knowledge bases and agents, even at a disconnected site, since 9.1 supports installing Private AI Services in air-gapped environments; and when the budget covers at least three GPU hosts in the first cluster plus NVIDIA AI Enterprise per GPU wherever vGPU or AI Enterprise software is used.

A leaner stack is enough when one or two GPU servers serve one team and a few models. A VM with a vGPU or passthrough GPU on the existing vSphere cluster, running an open-source engine such as vLLM, covers that case, and our renewal guide shows which vSphere editions include vGPU. Kubernetes-first teams can run their own platform on bare-metal GPU nodes, as our Kubernetes platform guide describes. An estate on vSphere Foundation or a standalone edition starts with the VM path above, because the requirements name a VCF subscription.

Neither path is a verdict on the platform; the deciding questions are scale, how many teams share the GPUs, and who will operate the result.

What we supply

Eurokommerz supplies GPU servers built to order with RTX PRO 6000 Server Edition, H200 NVL, L40S or L4 cards, sized by engineers per workload, with manufacturer warranty and an EU contract and invoicing from Austria; confirm the chosen server and GPUs in Broadcom’s Compatibility Guide before they go into a workload domain. NVIDIA AI Enterprise and vGPU licences come on the same invoice as the AI servers, and each H200 NVL includes NVIDIA’s five-year AI Enterprise subscription. For the VMware side, our engineering partner Vixen.UNO delivers VMware optimisation: an audit of the estate and its licences, editions and subscriptions matched to the real workloads, and modernisation of vSphere, vSAN, NSX and VCF in agreed maintenance windows, with a rollback plan at every stage. For the AI layer, the same team delivers AI/ML integration: private LLMs with vLLM, Ollama or NVIDIA AI Enterprise, RAG assistants that respect each user’s access rights, and logging of queries and answers. The contract is with Eurokommerz, and the engineering is delivered by the Vixen.UNO team.

FAQ

What is VMware Private AI Foundation with NVIDIA?
A Broadcom solution on VMware Cloud Foundation, developed with NVIDIA, for AI workloads on ESX hosts with NVIDIA GPUs: deep learning VMs, GPU Kubernetes clusters on VKS and Private AI Services, with NVIDIA’s vGPU drivers, GPU Operator and NGC containers. Version 9.1 shipped on 12 May 2026 and 9.1.1 on 3 September 2026.
Is Private AI Foundation included in a VCF subscription?
It can be. Broadcom’s VCF 9.1 licensing documentation says that since 3 November 2025, based on your VCF subscription, you might also receive a licence for it, and the VCF 9.1 FAQ lists VCF Private AI Services as a component. The Private AI Foundation requirements still list it as an add-on licence, so check the entitlement before ordering.
Do I need NVIDIA AI Enterprise for Private AI Foundation?
For vGPU, yes: Broadcom requires it for the host driver on the ESX hosts and the guest OS drivers, and NVIDIA licenses vGPU for Compute only through AI Enterprise, per physical GPU. Broadcom’s 9.1 documentation lets VMs and Kubernetes nodes use GPUs through DirectPath without an AI Enterprise licence; NVIDIA still requires one for every GPU in a server that hosts AI Enterprise software, such as NIM in production.
How many GPU hosts does Private AI Foundation need?
Broadcom’s 9.1 requirements ask for at least three GPU-enabled ESX hosts in the initial cluster of the workload domain, and that domain needs a Supervisor. The management domain runs vCenter, NSX Manager, VCF Operations, VCF Automation and Data Services Manager.
Which NVIDIA GPUs does Private AI Foundation support?
Broadcom points to its Compatibility Guide for GPUs and Accelerators for AI/ML Compute rather than listing models. On NVIDIA’s side, the vGPU documentation for VCF 9.x lists, among others, the L4, L40S and RTX PRO 6000 Blackwell Server Edition, and NVIDIA AI Enterprise 8.2 also lists the H200 NVL; the RTX PRO workstation editions are neither on NVIDIA’s vGPU list for vSphere nor in the AI Enterprise 8.2 support matrix.
Which inference engines does Private AI Services use?
Model Runtime uses open-source engines behind an OpenAI-compatible API: vLLM for models on GPU, Infinity for embedding models on GPU or CPU, and llama.cpp for models on CPU. Private AI Services 3.0 of 3 September 2026 lists vLLM 0.20.0 among its inference engines and supports PostgreSQL 16.8 with pgvector 0.8.0; the pgvector database is provisioned through Data Services Manager.

Send us your VCF version, the host list and the GPUs you run or plan, and tell us which AI workloads should land on them. We will answer with the licence lines, a GPU server configuration and a first assessment call. We reply within one business day.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry – see our privacy policy.

request@eurokommerz.at  ·  +43 1 585 1405 50  ·  Jordangasse 7, 1010 Vienna