BLOG · GUIDE ·

H200 NVL MIG profiles: the seven instance sizes, valid combinations and MIG-backed vGPU

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • NVIDIA’s MIG user guide (11 September 2026) lists the H200 NVL with up to seven MIG instances and seven H200 profiles: 1g.18gb up to 7 times, 1g.18gb+me once, 1g.35gb up to 4, 2g.35gb up to 3, 3g.71gb up to 2, 4g.71gb and 7g.141gb once each
  • The card has eight memory slices and seven SM slices: 1g.18gb takes 1/8 of the memory and 1/7 of the SMs, 1g.35gb takes 1/4 of the memory with 1/7 of the SMs, and both 3g.71gb and 4g.71gb take half the memory
  • NVIDIA’s H200 page gives the H200 NVL “Up to 7 MIGs @16.5GB each”, and the static MIG Manager configuration in NVIDIA’s GPU Operator repository lists 2 × 1g.18gb, 1 × 2g.35gb and 1 × 3g.71gb as the balanced layout for the H200 NVL
  • NVIDIA AI Enterprise 8.2 lists seven MIG-backed vGPU types for the H200 NVL, from H200-1-18C (up to 7 per GPU) to H200-7-141C, all compute (C-series) types that require an NVIDIA AI Enterprise licence
  • MIG on the H200 NVL is for compute only, since NVIDIA’s guide lists graphics API support in MIG for the RTX PRO 6000 Blackwell only, and NCCL does not run in MIG instances, so multi-GPU jobs take whole cards

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

H200 NVL MIG profiles at a glance

The H200 NVL splits into up to seven MIG instances. NVIDIA’s MIG user guide, last updated on 11 September 2026, lists the H200 NVL among the supported GPUs with compute capability 9.0, 141 GB and a maximum of seven instances. Its table of H200 MIG profiles for the “H200 141GB product” names seven GPU instance profiles, from 1g.18gb, available up to seven times, to 7g.141gb, which takes the whole card. The table does not separate the NVL from the SXM board, while NVIDIA’s vGPU reference lists the same seven profiles for the H200 NVL by name.

PROFILEMEMORY, SM SHAREENGINESMAX COUNTMIG-BACKED VGPU
MIG 1g.18gb1/8 memory and L2, 1/7 SMs1 NVDEC, 1 JPEG, 1 copy engine7H200-1-18C
MIG 1g.18gb+me1/8 memory and L2, 1/7 SMs1 NVDEC, 1 JPEG, 1 OFA, 1 copy engine1H200-1-18CME
MIG 1g.35gb1/4 memory, 1/8 L2, 1/7 SMs1 NVDEC, 1 JPEG, 1 copy engine4H200-1-35C
MIG 2g.35gb2/8 memory and L2, 2/7 SMs2 NVDEC, 2 JPEG, 2 copy engines3H200-2-35C
MIG 3g.71gb4/8 memory and L2, 3/7 SMs3 NVDEC, 3 JPEG, 3 copy engines2H200-3-71C
MIG 4g.71gb4/8 memory and L2, 4/7 SMs4 NVDEC, 4 JPEG, 4 copy engines1H200-4-71C
MIG 7g.141gbfull memory and L2, 7/7 SMs7 NVDEC, 7 JPEG, 1 OFA, 8 copy engines1H200-7-141C

NVIDIA MIG user guide, “Supported MIG Profiles”, Table 11 “GPU Instance Profiles on H200” and “Supported GPUs”, updated 11 September 2026; vGPU types from NVIDIA AI Enterprise 8.2, “Hopper H200 vGPU Types”, Table 93, updated 2 September 2026.

The gigabytes in a profile name are nominal. NVIDIA’s H200 product page gives the H200 NVL “Up to 7 MIGs @16.5GB each” and the H200 SXM “Up to 7 MIGs @18GB each”, so plan a 1g.18gb instance on the NVL at 16.5 GB and read the exact figure from nvidia-smi mig -lgip on the card.

How the slices add up: memory, SMs and engines

NVIDIA’s guide builds every instance from GPU slices, each combining one memory slice, “the smallest fraction of the GPU’s memory, including the corresponding memory controllers and cache”, with one SM slice. The H200 has eight memory slices and seven SM slices, which is why the fractions in the table run in eighths and sevenths. A 1g.18gb instance holds one of each. The 1g.35gb profile pairs two memory slices with a single SM slice, so it holds as much memory as a 2g.35gb but half its compute. The 3g.71gb and 4g.71gb profiles both take half of the memory and differ only in SMs and engines.

Decoders and copy engines grow with the SM share, and the full card has eight copy engines. Each instance has its own video decoders (NVDEC), JPEG decoders and copy engines; the H200 table lists no video encoders (NVENC). The optical flow accelerator (OFA) appears only in 1g.18gb+me and in the full 7g.141gb, and the guide notes that a single 1g profile can include the media extensions, which is why the count is one. The guide also defines the variants -me and +me.all and the graphics variant +gfx, which it marks as “new in GB20X”, the chips of the RTX PRO Blackwell cards. None of the three appears in the H200 table.

Each instance has its own memory controllers, cache and SMs, so the isolation is in hardware and one team’s job cannot take memory or bandwidth from another. NVIDIA’s guide sets three limits that matter for the H200: “NCCL is currently not supported with MIG”, “CUDA IPC across GPU instances is not supported”, and “No graphics APIs are supported”, with an exception for the RTX PRO 6000 Blackwell only.

Which H200 MIG profiles can coexist

A layout is valid when its instances fit the eight memory slices and seven SM slices without overlapping. NVIDIA’s concepts page describes building valid combinations “such that no two profiles overlap vertically” in its placement diagram. It also notes that drivers before R510 did not allow a (4 memory, 4 compute) instance next to a (4 memory, 3 compute) instance, and that “This restriction no longer applies on newer drivers”, so 4g.71gb plus 3g.71gb is a valid split of one H200 NVL.

The uniform layouts follow from the maximum counts in the table: seven 1g.18gb, four 1g.35gb, three 2g.35gb, two 3g.71gb, or one full instance. For mixed layouts, the static MIG Manager ConfigMap in NVIDIA’s GPU Operator repository, read on 10 October 2026, defines the “all-balanced” layout for the H200 141GB and the H200 NVL as two 1g.18gb, one 2g.35gb and one 3g.71gb. The guide’s placement example on an earlier card shows that a 3g instance starts only at slice 0 or 4, and smaller instances fill what the larger ones leave, so create the large instances first. Before you fix a mixed layout, list what the card accepts.

  1. Enable MIG with nvidia-smi -i <GPU> -mig 1; on Hopper this needs no GPU reset, according to NVIDIA’s guide.
  2. List the profiles with nvidia-smi mig -lgip and the possible placements with nvidia-smi mig -lgipp.
  3. Create the instances with nvidia-smi mig -cgi, followed by a comma-separated list of profile names and -C for the compute instances, for example nvidia-smi mig -cgi 3g.71gb,2g.35gb -C.
  4. Check the result with nvidia-smi -L, which lists each MIG device with its UUID.

Our MIG runbook for RTX PRO Blackwell cards covers these commands in detail; the H200 NVL skips its vBIOS and display-mode steps. MIG mode on Hopper “is only persistent as long as the driver is resident”, so the layout has to be recreated after every reboot.

What each H200 MIG profile runs

The choice of profile starts from the model’s weights plus room for its KV cache, then from how much compute the workload needs. The checkpoint sizes below are the files on Hugging Face; the cache for each conversation comes on top.

WORKLOADEXAMPLE AND WEIGHTSPROFILE
Notebooks, small experimentsmodels up to a few GB1g.18gb
RAG embedding and rerankerQwen3-Embedding-0.6B, about 1.2 GB in BF161g.18gb
Larger embedding modelQwen3-Embedding-8B, about 16 GB in BF161g.35gb or 2g.35gb
Team assistant, 14B modelQwen3-14B FP8, 16.3 GB2g.35gb
Department assistantQwen3-32B FP8, 34.3 GB3g.71gb or 4g.71gb
70B model in FP8Llama 3.3 70B FP8, 72.7 GB7g.141gb or whole card
Video with optical flowOFA pipelines, one NVDEC enough1g.18gb+me
Multi-GPU fine-tuningNCCL across cardswhole cards, MIG off

Checkpoint sizes from Hugging Face (Qwen/Qwen3-14B-FP8, Qwen/Qwen3-32B-FP8, nvidia/Llama-3.3-70B-Instruct-FP8, summed safetensors files, read 10 October 2026); embedding weights at two bytes per parameter; profiles are our reading of NVIDIA’s profile table.

A 1g.18gb instance at 16.5 GB holds a small model or a retrieval pair with room for batches, while the 16.3 GB Qwen3-14B checkpoint would leave it almost no cache. In a 35 GB instance the same model has about 18 GB of nominal memory beside its weights, less the serving engine’s own overhead, and an 8K conversation takes 0.625 GiB in an FP8 cache, from its 40 layers and 8 KV heads of 128 dimensions. Choose 2g.35gb when the endpoint serves many users at once and 1g.35gb when memory matters more than speed. Qwen3-32B in FP8 needs 34.3 GB for its weights and 1 GiB per 8K conversation in FP8, so it belongs in a half-card instance. Llama 3.3 70B in FP8, at 72.7 GB, does not fit 71 GB and takes the full card. Our embedding and reranker sizing guide explains the retrieval models.

We supply the H200 NVL as cards and in AI servers built to order. Tell us which models should run in which instances, with the users per model, and we reply with a configuration and quote within one business day.

MIG with Kubernetes and Slurm on the H200 NVL

In Kubernetes, the GPU Operator’s MIG Manager applies a named layout from the node label nvidia.com/mig.config, for example all-1g.18gb or all-balanced. NVIDIA’s Operator documentation, updated 23 September 2026, states that “MIG Manager requires that no user workloads are running on the GPUs being configured”, so plan layout changes as maintenance. Since Operator v26.3.0, MIG Manager generates each node’s layouts at start-up from the MIG profiles its cards report, one per profile plus all-balanced, so an H200 NVL node gets its layouts from its own profile table. On older drivers that cannot report the profiles, the Operator uses a static ConfigMap instead, and a custom ConfigMap with the key config.yaml sets any layout the generated list lacks.

A layout with more than one profile, or a node with some cards whole, needs the device plugin’s mixed strategy, under which the smallest instance appears as nvidia.com/mig-1g.18gb. Our Kubernetes GPU sharing guide explains the single and mixed strategies. Slurm expects the instances to exist before it starts, which our guide to GPU servers for research labs covers.

MIG-backed vGPU on the H200 NVL

As of October 2026, NVIDIA’s vGPU documentation lists MIG-backed vGPU for the H200 NVL. The page “Hopper H200 vGPU Types” of NVIDIA AI Enterprise 8.2, updated on 2 September 2026, gives seven MIG-backed types for the “NVIDIA H200 PCIe 141 GB (H200 NVL)”, one per GPU instance profile, shown in the first table. H200-1-18C runs up to seven per GPU, and H200-7-141C gives one VM the whole card. A footnote states that H200-1-35C and H200-1-18CME “are supported on ESXi, starting with vSphere 8.0 update 3.” The SXM board’s types carry the prefix H200X instead.

All of them are vGPU for Compute types, and the page gives “Required license edition: NVIDIA AI Enterprise”. Each H200 NVL includes a five-year NVIDIA AI Enterprise subscription, activated with its serial number. The same page lists time-sliced types for the H200 NVL from H200-4C, up to 32 per GPU, to H200-141C. The H200 NVL table gives one vGPU per GPU instance, and we found no time-sliced vGPU types inside MIG instances listed for this card. Our comparison of MIG and vGPU per card covers the hypervisors and licences.

H200 NVL MIG vs RTX PRO 6000 MIG

The RTX PRO 6000 Blackwell splits into at most four instances, 1g.24gb up to four times, 2g.48gb twice or 4g.96gb once, each also as a +gfx variant for graphics APIs. On the Server Edition, MIG instances can host graphics vGPUs for virtual workstations. The H200 NVL splits into seven smaller instances for compute only, with 141 GB of HBM3e at 4.8 TB/s and NVLink bridges for whole cards. For CAD and desktops in MIG instances, choose the RTX PRO 6000 Server Edition; for many isolated compute users, or large models on whole cards, the H200 NVL. For an existing H100 NVL estate, our H200 NVL vs H100 NVL comparison lists the older card’s profiles.

Worked example: four H200 NVL cards, two partitioned for 20 data scientists

As an example, a team of 20 data scientists shares one server with four H200 NVL cards. We assume that at most about ten notebooks run at the same moment and that the team also needs a shared model endpoint and room for fine-tuning.

Card 1 runs seven 1g.18gb instances for notebooks and small experiments. Card 2 runs the balanced layout from NVIDIA’s MIG Manager ConfigMap: two more 1g.18gb instances, a 2g.35gb with Qwen3-14B in FP8 as the team’s assistant and a 3g.71gb for larger evaluation runs. This gives nine notebook instances for the peak of about ten, and a scheduler queues the rest. Cards 3 and 4 stay whole, joined by a 2-way NVLink bridge at 900 GB/s, for fine-tuning and any job that needs NCCL across cards.

The node then needs the mixed strategy in Kubernetes, since it carries several profiles and whole cards. Slurm needs the instances recreated at start-up before it runs. If notebook demand grows, card 2 can move to seven 1g.18gb instances in a maintenance window, for fourteen in total, with the assistant and the evaluation runs moved to card 3 or 4. Each card includes its five-year AI Enterprise subscription for MIG-backed vGPU or NIM. These figures are an example, not a sizing for your team.

We build four-card H200 NVL servers to order and check the rack, power and airflow before we quote. Describe your team, its models and its peak use in the form below, and we reply with a configuration and quote within one business day.

What we supply

We supply the NVIDIA H200 NVL with its NVLink bridges and the RTX PRO 6000 Server Edition, as cards or in AI servers built to order, which are assembled and burn-in tested, with manufacturer warranty on every component. NVIDIA AI Enterprise and vGPU licences come on the same EU contract and invoice, and the five-year subscription included with each H200 NVL is part of the quote. Operating system, drivers, CUDA and a container runtime are installed on request. The full platform on top, with Kubernetes, models and RAG, is our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What MIG profiles does the H200 NVL support?
NVIDIA’s MIG user guide lists seven H200 profiles: 1g.18gb up to seven times, 1g.18gb+me once, 1g.35gb up to four, 2g.35gb up to three, 3g.71gb up to two, and 4g.71gb and 7g.141gb once each. The guide’s list of supported GPUs names the H200 NVL with a maximum of seven instances, and NVIDIA’s vGPU reference lists the same profiles for the H200 NVL.
How much memory does an H200 NVL MIG 1g.18gb instance have?
The name says 18 GB, but NVIDIA’s H200 product page gives the H200 NVL “Up to 7 MIGs @16.5GB each”, against 18 GB on the SXM board. The instance holds one eighth of the memory and one seventh of the SMs, and nvidia-smi mig -lgip shows the exact memory on the card.
What is the difference between 1g.35gb and 2g.35gb on the H200?
Both hold a quarter of the card’s memory. The 1g.35gb profile has one SM slice, one decoder and one copy engine, while 2g.35gb has two of each and twice the compute. Use 1g.35gb when a model needs memory but little compute, and 2g.35gb for an endpoint with many users at once.
Which H200 MIG profiles can be combined on one card?
Any set that fits the eight memory slices and seven SM slices without overlapping. The static MIG Manager configuration in NVIDIA’s GPU Operator repository defines a balanced layout of two 1g.18gb, one 2g.35gb and one 3g.71gb for the H200 NVL, and 4g.71gb plus 3g.71gb is valid on drivers from R510. nvidia-smi mig -lgipp lists the placements the card accepts.
Does the H200 NVL support MIG-backed vGPU?
Yes. NVIDIA AI Enterprise 8.2 lists seven MIG-backed types for the H200 NVL, from H200-1-18C, up to seven per GPU, to H200-7-141C, and states that they require an NVIDIA AI Enterprise licence. H200-1-35C and H200-1-18CME are supported on ESXi from vSphere 8.0 update 3.
H200 MIG vs RTX PRO 6000 MIG: which is better for my workloads?
The H200 NVL splits into up to seven compute-only instances, the smallest 1g.18gb, and NVIDIA’s guide lists no graphics API support for them. The RTX PRO 6000 splits into up to four instances of 24 GB with +gfx variants for graphics, and its Server Edition hosts graphics vGPUs in them. Many isolated compute users point to the H200 NVL, virtual workstations to the RTX PRO 6000 Server Edition.

Send us the workloads you plan to place in MIG instances, the models and their precision, the number of users, and whether they run on bare metal, Kubernetes, Slurm or vGPU. We reply within one business day with a configuration and quote for H200 NVL cards and the server, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna