BLOG · GUIDE ·

Mixture of experts VRAM: total vs active parameters, expert parallelism and CPU offload

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • A mixture-of-experts model keeps all its experts in memory and routes each token to a few of them, so total parameters set the GPU memory and active parameters set the compute and the weights read per token
  • Meta’s Llama 4 Maverick has 400B parameters and 17B active; its FP8 release of about 417 GB needs four H200 NVL or eight RTX PRO 6000, while one user’s token reads at least 17 GB of weights by our arithmetic
  • The KV cache follows the attention layout, not the experts: a 32K conversation takes 1.125 GiB of 16-bit cache on gpt-oss-120b, 2.14 GiB on Kimi K2 and about 0.12 GiB of FP8 cache on DeepSeek-V4-Flash
  • Expert parallelism places whole experts on different GPUs and sends tokens to them in every MoE layer; in vLLM the expert-parallel size is the tensor-parallel size times the data-parallel size, and on the RTX PRO 6000 that traffic crosses PCIe
  • llama.cpp keeps expert weights in system memory with --cpu-moe or --n-cpu-moe, and vLLM offloads weights with --cpu-offload-gb, which vLLM says “requires fast CPU-GPU interconnect”, so speed drops when experts leave the card

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Mixture of experts VRAM: total parameters set the memory

A mixture-of-experts (MoE) model needs GPU memory for all of its parameters, because any token can be routed to any expert. Only the routed experts compute for a given token, so the active parameters set the compute per token and, with the card’s memory bandwidth, the speed one user sees. Meta’s Llama 4 post of 5 April 2025 states that “while all parameters are stored in memory, only a subset of the total parameters are activated while serving these models”.

For a hardware buyer this means sizing memory from the total parameter count and the checkpoint size, then adding the KV cache per conversation, exactly as for a dense model. The active count does not reduce the memory needed. It changes how fast the model generates on a given card and how much the cards exchange when the model is split. Our LLM hardware requirements by model gives the per-model sizing; this article explains the mechanism behind those figures.

How current MoE models route each token

Each MoE layer replaces the dense feed-forward block with many smaller expert blocks and a router that picks a few of them per token. Attention, embeddings and the router stay dense and run for every token. The models in our sizing pages differ mainly in how many experts they have and how many they pick.

gpt-oss-120b has 36 layers with 128 experts each and routes every token to 4, as its config.json lists, and OpenAI’s card gives “117B parameters with 5.1B active parameters”. Llama 4 Maverick has “128 routed experts and a shared expert”, and in Meta’s words “Each token is sent to the shared expert and also to one of the 128 routed experts.” Moonshot AI’s Kimi-K2.6 card lists 384 experts, 8 selected per token and one shared expert, for 1T parameters and 32B active. Qwen3.8-Flash-Next activates “10 Routed + 1 Shared” of 512 experts, and GLM-5.3’s config.json gives 256 routed experts with 8 active per token and one shared expert.

The experts hold most of the weights. OpenAI’s gpt-oss model card of 5 August 2025 states that “The MoE weights are responsible for 90+% of the total parameter count”, which is why these releases quantise the experts first: gpt-oss and DeepSeek-V4-Flash ship FP4 experts, Kimi-K2.6 INT4 experts, while attention stays in BF16 or FP8. The RTX PRO 6000 and DGX Spark compute FP4 directly, and the H200 NVL holds 4-bit weights but computes them at higher precision, as our guide to FP8, NVFP4 and MXFP4 explains.

Total and active parameters of current MoE models

MODELTOTALACTIVECHECKPOINTSMALLEST, ONE USER
gpt-oss-20b21B3.6B13.8 GB, MXFP4 expertsone Spark or one card
gpt-oss-120b117B5.1B65.3 GB, MXFP4 expertsone Spark or one card
GLM-4.7-Flash30B3B62.5 GB, BF16one Spark or one card
Mistral Small 4119B6.5B120.9 GB, FP82 RTX PRO 6000 or 1 H200 NVL
Llama 4 Scout109B17B111.6 GB, FP82 RTX PRO 6000 or 1 H200 NVL
Qwen3.8-Flash-Next125B plus 51B n-gram, 4B MTP6B172.78 GiB (186 GB), FP84 RTX PRO 6000 or 2 H200 NVL
DeepSeek-V4-Flash-0731284B (DeepSeek), 304B (NVIDIA)13B167 GB, FP4 experts2 RTX PRO 6000 or 2 H200 NVL, weight-only
Llama 4 Maverick400B17Babout 417 GB, FP88 RTX PRO 6000 or 4 H200 NVL
DeepSeek-V3.2685B with MTP37B690 GB, FP88 RTX PRO 6000 or 8 H200 NVL
GLM-5.3753Bnot stated756 GB, FP88 H200 NVL
Kimi-K2.61T32B595.2 GB, INT4 experts8 H200 NVL or 8 RTX PRO 6000

Parameters from the vendors’ model cards and, for DeepSeek-V4-Flash, from DeepSeek’s release note and NVIDIA’s NVFP4 card; checkpoints and configurations from our model pages (read on 9 and 10 October 2026). “One card” means one RTX PRO 6000 or one H200 NVL. Fits are our estimates for one conversation of 32K, see our hub and the DeepSeek, Kimi K2, Qwen, Llama 4 and GLM pages; no Moonshot AI or vLLM document runs Kimi K2 on the RTX PRO 6000.

For DeepSeek-V4-Flash the published totals differ. DeepSeek’s release note of 24 April 2026 gives “284B total / 13B active params”, and its change log of 31 July 2026 says that the -0731 release “keeps the same model architecture and size as DeepSeek-V4-Flash-Preview”. NVIDIA’s card for its NVFP4 version of -0731, released on 31 August 2026, states “304B in total and 13B activated”. Our sizing starts from the 167 GB checkpoint, which does not depend on either count.

Total and active counts can lie far apart. Kimi-K2.6 computes with 32B parameters per token, less than a dense 70B model, yet its 595.2 GB of INT4 weights need eight H200 NVL, or by memory eight RTX PRO 6000, as our Kimi K2 hardware requirements set out. gpt-oss-120b sits at the other end: 5.1B active and a 65.3 GB checkpoint that fits one card.

We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Tell us which MoE model you plan to run, at what precision and for how many conversations at peak, and we reply with a configuration and quote.

What active parameters change: speed per user and cache

When one user generates text, each new token requires reading the weights it uses from GPU memory. For an MoE model those are the attention weights, any dense or shared-expert layers and the routed experts. Llama 4 Maverick has 17B active parameters, so in FP8 it reads at least 17 GB of weights per token by our arithmetic, while it holds about 417 GB. A token reads more than that, because some weights stay in 16-bit: Meta’s FP8 release, at 416.8 GB, is about 15 GB larger than half of its 803.2 GB BF16 release.

NVIDIA’s FP8 checkpoint of the dense Llama 3.3 70B is 72.7 GB, fits one H200 NVL, and each token reads nearly all of it. Maverick therefore needs four times as many H200 NVL as the 70B model and, by our arithmetic, reads less than half as many bytes per token.

This advantage is largest for a single request. With many requests in flight, the tokens of one decode step go to different experts, so a step reads far more than one token’s share of the experts. NVIDIA’s blog on mixture of experts makes the same observation for training: “Given that tokens are batched together in training, most if not all experts are used.” By our reading the per-user speed of an MoE model falls faster with load than the active count suggests. NVIDIA’s TensorRT-LLM blog on expert parallelism, last updated on 26 September 2026, says that large-scale EP shortens the MoE computation “by reducing expert weight loading pressure”.

The KV cache does not depend on the experts at all. It follows the attention layout: our pages give 1.125 GiB of 16-bit cache per 32K conversation on gpt-oss-120b, which keeps the full context in half its layers, 2.14 GiB of 16-bit cache on Kimi K2 with latent attention, and about 0.12 GiB of FP8 cache on DeepSeek-V4-Flash with compressed attention. Users per card therefore depend on the attention design as much as on the weights.

The 128 GB of unified memory of a DGX Spark holds gpt-oss-120b, and its 273 GB/s of memory bandwidth, lower than that of the RTX PRO 6000 or the H200 NVL, suits models that read few weights per token. By our reading, MoE models with small active counts are the ones that make sense on one Spark, and our hub notes that speed limits a team there before memory does.

Dense vs MoE models: what changes in GPU sizing

SIZING QUESTIONDENSE MODELMOE MODEL
GPU memory for weightsall parametersall parameters, experts included
Reads per token, one userall parametersdense parts plus routed experts
Reads per step, many usersall parametersup to all experts, by our reading
KV cache per conversationattention layoutattention layout, not the experts
Ways to splittensor, pipelinetensor, pipeline, expert
CPU offload targetwhole layersexpert weights first

Meta’s Llama 4 post (5 April 2025), NVIDIA’s blog on mixture of experts (14 March 2024), vLLM’s expert parallel deployment documentation (2 October 2026) and llama.cpp’s server documentation, read on 10 October 2026; rows marked “by our reading” are our interpretation.

Expert parallelism on RTX PRO 6000 and H200 NVL

vLLM’s documentation describes expert parallelism (EP) as a layout that “allows experts in Mixture-of-Experts (MoE) models to be deployed on separate GPUs”. It is switched on with --enable-expert-parallel, and the documentation gives “EP_SIZE = TP_SIZE × DP_SIZE”. With a tensor-parallel size of 1, the attention weights are replicated across the data-parallel ranks; above 1, they are sharded within each data-parallel group. Every MoE layer then sends tokens to the cards that hold their experts and collects the results, the dispatch and combine our article on tensor, pipeline and expert parallelism over PCIe and NVLink describes.

On the RTX PRO 6000, which has no NVLink on any edition, these all-to-all exchanges cross PCIe Gen5. NVIDIA’s H200 page lists a 2- or 4-way NVLink bridge at “900GB/s per GPU” for the H200 NVL beside 128 GB/s for PCIe Gen5, and in an eight-card server the bridges form two domains of four joined over PCIe, by our reading. vLLM’s page lists its DeepEP all-to-all backends for multi-node prefill and decode, and its FlashInfer ones for multi-node NVLink systems. DeepEP’s own README requires “NVLink for intranode communication” and offers PCIe kernels only in an experimental branch, so on RTX PRO 6000 servers we would plan with vLLM’s standard allgather_reducescatter backend.

The recipes on our model pages show EP on both platforms. vLLM’s DeepSeek-V4-Flash recipe sets --enable-expert-parallel on eight RTX PRO 6000 and states for the -0731 checkpoint that “Serving without speculative decoding was verified on 8× RTX PRO 6000 (PCIe, no NVLink).”, as our DeepSeek hardware requirements quote. Its DeepSeek-V3.2 recipe prefers -dp 8 --enable-expert-parallel, which gives each card its own KV cache but repeats the non-expert weights on every card.

Experts are not used evenly. NVIDIA’s TensorRT-LLM blog reports that “EP level workload imbalance issue is common for large-scale EP inference”, caused when hot experts sit on the same rank. vLLM’s Expert Parallel Load Balancer, --enable-eplb, “collects load statistics with every forward pass and periodically rebalances expert distribution”.

We build servers with four or eight H200 NVL and their NVLink bridges, or with RTX PRO 6000 Server Edition cards, and we check the rack, power and airflow before we quote. Describe the model, its precision and your rack position in the form below.

Offloading MoE experts to CPU memory: llama.cpp and vLLM

When the experts do not fit the cards, both common engines can keep weights in system memory. llama.cpp’s server documentation lists --cpu-moe, to “keep all Mixture of Experts (MoE) weights in the CPU”, and --n-cpu-moe N for the experts of the first N layers. The more general --override-tensor assigns tensors by name pattern to a buffer type. With all layers on the GPU (-ngl all), attention and the shared layers stay there and the experts sit in system memory.

vLLM offers --cpu-offload-gb, “The space in GiB to offload to CPU, per GPU”, and --cpu-offload-params, whose documentation shows that the name segment “experts” matches an expert weight. The documentation warns that offloading “requires fast CPU-GPU interconnect, as part of the model is loaded from CPU memory to GPU memory on the fly in each model forward pass”.

Generation slows down with offload. Weights on the card are read at 1,597 GB/s on the RTX PRO 6000 Server Edition and 4.8 TB/s on the H200 NVL, as NVIDIA states. Offloaded weights in vLLM come over a PCIe 5.0 x16 link that NVIDIA rates at 64 GB/s in each direction. In llama.cpp the CPU works on them at system memory speed, or, with --op-offload on by default, llama.cpp may “offload host tensor operations to device”, which again moves weights over PCIe. Offload therefore suits tests and single users, by our reading. The offloads in the vLLM recipes behind our model pages move large lookup tables rather than experts, the n-gram table of Qwen3.8-Flash-Next and the Engram tables of DeepSeek-V4.1-Flash.

What we supply

We supply the RTX PRO 6000 in its Workstation, Max-Q and Server editions, the H200 NVL with its two-way and four-way NVLink bridges, and the DGX Spark Founders Edition, as cards or in AI servers built to order. For an MoE model we size the cards from the checkpoint and the cache per conversation, and choose between replicas, a split over PCIe or an NVLink domain. Everything comes on one EU contract and invoice with manufacturer warranty, and our professional GPU range lists every card. Serving the model with vLLM, RAG and MLOps on top is our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

How much VRAM does a mixture of experts model need?
It needs memory for all of its parameters, because every expert must be on the GPUs, plus a KV cache for each conversation in flight. Llama 4 Maverick, with 400B parameters and 17B active, is about 417 GB in FP8 and needs four H200 NVL or eight RTX PRO 6000 by our estimate. The active parameter count does not reduce the memory needed.
What is the difference between active parameters and total parameters?
Total parameters are all the weights the model holds, every expert included. Active parameters are the weights one token uses: attention and other dense parts plus the experts the router picks, for example 5.1B of 117B in gpt-oss-120b. Total parameters set memory, active parameters set compute and the weights read per token.
Are MoE models faster than dense models?
For one user, an MoE model reads only its active weights per token, so it generates faster than a dense model of the same total size on the same cards. With many requests in flight, the tokens of one step route to different experts, so more of the experts are read and by our reading the advantage shrinks. The model still needs memory for all its parameters.
What is expert parallelism?
Expert parallelism places whole experts on different GPUs and sends each token to the GPUs that hold its experts, in every MoE layer. In vLLM it is switched on with --enable-expert-parallel, and the expert-parallel size equals the tensor-parallel size times the data-parallel size. On the RTX PRO 6000 that traffic crosses PCIe, while two or four H200 NVL can carry it over an NVLink bridge.
Can MoE experts be offloaded to CPU memory?
Yes. llama.cpp keeps expert weights in system memory with --cpu-moe or --n-cpu-moe, and vLLM offloads weights with --cpu-offload-gb or --cpu-offload-params. vLLM states that this requires a fast CPU-GPU interconnect because offloaded weights are loaded to the GPU in every forward pass, so generation slows down.
Do MoE models need NVLink?
They run without it, and vLLM’s recipe reports serving of DeepSeek-V4-Flash-0731 with expert parallelism, without speculative decoding, verified on eight RTX PRO 6000 over PCIe without NVLink. NVLink matters where expert and tensor-parallel traffic limit speed, and the H200 NVL bridge joins two or four cards at 900 GB/s per GPU, against 128 GB/s for PCIe Gen5. A model that fits one card needs no link at all when it runs as replicas, one copy per card.

Send us the MoE model you plan to run, its precision, your context length and the peak number of conversations in flight. We reply within one business day with a configuration and a quote in writing, and we check the rack, power and airflow before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna