Mixture of experts VRAM: total vs active parameters, expert parallelism and CPU offload
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- A mixture-of-experts model keeps all its experts in memory and routes each token to a few of them, so total parameters set the GPU memory and active parameters set the compute and the weights read per token
- Meta’s Llama 4 Maverick has 400B parameters and 17B active; its FP8 release of about 417 GB needs four H200 NVL or eight RTX PRO 6000, while one user’s token reads at least 17 GB of weights by our arithmetic
- The KV cache follows the attention layout, not the experts: a 32K conversation takes 1.125 GiB of 16-bit cache on gpt-oss-120b, 2.14 GiB on Kimi K2 and about 0.12 GiB of FP8 cache on DeepSeek-V4-Flash
- Expert parallelism places whole experts on different GPUs and sends tokens to them in every MoE layer; in vLLM the expert-parallel size is the tensor-parallel size times the data-parallel size, and on the RTX PRO 6000 that traffic crosses PCIe
- llama.cpp keeps expert weights in system memory with --cpu-moe or --n-cpu-moe, and vLLM offloads weights with --cpu-offload-gb, which vLLM says “requires fast CPU-GPU interconnect”, so speed drops when experts leave the card
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Mixture of experts VRAM: total parameters set the memory
A mixture-of-experts (MoE) model needs GPU memory for all of its parameters, because any token can be routed to any expert. Only the routed experts compute for a given token, so the active parameters set the compute per token and, with the card’s memory bandwidth, the speed one user sees. Meta’s Llama 4 post of 5 April 2025 states that “while all parameters are stored in memory, only a subset of the total parameters are activated while serving these models”.
For a hardware buyer this means sizing memory from the total parameter count and the checkpoint size, then adding the KV cache per conversation, exactly as for a dense model. The active count does not reduce the memory needed. It changes how fast the model generates on a given card and how much the cards exchange when the model is split. Our LLM hardware requirements by model gives the per-model sizing; this article explains the mechanism behind those figures.
How current MoE models route each token
Each MoE layer replaces the dense feed-forward block with many smaller expert blocks and a router that picks a few of them per token. Attention, embeddings and the router stay dense and run for every token. The models in our sizing pages differ mainly in how many experts they have and how many they pick.
gpt-oss-120b has 36 layers with 128 experts each and routes every token to 4, as its config.json lists, and OpenAI’s card gives “117B parameters with 5.1B active parameters”. Llama 4 Maverick has “128 routed experts and a shared expert”, and in Meta’s words “Each token is sent to the shared expert and also to one of the 128 routed experts.” Moonshot AI’s Kimi-K2.6 card lists 384 experts, 8 selected per token and one shared expert, for 1T parameters and 32B active. Qwen3.8-Flash-Next activates “10 Routed + 1 Shared” of 512 experts, and GLM-5.3’s config.json gives 256 routed experts with 8 active per token and one shared expert.
The experts hold most of the weights. OpenAI’s gpt-oss model card of 5 August 2025 states that “The MoE weights are responsible for 90+% of the total parameter count”, which is why these releases quantise the experts first: gpt-oss and DeepSeek-V4-Flash ship FP4 experts, Kimi-K2.6 INT4 experts, while attention stays in BF16 or FP8. The RTX PRO 6000 and DGX Spark compute FP4 directly, and the H200 NVL holds 4-bit weights but computes them at higher precision, as our guide to FP8, NVFP4 and MXFP4 explains.
Total and active parameters of current MoE models
| MODEL | TOTAL | ACTIVE | CHECKPOINT | SMALLEST, ONE USER |
|---|---|---|---|---|
| gpt-oss-20b | 21B | 3.6B | 13.8 GB, MXFP4 experts | one Spark or one card |
| gpt-oss-120b | 117B | 5.1B | 65.3 GB, MXFP4 experts | one Spark or one card |
| GLM-4. | 30B | 3B | 62.5 GB, BF16 | one Spark or one card |
| Mistral Small 4 | 119B | 6.5B | 120.9 GB, FP8 | 2 RTX PRO 6000 or 1 H200 NVL |
| Llama 4 Scout | 109B | 17B | 111.6 GB, FP8 | 2 RTX PRO 6000 or 1 H200 NVL |
| Qwen3. | 125B plus 51B n-gram, 4B MTP | 6B | 172.78 GiB (186 GB), FP8 | 4 RTX PRO 6000 or 2 H200 NVL |
| Deep | 284B (DeepSeek), 304B (NVIDIA) | 13B | 167 GB, FP4 experts | 2 RTX PRO 6000 or 2 H200 NVL, weight-only |
| Llama 4 Maverick | 400B | 17B | about 417 GB, FP8 | 8 RTX PRO 6000 or 4 H200 NVL |
| DeepSeek-V3. | 685B with MTP | 37B | 690 GB, FP8 | 8 RTX PRO 6000 or 8 H200 NVL |
| GLM-5.3 | 753B | not stated | 756 GB, FP8 | 8 H200 NVL |
| Kimi-K2.6 | 1T | 32B | 595.2 GB, INT4 experts | 8 H200 NVL or 8 RTX PRO 6000 |
Parameters from the vendors’ model cards and, for DeepSeek-V4-Flash, from DeepSeek’s release note and NVIDIA’s NVFP4 card; checkpoints and configurations from our model pages (read on 9 and 10 October 2026). “One card” means one RTX PRO 6000 or one H200 NVL. Fits are our estimates for one conversation of 32K, see our hub and the DeepSeek, Kimi K2, Qwen, Llama 4 and GLM pages; no Moonshot AI or vLLM document runs Kimi K2 on the RTX PRO 6000.
For DeepSeek-V4-Flash the published totals differ. DeepSeek’s release note of 24 April 2026 gives “284B total / 13B active params”, and its change log of 31 July 2026 says that the -0731 release “keeps the same model architecture and size as DeepSeek-V4-Flash-Preview”. NVIDIA’s card for its NVFP4 version of -0731, released on 31 August 2026, states “304B in total and 13B activated”. Our sizing starts from the 167 GB checkpoint, which does not depend on either count.
Total and active counts can lie far apart. Kimi-K2.6 computes with 32B parameters per token, less than a dense 70B model, yet its 595.2 GB of INT4 weights need eight H200 NVL, or by memory eight RTX PRO 6000, as our Kimi K2 hardware requirements set out. gpt-oss-120b sits at the other end: 5.1B active and a 65.3 GB checkpoint that fits one card.
We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Tell us which MoE model you plan to run, at what precision and for how many conversations at peak, and we reply with a configuration and quote.
What active parameters change: speed per user and cache
When one user generates text, each new token requires reading the weights it uses from GPU memory. For an MoE model those are the attention weights, any dense or shared-expert layers and the routed experts. Llama 4 Maverick has 17B active parameters, so in FP8 it reads at least 17 GB of weights per token by our arithmetic, while it holds about 417 GB. A token reads more than that, because some weights stay in 16-bit: Meta’s FP8 release, at 416.8 GB, is about 15 GB larger than half of its 803.2 GB BF16 release.
NVIDIA’s FP8 checkpoint of the dense Llama 3.3 70B is 72.7 GB, fits one H200 NVL, and each token reads nearly all of it. Maverick therefore needs four times as many H200 NVL as the 70B model and, by our arithmetic, reads less than half as many bytes per token.
This advantage is largest for a single request. With many requests in flight, the tokens of one decode step go to different experts, so a step reads far more than one token’s share of the experts. NVIDIA’s blog on mixture of experts makes the same observation for training: “Given that tokens are batched together in training, most if not all experts are used.” By our reading the per-user speed of an MoE model falls faster with load than the active count suggests. NVIDIA’s TensorRT-LLM blog on expert parallelism, last updated on 26 September 2026, says that large-scale EP shortens the MoE computation “by reducing expert weight loading pressure”.
The KV cache does not depend on the experts at all. It follows the attention layout: our pages give 1.125 GiB of 16-bit cache per 32K conversation on gpt-oss-120b, which keeps the full context in half its layers, 2.14 GiB of 16-bit cache on Kimi K2 with latent attention, and about 0.12 GiB of FP8 cache on DeepSeek-V4-Flash with compressed attention. Users per card therefore depend on the attention design as much as on the weights.
The 128 GB of unified memory of a DGX Spark holds gpt-oss-120b, and its 273 GB/s of memory bandwidth, lower than that of the RTX PRO 6000 or the H200 NVL, suits models that read few weights per token. By our reading, MoE models with small active counts are the ones that make sense on one Spark, and our hub notes that speed limits a team there before memory does.
Dense vs MoE models: what changes in GPU sizing
| SIZING QUESTION | DENSE MODEL | MOE MODEL |
|---|---|---|
| GPU memory for weights | all parameters | all parameters, experts included |
| Reads per token, one user | all parameters | dense parts plus routed experts |
| Reads per step, many users | all parameters | up to all experts, by our reading |
| KV cache per conversation | attention layout | attention layout, not the experts |
| Ways to split | tensor, pipeline | tensor, pipeline, expert |
| CPU offload target | whole layers | expert weights first |
Meta’s Llama 4 post (5 April 2025), NVIDIA’s blog on mixture of experts (14 March 2024), vLLM’s expert parallel deployment documentation (2 October 2026) and llama.cpp’s server documentation, read on 10 October 2026; rows marked “by our reading” are our interpretation.
Expert parallelism on RTX PRO 6000 and H200 NVL
vLLM’s documentation describes expert parallelism (EP) as a layout that “allows experts in Mixture-of-Experts (MoE) models to be deployed on separate GPUs”. It is switched on with --enable-expert-parallel, and the documentation gives “EP_SIZE = TP_SIZE × DP_SIZE”. With a tensor-parallel size of 1, the attention weights are replicated across the data-parallel ranks; above 1, they are sharded within each data-parallel group. Every MoE layer then sends tokens to the cards that hold their experts and collects the results, the dispatch and combine our article on tensor, pipeline and expert parallelism over PCIe and NVLink describes.
On the RTX PRO 6000, which has no NVLink on any edition, these all-to-all exchanges cross PCIe Gen5. NVIDIA’s H200 page lists a 2- or 4-way NVLink bridge at “900GB/s per GPU” for the H200 NVL beside 128 GB/s for PCIe Gen5, and in an eight-card server the bridges form two domains of four joined over PCIe, by our reading. vLLM’s page lists its DeepEP all-to-all backends for multi-node prefill and decode, and its FlashInfer ones for multi-node NVLink systems. DeepEP’s own README requires “NVLink for intranode communication” and offers PCIe kernels only in an experimental branch, so on RTX PRO 6000 servers we would plan with vLLM’s standard allgather_reducescatter backend.
The recipes on our model pages show EP on both platforms. vLLM’s DeepSeek-V4-Flash recipe sets --enable-expert-parallel on eight RTX PRO 6000 and states for the -0731 checkpoint that “Serving without speculative decoding was verified on 8× RTX PRO 6000 (PCIe, no NVLink).”, as our DeepSeek hardware requirements quote. Its DeepSeek-V3.2 recipe prefers -dp 8 --enable-expert-parallel, which gives each card its own KV cache but repeats the non-expert weights on every card.
Experts are not used evenly. NVIDIA’s TensorRT-LLM blog reports that “EP level workload imbalance issue is common for large-scale EP inference”, caused when hot experts sit on the same rank. vLLM’s Expert Parallel Load Balancer, --enable-eplb, “collects load statistics with every forward pass and periodically rebalances expert distribution”.
We build servers with four or eight H200 NVL and their NVLink bridges, or with RTX PRO 6000 Server Edition cards, and we check the rack, power and airflow before we quote. Describe the model, its precision and your rack position in the form below.
Offloading MoE experts to CPU memory: llama.cpp and vLLM
When the experts do not fit the cards, both common engines can keep weights in system memory. llama.cpp’s server documentation lists --cpu-moe, to “keep all Mixture of Experts (MoE) weights in the CPU”, and --n-cpu-moe N for the experts of the first N layers. The more general --override-tensor assigns tensors by name pattern to a buffer type. With all layers on the GPU (-ngl all), attention and the shared layers stay there and the experts sit in system memory.
vLLM offers --cpu-offload-gb, “The space in GiB to offload to CPU, per GPU”, and --cpu-offload-params, whose documentation shows that the name segment “experts” matches an expert weight. The documentation warns that offloading “requires fast CPU-GPU interconnect, as part of the model is loaded from CPU memory to GPU memory on the fly in each model forward pass”.
Generation slows down with offload. Weights on the card are read at 1,597 GB/s on the RTX PRO 6000 Server Edition and 4.8 TB/s on the H200 NVL, as NVIDIA states. Offloaded weights in vLLM come over a PCIe 5.0 x16 link that NVIDIA rates at 64 GB/s in each direction. In llama.cpp the CPU works on them at system memory speed, or, with --op-offload on by default, llama.cpp may “offload host tensor operations to device”, which again moves weights over PCIe. Offload therefore suits tests and single users, by our reading. The offloads in the vLLM recipes behind our model pages move large lookup tables rather than experts, the n-gram table of Qwen3.8-Flash-Next and the Engram tables of DeepSeek-V4.1-Flash.
What we supply
We supply the RTX PRO 6000 in its Workstation, Max-Q and Server editions, the H200 NVL with its two-way and four-way NVLink bridges, and the DGX Spark Founders Edition, as cards or in AI servers built to order. For an MoE model we size the cards from the checkpoint and the cache per conversation, and choose between replicas, a split over PCIe or an NVLink domain. Everything comes on one EU contract and invoice with manufacturer warranty, and our professional GPU range lists every card. Serving the model with vLLM, RAG and MLOps on top is our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
How much VRAM does a mixture of experts model need?
What is the difference between active parameters and total parameters?
Are MoE models faster than dense models?
What is expert parallelism?
Can MoE experts be offloaded to CPU memory?
Do MoE models need NVLink?
Send us the MoE model you plan to run, its precision, your context length and the peak number of conversations in flight. We reply within one business day with a configuration and a quote in writing, and we check the rack, power and airflow before we quote.
Talk to an expertWe reply within one business day