Llama 4 hardware requirements: Scout and Maverick VRAM, GPUs for 1 to 100 users and the licence
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Llama 4 Scout (109B parameters, 17B active, 16 experts) fits one RTX PRO 6000 or one DGX Spark for one user with NVIDIA’s 65.3 GB NVFP4 checkpoint, or one H200 NVL with NVIDIA’s 111.6 GB FP8 checkpoint
- Llama 4 Maverick (400B parameters, 17B active, 128 experts) is 416.8 GB in Meta’s FP8 release and needs four H200 NVL or eight RTX PRO 6000 for one to 20 users
- By our estimate, with an FP8 KV cache, a 32,768-token conversation needs up to 3 GiB on either model; 100 such conversations need eight RTX PRO 6000 or four H200 NVL for Scout, and eight H200 NVL for Maverick
- vLLM’s Scout recipe (24 September 2026) serves NVIDIA’s FP8 checkpoint on Hopper and the FP4 one only on GPUs of compute capability 10.0, with tensor-parallel-size 1 and an FP8 KV cache; set the maximum context yourself, as Scout’s card states 10M tokens
- Meta’s Llama 4 Acceptable Use Policy does not grant the Section 1(a) rights to Llama 4’s multimodal models to EU-domiciled individuals or companies with their principal place of business in the EU, except for end users of a product or service that incorporates them
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Llama 4 hardware requirements for Scout and Maverick
Llama 4 hardware requirements depend on the model and the checkpoint format. Llama 4 Scout fits one RTX PRO 6000 or one DGX Spark for one user with NVIDIA’s 65.3 GB NVFP4 checkpoint, or one H200 NVL with NVIDIA’s 111.6 GB FP8 checkpoint. Twenty concurrent users at 32K tokens each need two RTX PRO 6000 or two H200 NVL, and 100 users need eight RTX PRO 6000 or four H200 NVL. Llama 4 Maverick in Meta’s FP8 release, 416.8 GB, needs four H200 NVL or eight RTX PRO 6000 for one to 20 users, and eight H200 NVL for 100.
These are our estimates for running Llama 4 on-premise with conversations of 32,768 tokens in flight at the same moment, based on the repositories on Hugging Face as read on 9 October 2026. They agree with the short table in our LLM hardware requirements by model overview.
Llama 4 Scout and Maverick: experts, attention and context
As of 9 October 2026, Meta’s Hugging Face organisation lists no Llama language model newer than Llama 4 Scout and Maverick, both released on 5 April 2025; the later entries are the Llama Guard 4 and Llama Prompt Guard 2 safety models. Meta’s release post said that Llama 4 Behemoth “is still training”, and we found no published weights for it.
Both models are mixtures of experts with 17B active parameters. Scout has 16 experts and 109B parameters in total, Maverick 128 experts and 400B, while the Transformers documentation and the Hugging Face listing give the larger model as 402B. In Maverick, Meta’s post says, each token goes to a shared expert and to one of the 128 routed experts. The GPUs therefore hold all 400B parameters, while each generated token reads the weights of about 17B.
Meta calls the attention design iRoPE, where the “i” stands for interleaved attention layers. vLLM’s launch post describes the pattern: “Llama 4 interleaves global attention (without RoPE) with chunked local attention (with RoPE) in a 1:3 ratio.” The Transformers configuration defaults the chunk to 8,192 tokens. The model cards state a context of 10M tokens for Scout and 1M for Maverick, and Meta’s post says Scout was “both pre-trained and post-trained with a 256K context length”.
Llama 4 checkpoints: BF16, FP8 and NVFP4 file sizes
| CHECKPOINT | PUBLISHED BY | FORMAT | SIZE | NATIVE ON |
|---|---|---|---|---|
| Scout Instruct | Meta | BF16 | 217.3 GB | all three |
| Scout Instruct FP8 | NVIDIA | FP8 | 111.6 GB | all three |
| Scout Instruct NVFP4 | NVIDIA | NVFP4 | 65.3 GB | RTX PRO 6000, Spark |
| Maverick Instruct | Meta | BF16 | 803.2 GB | all three |
| Maverick Instruct FP8 | Meta | FP8 | 416.8 GB | all three |
| Maverick Instruct FP8 | NVIDIA | FP8 | 404.5 GB | all three |
Sums of the safetensors files in each Hugging Face repository (meta-llama and nvidia), read on 9 October 2026; formats from the model cards and vLLM’s Scout recipe; “Native on” follows NVIDIA’s specifications for the H200 NVL, RTX PRO 6000 and DGX Spark. We found no NVFP4 checkpoint of Maverick from Meta or NVIDIA.
Meta released Scout in BF16 only, and its model card says it “can fit within a single H100 GPU with on-the-fly int4 quantization”. The memory cost of quantising at load time depends on the engine. NVIDIA’s NIM documentation for vision language models, version 1.5.0, last updated on 19 December 2025, says its Scout FP8 profile needs “the same memory as BF16” because the quantisation happens on the fly. NVIDIA’s card for the NVFP4 version says it reduces disk size and GPU memory “by approximately 3.3x”.
The H200 NVL computes FP8 but has no FP4 arithmetic, and vLLM’s Scout recipe names only the FP8 checkpoint for Hopper. The RTX PRO 6000 and the DGX Spark are Blackwell systems that compute FP4 directly, as our guide to FP8, NVFP4 and MXFP4 explains. The recipe’s launch script picks the FP4 checkpoint only on GPUs of compute capability 10.0, and NVIDIA lists the RTX PRO 6000 at 12.0 and the DGX Spark at 12.1. Check that your vLLM release runs NVFP4 mixture-of-experts layers on them, or size for the FP8 checkpoint.
KV cache per conversation for Llama 4
Our estimates take 90 per cent of the memory the driver reports, less 3 GiB per card, for weights and cache: 83.0 GiB per RTX PRO 6000 and 123.4 GiB per H200 NVL. For one DGX Spark (128 GB) we take 102 GB. The method is set out in our guide to how much VRAM an LLM needs.
The cache per token is 2 × layers × KV heads × head dimension × bytes per value. The configuration files in the Llama 4 repositories are gated, so we use the Transformers defaults, which the documentation describes as similar to Scout: 48 layers, 8 KV heads and a head dimension of 128. With the FP8 cache that vLLM’s recipe sets, that is 96 KiB per token, 3 GiB per 32,768-token conversation and 6 GiB in 16-bit. We assume the same layout for Maverick, as our overview does. Across cards the cache pools, since vLLM’s blog of 7 August 2026 states that for grouped-query attention “TP splits the KV cache by those heads first” and tensor parallelism of 2, 4 or 8 divides the 8 KV heads evenly, while pipeline parallelism divides the layers. Data-parallel attention copies do not pool, so check the “GPU KV cache size” line vLLM logs at startup.
These figures count all 48 layers at full length, although three in four attend only within 8,192-token chunks. vLLM’s design notes list Llama 4 as “3 local : 1 full” but do not say how many tokens the local layers keep, so treat our numbers as an upper bound. A conversation of 1M tokens (1,048,576) needs 96 GiB of FP8 cache by this method, more than the 83.0 GiB that one RTX PRO 6000 offers for weights and cache by our rule.
GPUs for 1, 20 and 100 users
| VARIANT, FORMAT | ONE DGX SPARK | RTX PRO 6000 | H200 NVL |
|---|---|---|---|
| Scout, NVFP4 | yes / no / no | 1 / 2 / 8 | use FP8 |
| Scout, FP8 | no / no / no | 2 / 2 / 8 | 1 / 2 / 4 |
| Scout, BF16 | no / no / no | 4 / 4 / 8 | 2 / 4 / 8 |
| Maverick, FP8 (Meta) | no / no / no | 8 / 8 / over 8 | 4 / 4 / 8 |
| Maverick, BF16 | no / no / no | over 8 | 8 / 8 / over 8 |
Our estimates, not measurements: cards needed for 1, 20 and 100 concurrent conversations of 32,768 tokens with an FP8 KV cache, in configurations of 1, 2, 4 or 8 cards; DGX Spark shows a memory fit only. Weights from the repository sizes above; cache layout from the Transformers defaults for Scout.
For most of these steps the cache sets the card count, because the weights fit on fewer cards. Scout in FP8 leaves about 19 GiB on one H200 NVL, room for six conversations, while two cards hold 47. On two RTX PRO 6000 the same checkpoint leaves 62.2 GiB, enough for 20 conversations with no margin, and the NVFP4 checkpoint leaves room for 35. One RTX PRO 6000 holds seven conversations with NVFP4 weights and one Spark about 11, so two separate copies on two cards serve 14, fewer than one copy split across both.
At 100 users, Scout in FP8 needs eight RTX PRO 6000, because four hold 76 conversations, while four H200 NVL hold 129. Maverick in Meta’s FP8 release leaves room for 92 conversations on eight RTX PRO 6000, and NVIDIA’s 404.5 GB version for 95, both short of 100. At 8,192 tokens per conversation the cache is a quarter of these figures, so the same cards hold about four times as many users.
We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the Llama 4 variant, your context length and peak conversations through the form below.
Serving Llama 4 with vLLM: recipe settings
vLLM’s “Quick Start Recipe for Llama 4 Scout on vLLM”, updated on 24 September 2026, serves NVIDIA’s checkpoint with --tensor-parallel-size 1 in vLLM v0.12.0, and for maximum throughput it advises: “Set TP to 1 for FP4 model and 2 for FP8 model.” Its configuration files for Blackwell and Hopper both set kv-cache-dtype: fp8, max-num-batched-tokens: 8192, async-scheduling: true and no-enable-prefix-caching: true. The recipe notes that the batch size “is limited by the amount of available GPU memory for the kv-cache after the weights are loaded.”
Set the maximum context yourself, for example --max-model-len 32768, since the recipe says it defaults to the model’s maximum and Scout’s card states 10M tokens. A smaller limit lets vLLM start on fewer cards. Images also count as input. vLLM’s launch post of April 2025 raises the limit per request with --limit-mm-per-prompt image=10, where one image is the default, and newer releases may expect a different syntax. Its long-context Scout command switches on attn_temperature_tuning.
NVIDIA’s NIM documentation for vision language models, version 1.5.0, states for NIM release 1.4.0 that “Llama 4 Maverick 17B 128E Instruct is only supported on a node of eight of one of the following GPUs” and lists the H200 NVL among them, so the four-card layout below applies to vLLM, not to that NIM. NIM in production needs an NVIDIA AI Enterprise licence, and each H200 NVL includes a five-year subscription, as our H200 NVL guide notes.
Llama 4 Maverick on four or eight cards: NVLink and PCIe
Four H200 NVL joined by a four-way NVLink bridge form one 564 GB domain, and Maverick in Meta’s FP8 release leaves about 105 GiB of it for the cache, room for 35 conversations at 32K. For 100 users, eight cards form two NVLink domains of four joined over PCIe, as our guide to H200 NVL card counts for large models explains. A common layout is tensor parallelism inside each domain and pipeline parallelism between the two, which leaves room for about 199 conversations.
Eight RTX PRO 6000 hold 768 GB but have no NVLink, so every split runs over PCIe 5.0. vLLM suggests pipeline parallelism for GPUs without NVLink, and for a mixture of experts such as Maverick, expert parallelism is a further option. Our article on one model on several GPUs over PCIe and NVLink compares the three methods. Eight Server Edition cards at up to 600 W each draw up to 4.8 kW before processors and fans.
We build servers with four or eight H200 NVL and their NVLink bridges, and we check the rack, power and airflow before we quote. Tell us the rack position, its power feed and how many users Maverick must serve.
Llama 4 licence terms and the EU clause
Both models come under the Llama 4 Community License Agreement, whose page gives a “Llama 4 Version Effective Date: April 5, 2025”. Its Section 2 begins: “If, on the Llama 4 version release date, the monthly active users of the products or services made available by or for Licensee, or Licensee’s affiliates, is greater than 700 million monthly active users in the preceding calendar month, you must request a license from Meta, which Meta may grant to you in its sole discretion, …” Anyone who distributes or makes available the models, or a product or service that contains them, must provide a copy of the agreement and prominently display “Built with Llama”. An AI model created or improved with the models or their outputs, and distributed or made available, must have “Llama” at the beginning of its name.
Section 1.b.iv makes the Llama 4 Acceptable Use Policy part of the licence. That policy states: “With respect to any multimodal models included in Llama 4, the rights granted under Section 1(a) of the Llama 4 Community License Agreement are not being granted to you if you are an individual domiciled in, or a company with a principal place of business in, the European Union.” It continues: “This restriction does not apply to end users of a product or service that incorporates any such multimodal models.” Meta’s model card calls the Llama 4 models “natively multimodal AI models”. The cards of NVIDIA’s FP8 and NVFP4 checkpoints state that their use is governed by the NVIDIA Open Model License and also name the Llama 4 Community License Agreement. Whether a clause applies to your company is a legal assessment for your legal department.
What we supply
We supply the DGX Spark Founders Edition, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, and the H200 NVL with its two-way and four-way NVLink bridges. They come as cards or in AI servers built to order, on one EU contract and invoice with manufacturer warranty. We return a configuration and a quote within one business day, sized for the Llama 4 variant, context length and users you give us, and we check the rack, power and airflow before we quote. Serving the model, RAG and MLOps on top of the hardware are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What are the hardware requirements for Llama 4?
How much VRAM does Llama 4 Scout need?
How much VRAM does Llama 4 Maverick need?
Can I run Llama 4 locally on one GPU?
Which GPU do I need for Llama 4 Scout?
Can companies in the EU use Llama 4?
Send us the Llama 4 variant, the checkpoint format, the context length and how many conversations must run at once. We reply within one business day with a configuration and a quote, and we check the rack, power and airflow before we quote.
Talk to an expertWe reply within one business day