BLOG · GUIDE ·

Llama 4 hardware requirements: Scout and Maverick VRAM, GPUs for 1 to 100 users and the licence

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Llama 4 Scout (109B parameters, 17B active, 16 experts) fits one RTX PRO 6000 or one DGX Spark for one user with NVIDIA’s 65.3 GB NVFP4 checkpoint, or one H200 NVL with NVIDIA’s 111.6 GB FP8 checkpoint
  • Llama 4 Maverick (400B parameters, 17B active, 128 experts) is 416.8 GB in Meta’s FP8 release and needs four H200 NVL or eight RTX PRO 6000 for one to 20 users
  • By our estimate, with an FP8 KV cache, a 32,768-token conversation needs up to 3 GiB on either model; 100 such conversations need eight RTX PRO 6000 or four H200 NVL for Scout, and eight H200 NVL for Maverick
  • vLLM’s Scout recipe (24 September 2026) serves NVIDIA’s FP8 checkpoint on Hopper and the FP4 one only on GPUs of compute capability 10.0, with tensor-parallel-size 1 and an FP8 KV cache; set the maximum context yourself, as Scout’s card states 10M tokens
  • Meta’s Llama 4 Acceptable Use Policy does not grant the Section 1(a) rights to Llama 4’s multimodal models to EU-domiciled individuals or companies with their principal place of business in the EU, except for end users of a product or service that incorporates them

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Llama 4 hardware requirements for Scout and Maverick

Llama 4 hardware requirements depend on the model and the checkpoint format. Llama 4 Scout fits one RTX PRO 6000 or one DGX Spark for one user with NVIDIA’s 65.3 GB NVFP4 checkpoint, or one H200 NVL with NVIDIA’s 111.6 GB FP8 checkpoint. Twenty concurrent users at 32K tokens each need two RTX PRO 6000 or two H200 NVL, and 100 users need eight RTX PRO 6000 or four H200 NVL. Llama 4 Maverick in Meta’s FP8 release, 416.8 GB, needs four H200 NVL or eight RTX PRO 6000 for one to 20 users, and eight H200 NVL for 100.

These are our estimates for running Llama 4 on-premise with conversations of 32,768 tokens in flight at the same moment, based on the repositories on Hugging Face as read on 9 October 2026. They agree with the short table in our LLM hardware requirements by model overview.

Llama 4 Scout and Maverick: experts, attention and context

As of 9 October 2026, Meta’s Hugging Face organisation lists no Llama language model newer than Llama 4 Scout and Maverick, both released on 5 April 2025; the later entries are the Llama Guard 4 and Llama Prompt Guard 2 safety models. Meta’s release post said that Llama 4 Behemoth “is still training”, and we found no published weights for it.

Both models are mixtures of experts with 17B active parameters. Scout has 16 experts and 109B parameters in total, Maverick 128 experts and 400B, while the Transformers documentation and the Hugging Face listing give the larger model as 402B. In Maverick, Meta’s post says, each token goes to a shared expert and to one of the 128 routed experts. The GPUs therefore hold all 400B parameters, while each generated token reads the weights of about 17B.

Meta calls the attention design iRoPE, where the “i” stands for interleaved attention layers. vLLM’s launch post describes the pattern: “Llama 4 interleaves global attention (without RoPE) with chunked local attention (with RoPE) in a 1:3 ratio.” The Transformers configuration defaults the chunk to 8,192 tokens. The model cards state a context of 10M tokens for Scout and 1M for Maverick, and Meta’s post says Scout was “both pre-trained and post-trained with a 256K context length”.

Llama 4 checkpoints: BF16, FP8 and NVFP4 file sizes

CHECKPOINTPUBLISHED BYFORMATSIZENATIVE ON
Scout InstructMetaBF16217.3 GBall three
Scout Instruct FP8NVIDIAFP8111.6 GBall three
Scout Instruct NVFP4NVIDIANVFP465.3 GBRTX PRO 6000, Spark
Maverick InstructMetaBF16803.2 GBall three
Maverick Instruct FP8MetaFP8416.8 GBall three
Maverick Instruct FP8NVIDIAFP8404.5 GBall three

Sums of the safetensors files in each Hugging Face repository (meta-llama and nvidia), read on 9 October 2026; formats from the model cards and vLLM’s Scout recipe; “Native on” follows NVIDIA’s specifications for the H200 NVL, RTX PRO 6000 and DGX Spark. We found no NVFP4 checkpoint of Maverick from Meta or NVIDIA.

Meta released Scout in BF16 only, and its model card says it “can fit within a single H100 GPU with on-the-fly int4 quantization”. The memory cost of quantising at load time depends on the engine. NVIDIA’s NIM documentation for vision language models, version 1.5.0, last updated on 19 December 2025, says its Scout FP8 profile needs “the same memory as BF16” because the quantisation happens on the fly. NVIDIA’s card for the NVFP4 version says it reduces disk size and GPU memory “by approximately 3.3x”.

The H200 NVL computes FP8 but has no FP4 arithmetic, and vLLM’s Scout recipe names only the FP8 checkpoint for Hopper. The RTX PRO 6000 and the DGX Spark are Blackwell systems that compute FP4 directly, as our guide to FP8, NVFP4 and MXFP4 explains. The recipe’s launch script picks the FP4 checkpoint only on GPUs of compute capability 10.0, and NVIDIA lists the RTX PRO 6000 at 12.0 and the DGX Spark at 12.1. Check that your vLLM release runs NVFP4 mixture-of-experts layers on them, or size for the FP8 checkpoint.

KV cache per conversation for Llama 4

Our estimates take 90 per cent of the memory the driver reports, less 3 GiB per card, for weights and cache: 83.0 GiB per RTX PRO 6000 and 123.4 GiB per H200 NVL. For one DGX Spark (128 GB) we take 102 GB. The method is set out in our guide to how much VRAM an LLM needs.

The cache per token is 2 × layers × KV heads × head dimension × bytes per value. The configuration files in the Llama 4 repositories are gated, so we use the Transformers defaults, which the documentation describes as similar to Scout: 48 layers, 8 KV heads and a head dimension of 128. With the FP8 cache that vLLM’s recipe sets, that is 96 KiB per token, 3 GiB per 32,768-token conversation and 6 GiB in 16-bit. We assume the same layout for Maverick, as our overview does. Across cards the cache pools, since vLLM’s blog of 7 August 2026 states that for grouped-query attention “TP splits the KV cache by those heads first” and tensor parallelism of 2, 4 or 8 divides the 8 KV heads evenly, while pipeline parallelism divides the layers. Data-parallel attention copies do not pool, so check the “GPU KV cache size” line vLLM logs at startup.

These figures count all 48 layers at full length, although three in four attend only within 8,192-token chunks. vLLM’s design notes list Llama 4 as “3 local : 1 full” but do not say how many tokens the local layers keep, so treat our numbers as an upper bound. A conversation of 1M tokens (1,048,576) needs 96 GiB of FP8 cache by this method, more than the 83.0 GiB that one RTX PRO 6000 offers for weights and cache by our rule.

GPUs for 1, 20 and 100 users

VARIANT, FORMATONE DGX SPARKRTX PRO 6000H200 NVL
Scout, NVFP4yes / no / no1 / 2 / 8use FP8
Scout, FP8no / no / no2 / 2 / 81 / 2 / 4
Scout, BF16no / no / no4 / 4 / 82 / 4 / 8
Maverick, FP8 (Meta)no / no / no8 / 8 / over 84 / 4 / 8
Maverick, BF16no / no / noover 88 / 8 / over 8

Our estimates, not measurements: cards needed for 1, 20 and 100 concurrent conversations of 32,768 tokens with an FP8 KV cache, in configurations of 1, 2, 4 or 8 cards; DGX Spark shows a memory fit only. Weights from the repository sizes above; cache layout from the Transformers defaults for Scout.

For most of these steps the cache sets the card count, because the weights fit on fewer cards. Scout in FP8 leaves about 19 GiB on one H200 NVL, room for six conversations, while two cards hold 47. On two RTX PRO 6000 the same checkpoint leaves 62.2 GiB, enough for 20 conversations with no margin, and the NVFP4 checkpoint leaves room for 35. One RTX PRO 6000 holds seven conversations with NVFP4 weights and one Spark about 11, so two separate copies on two cards serve 14, fewer than one copy split across both.

At 100 users, Scout in FP8 needs eight RTX PRO 6000, because four hold 76 conversations, while four H200 NVL hold 129. Maverick in Meta’s FP8 release leaves room for 92 conversations on eight RTX PRO 6000, and NVIDIA’s 404.5 GB version for 95, both short of 100. At 8,192 tokens per conversation the cache is a quarter of these figures, so the same cards hold about four times as many users.

We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the Llama 4 variant, your context length and peak conversations through the form below.

Serving Llama 4 with vLLM: recipe settings

vLLM’s “Quick Start Recipe for Llama 4 Scout on vLLM”, updated on 24 September 2026, serves NVIDIA’s checkpoint with --tensor-parallel-size 1 in vLLM v0.12.0, and for maximum throughput it advises: “Set TP to 1 for FP4 model and 2 for FP8 model.” Its configuration files for Blackwell and Hopper both set kv-cache-dtype: fp8, max-num-batched-tokens: 8192, async-scheduling: true and no-enable-prefix-caching: true. The recipe notes that the batch size “is limited by the amount of available GPU memory for the kv-cache after the weights are loaded.”

Set the maximum context yourself, for example --max-model-len 32768, since the recipe says it defaults to the model’s maximum and Scout’s card states 10M tokens. A smaller limit lets vLLM start on fewer cards. Images also count as input. vLLM’s launch post of April 2025 raises the limit per request with --limit-mm-per-prompt image=10, where one image is the default, and newer releases may expect a different syntax. Its long-context Scout command switches on attn_temperature_tuning.

NVIDIA’s NIM documentation for vision language models, version 1.5.0, states for NIM release 1.4.0 that “Llama 4 Maverick 17B 128E Instruct is only supported on a node of eight of one of the following GPUs” and lists the H200 NVL among them, so the four-card layout below applies to vLLM, not to that NIM. NIM in production needs an NVIDIA AI Enterprise licence, and each H200 NVL includes a five-year subscription, as our H200 NVL guide notes.

Llama 4 Maverick on four or eight cards: NVLink and PCIe

Four H200 NVL joined by a four-way NVLink bridge form one 564 GB domain, and Maverick in Meta’s FP8 release leaves about 105 GiB of it for the cache, room for 35 conversations at 32K. For 100 users, eight cards form two NVLink domains of four joined over PCIe, as our guide to H200 NVL card counts for large models explains. A common layout is tensor parallelism inside each domain and pipeline parallelism between the two, which leaves room for about 199 conversations.

Eight RTX PRO 6000 hold 768 GB but have no NVLink, so every split runs over PCIe 5.0. vLLM suggests pipeline parallelism for GPUs without NVLink, and for a mixture of experts such as Maverick, expert parallelism is a further option. Our article on one model on several GPUs over PCIe and NVLink compares the three methods. Eight Server Edition cards at up to 600 W each draw up to 4.8 kW before processors and fans.

We build servers with four or eight H200 NVL and their NVLink bridges, and we check the rack, power and airflow before we quote. Tell us the rack position, its power feed and how many users Maverick must serve.

Llama 4 licence terms and the EU clause

Both models come under the Llama 4 Community License Agreement, whose page gives a “Llama 4 Version Effective Date: April 5, 2025”. Its Section 2 begins: “If, on the Llama 4 version release date, the monthly active users of the products or services made available by or for Licensee, or Licensee’s affiliates, is greater than 700 million monthly active users in the preceding calendar month, you must request a license from Meta, which Meta may grant to you in its sole discretion, …” Anyone who distributes or makes available the models, or a product or service that contains them, must provide a copy of the agreement and prominently display “Built with Llama”. An AI model created or improved with the models or their outputs, and distributed or made available, must have “Llama” at the beginning of its name.

Section 1.b.iv makes the Llama 4 Acceptable Use Policy part of the licence. That policy states: “With respect to any multimodal models included in Llama 4, the rights granted under Section 1(a) of the Llama 4 Community License Agreement are not being granted to you if you are an individual domiciled in, or a company with a principal place of business in, the European Union.” It continues: “This restriction does not apply to end users of a product or service that incorporates any such multimodal models.” Meta’s model card calls the Llama 4 models “natively multimodal AI models”. The cards of NVIDIA’s FP8 and NVFP4 checkpoints state that their use is governed by the NVIDIA Open Model License and also name the Llama 4 Community License Agreement. Whether a clause applies to your company is a legal assessment for your legal department.

What we supply

We supply the DGX Spark Founders Edition, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, and the H200 NVL with its two-way and four-way NVLink bridges. They come as cards or in AI servers built to order, on one EU contract and invoice with manufacturer warranty. We return a configuration and a quote within one business day, sized for the Llama 4 variant, context length and users you give us, and we check the rack, power and airflow before we quote. Serving the model, RAG and MLOps on top of the hardware are our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What are the hardware requirements for Llama 4?
Llama 4 Scout fits one RTX PRO 6000 or one DGX Spark for one user with NVIDIA’s 65.3 GB NVFP4 checkpoint, or one H200 NVL with NVIDIA’s 111.6 GB FP8 checkpoint. Llama 4 Maverick in Meta’s 416.8 GB FP8 release needs four H200 NVL or eight RTX PRO 6000. These are our estimates with 90 per cent of the memory the driver reports, less 3 GiB per card, and up to 3 GiB of FP8 cache per 32K conversation.
How much VRAM does Llama 4 Scout need?
Scout’s weights take 217.3 GB in Meta’s BF16 release, 111.6 GB in NVIDIA’s FP8 checkpoint and 65.3 GB in NVIDIA’s NVFP4 checkpoint. Each conversation of 32,768 tokens adds up to 3 GiB of FP8 cache by our estimate, so 20 users need about 60 GiB on top of the weights. In FP8, that means one H200 NVL for a few users and two for 20.
How much VRAM does Llama 4 Maverick need?
Maverick is 803.2 GB in BF16, 416.8 GB in Meta’s FP8 release and 404.5 GB in NVIDIA’s FP8 version, as read on 9 October 2026. Four H200 NVL with a four-way NVLink bridge hold Meta’s FP8 weights with about 105 GiB left for the cache, room for 35 conversations at 32K by our estimate. Eight RTX PRO 6000 also hold it, with all traffic between the cards over PCIe.
Can I run Llama 4 locally on one GPU?
Scout runs on one GPU in two ways: NVIDIA’s NVFP4 checkpoint on one RTX PRO 6000 or one DGX Spark, or NVIDIA’s FP8 checkpoint on one H200 NVL. One RTX PRO 6000 then holds about seven conversations of 32K, one Spark about 11 and one H200 NVL about six, by our estimate. Maverick does not fit one GPU or one Spark in any published format.
Which GPU do I need for Llama 4 Scout?
NVIDIA’s NVFP4 checkpoint needs a Blackwell GPU such as the RTX PRO 6000 or the DGX Spark, which compute FP4 directly, while the Hopper-based H200 NVL does not; check that your vLLM release runs that checkpoint on them. On the H200 NVL, vLLM’s Scout recipe uses NVIDIA’s FP8 checkpoint, which fits one card. For 20 users at 32K, plan two RTX PRO 6000 or two H200 NVL.
Can companies in the EU use Llama 4?
Meta’s Llama 4 Acceptable Use Policy states that the Section 1(a) rights for any multimodal models in Llama 4 are not granted to individuals domiciled in, or companies with a principal place of business in, the European Union. It adds that this restriction does not apply to end users of a product or service that incorporates such models, and Meta’s model card calls the Llama 4 models natively multimodal. Whether the clause applies to your company is a legal assessment for your legal department.

Send us the Llama 4 variant, the checkpoint format, the context length and how many conversations must run at once. We reply within one business day with a configuration and a quote, and we check the rack, power and airflow before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna