BLOG · GUIDE ·

Mistral hardware requirements: Mistral Small 4, Devstral, Medium 3.5 and Large 3 on-premise

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • By our memory estimate, Mistral Small 4 (119B, 6.5B active) runs for one user on one RTX PRO 6000 or one DGX Spark in its 70.8 GB NVFP4 version, and on two RTX PRO 6000 or one H200 NVL in its 120.9 GB FP8 release; Mistral’s own stated minimum for production is larger
  • Ministral 3 14B (15.7 GB) and Devstral Small 2 (25.8 GB) fit one card or one Spark, but their cache takes 5 GiB per 32K conversation, so 20 users need two RTX PRO 6000 by our estimate
  • The dense 128B models, Mistral Medium 3.5 and Devstral 2, need 11 GiB of 16-bit cache per 32K conversation: two cards for one user, four H200 NVL or eight RTX PRO 6000 for 20
  • Mistral Large 3 (675B) needs eight H200 NVL in FP8 (682 GB) for 20 users at 32K, or four H200 NVL for one user in its 403 GB NVFP4 version, which vLLM’s recipe lists for contexts under 64K tokens
  • Small 4, Large 3, Ministral 3 and Devstral Small 2 are Apache 2.0; the modified MIT licence of Medium 3.5 and Devstral 2 grants no rights to companies above a monthly revenue threshold

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Mistral hardware requirements in brief

Mistral Small 4, Mistral AI’s 119B open model of March 2026, runs for one user on one RTX PRO 6000 or one DGX Spark in its 70.8 GB NVFP4 version, and on two RTX PRO 6000 or one H200 NVL in its 120.9 GB FP8 release. Ministral 3 14B and Devstral Small 2 fit one card of either type. The dense 128B models, Mistral Medium 3.5 and Devstral 2, need two cards for one user and four H200 NVL or eight RTX PRO 6000 for 20. Mistral Large 3, with 675B parameters, needs eight H200 NVL in FP8 or four in its NVFP4 version.

These are our memory estimates for conversations of 32,768 tokens, from Mistral’s Hugging Face files as read on 9 October 2026. Mistral’s own stated minimums are larger, as explained below. Mistral AI calls itself “a French company incorporated in Paris” in its privacy policy, and a self-hosted Mistral model runs on hardware under your control. Our LLM hardware requirements by model compares Mistral Small 4 with other model families.

Open-weight Mistral models in October 2026

MODELPARAMETERSCHECKPOINTCONTEXTLICENCE
Ministral 3 14B13.5B plus 0.4B visionFP8, 15.7 GB256KApache 2.0
Devstral Small 224BFP8, 25.8 GB256KApache 2.0
Mistral Small 4119B MoE, 6.5B activeFP8 120.9 GB; NVFP4 70.8 GB256KApache 2.0
Devstral 2123B, denseFP8, about 128 GB256Kmodified MIT
Mistral Medium 3.5128B, denseFP8, 133.6 GB256Kmodified MIT
Mistral Large 3675B MoE, 41B activeFP8 682 GB; NVFP4 403 GB256KApache 2.0

Model cards and file lists on huggingface.co/mistralai, read on 9 October 2026; “dense” from Mistral’s release posts of 9 December 2025 (Devstral 2) and 22 May 2026 (Medium 3.5). Sizes count one copy of the weights, since several repositories hold the same weights twice, as Mistral’s consolidated files and as Hugging Face files.

Mistral Small 4 has 128 experts, four of them active per token. Its card states 6.5B activated parameters, while Mistral’s release post of 16 March 2026 gives “6B active parameters per token (8B including embedding and output layers)”. All six cards state function calling, and all but the Devstral 2 card state image input. Ministral 3 also comes in 3B and 8B sizes.

The FP8 weights carry one scale per tensor, except those of Mistral Large 3, which use 128 × 128 blocks; TensorRT-LLM’s support matrix lists only per-tensor FP8 for the RTX PRO 6000. Mistral publishes NVFP4 versions of Small 4 and Large 3 only. The RTX PRO 6000 and the DGX Spark compute FP4 directly, while the H200 NVL computes FP8 but not FP4. On the A100 and H100, vLLM’s Mistral Large 3 recipe (24 September 2026) says vLLM can “fall back to Marlin FP4”, with a memory gain but no speed-up over FP8, and the H200 NVL is a Hopper card too. Our guide to FP8, NVFP4 and MXFP4 explains the formats.

Mistral announced Mistral Large 4 on 6 October 2026, with 1 trillion parameters and 52B active, and wrote: “We will release the weights by the end of the month.” Mistral states no format, size or context for it, so we do not size it here. Mistral’s documentation lists all Magistral versions under “Deprecated & retired models”.

KV cache per conversation: MLA and grouped-query attention

The cache per token comes from each model’s config.json, and for Large 3 from its params.json. Mistral Small 4 and Mistral Large 3 use multi-head latent attention (MLA), which stores one compressed vector per layer and token. Small 4 has 36 layers with a 256-wide latent and a 64-wide positional part, 22.5 KiB per token in 16-bit or 0.70 GiB per 32K conversation. Large 3 has 61 layers with 512 plus 64, 68.6 KiB per token or 2.14 GiB per conversation.

The dense models use grouped-query attention with eight KV heads of dimension 128. Ministral 3 14B and Devstral Small 2 have 40 layers, which gives 160 KiB per token and 5 GiB per 32K conversation. Medium 3.5 and Devstral 2 have 88 layers, 352 KiB per token and 11 GiB per conversation, about 16 times the cache of Small 4.

For grouped-query attention, vLLM’s blog post of 7 August 2026 states that “TP splits the KV cache by those heads first”, so each of two cards holds half the cache, up to eight cards for these models’ eight KV heads. For MLA it states that “the latent KV cache is replicated in full across every TP rank”, so each card holds the whole cache. Our budget follows the rule in our guide to how much VRAM an LLM needs: 90 per cent of the memory the driver reports, less 3 GiB per card, which leaves 83.0 GiB per RTX PRO 6000 and 123.4 GiB per H200 NVL. For one DGX Spark (128 GB) we take 102 GB, and the cache is counted in 16-bit, vLLM’s default.

GPUs for 1, 20 and 100 users of each Mistral model

MODEL, FORMATONE DGX SPARKRTX PRO 6000H200 NVL
Ministral 3 14B, FP8yes / no / no1 / 2 / 81 / 1 / 8
Devstral Small 2, FP8yes / no / no1 / 2 / 81 / 2 / 8
Mistral Small 4, NVFP4yes / yes / no1 / 1 / 41 / 1 / 2, weight-only
Mistral Small 4, FP8no / no / no2 / 2 / 81 / 2 / 4
Devstral 2, FP8no / no / no2 / 8 / over 82 / 4 / over 8
Mistral Medium 3.5, FP8no / no / no2 / 8 / over 82 / 4 / over 8
Mistral Large 3, NVFP4no / no / no8 / over 8 / over 84 / 8 / over 8, weight-only
Mistral Large 3, FP8no / no / no8 / over 8 / over 88 / 8 / over 8

Our estimates, not measurements: cards for 1, 20 and 100 concurrent conversations of 32,768 tokens with a 16-bit cache, in 1, 2, 4 or 8 cards with tensor parallelism and further copies on further cards; MLA cache in full on every card, grouped-query cache split. Every conversation counts at its full length at once, an upper bound. DGX Spark shows a memory fit only. Hugging Face files read on 9 October 2026.

We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the Mistral model, its format, your context length and peak requests through the form below.

Mistral Small 4 in FP8 and NVFP4

The FP8 release leaves 10.8 GiB on one H200 NVL, room for 15 conversations at 32K, so 20 users need a second card. Two RTX PRO 6000 hold half the weights each and the full cache on both, leaving 26.7 GiB per card for 38 conversations. At the model card’s --gpu_memory_utilization 0.8, each keeps about 17 GiB, room for 24.

Mistral’s release post of 16 March 2026 states: “Minimum infrastructure: 4x NVIDIA HGX H100, 2x NVIDIA HGX H200, or 1x NVIDIA DGX B200.” Its next line recommends 4x HGX H100, 4x HGX H200 or 2x DGX B200 “for optimal performance”. Mistral states no context length, load or speed for either. Our figures count memory only, with no allowance for throughput under load, so for production serving Mistral’s minimum is the vendor’s own starting point.

The NVFP4 version is 70.8 GB and leaves 17.1 GiB on one RTX PRO 6000, room for 24 conversations at 32K. One DGX Spark holds it with 20 conversations in about 80 GiB. The Spark’s memory bandwidth is 273 GB/s, against 1,597 GB/s on the RTX PRO 6000 Server Edition, so on a Spark speed limits a team before memory does.

For 100 users of the FP8 model on RTX PRO 6000, the table counts four copies on two cards each, because larger splits keep the full cache on every card and stop below 100. vLLM’s data-parallel deployment guide says that for MoE models with MLA it “can be advantageous to use data parallel for the attention layers”, so each card holds only its own conversations’ cache. Four RTX PRO 6000 would then hold 100 by our estimate, but we found no document that tests this layout on Small 4.

Devstral 2 and Mistral Medium 3.5: two dense 128B models

Neither fits one DGX Spark or, with a 32K conversation, one H200 NVL. Two RTX PRO 6000 hold either one with room for three to four conversations at 32K, and two H200 NVL for 11. Twenty users need four H200 NVL (33 conversations) or eight RTX PRO 6000 (49), since four RTX PRO 6000 hold only 18 or 19. Mistral wrote on 9 December 2025 that Devstral 2 “requires a minimum of 4 H100-class GPUs for deployment”, and on 22 May 2026 that for Medium 3.5 self-hosting is “possible on as few as four GPUs”. Neither statement gives a context or load, while our two-card figure counts memory only, so plan production serving on at least four GPUs.

An FP8 cache, set with --kv-cache-dtype fp8 in vLLM, halves the 11 GiB per conversation. By our estimate eight H200 NVL then hold 100 conversations of 32K, while eight RTX PRO 6000 stop just short of 100. A dense model reads all its weights for every generated token, so the 4.8 TB/s of the H200 NVL sets a higher ceiling per user than the RTX PRO 6000. Our guide to GPUs for a private coding assistant sizes Devstral Small 2 for developer teams.

Mistral Large 3 on four or eight H200 NVL

vLLM’s recipe states that “The Mistral-Large-3-Instruct FP8 format can be used on one 8xH200 node”. The 682 GB FP8 checkpoint leaves 44.0 GiB per card on eight H200 NVL, enough for 20 conversations at 32K. Eight H200 NVL form two NVLink domains of four joined over PCIe, as our guide to H200 NVL card counts for large models explains. Eight RTX PRO 6000 hold the FP8 weights with 3.6 GiB per card to spare, one conversation at 32K, and need an engine that runs block-scaled FP8 on that card.

The NVFP4 version is 403 GB. Four H200 NVL hold it weight-only with room for 13 conversations at 32K, and eight RTX PRO 6000 compute it natively with room for 16. vLLM’s recipe lists it “for < 64k context length” and points to FP8 beyond. Mistral’s card notes a drop in performance above 64K, and an update on it dated 9 September 2026 says the recalibrated weights “should lead to similar long-context performance as original Mistral-Large-3”. Test the NVFP4 version at your own context length before you size for it.

We build servers with four or eight H200 NVL and their NVLink bridges, and we check the rack, power and airflow before we quote. Describe the model, your rack position and its power feed in the form below.

vLLM settings in Mistral’s model cards

The cards’ vLLM commands set --tensor-parallel-size 2 for Small 4 and Devstral Small 2, and 8 for Medium 3.5, Devstral 2 and Large 3 in FP8. The Small 4 and Devstral Small 2 cards set --max-model-len 262144, and vLLM’s recipes note that a lower value saves memory. The Small 4 card selects an MLA attention backend, FLASH_ATTN_MLA for FP8 and TRITON_MLA for NVFP4. The RTX PRO 6000 has no NVLink, so every split runs over PCIe, and our guide to one model on several GPUs explains what tensor parallelism sends per token.

Licences of the open Mistral models

Mistral Small 4, Mistral Large 3, the Ministral 3 models and Devstral Small 2 are published under Apache 2.0. Mistral Medium 3.5 and Devstral 2 come under a “Modified MIT License” whose second section begins: “You are not authorized to exercise any rights under this license if the global consolidated monthly revenue of your company” (or of your employer) exceeds a threshold set in the licence. As of 9 October 2026, Mistral names no licence for Large 4, which its documentation labels “Open”. Whether a clause applies to your company is a legal assessment for your legal department. Language coverage is in our guide to open LLMs for European languages.

What we supply

We supply the DGX Spark Founders Edition, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, and the H200 NVL with its NVLink bridges, as cards or in AI servers built to order. We size the server from the model, its format, your context and peak requests. Everything comes on one EU contract and invoice with manufacturer warranty; our professional GPU range lists every card. The models, RAG and MLOps on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What are the hardware requirements for Mistral Small 4?
Mistral Small 4 has 119B parameters with 6.5B active per token and ships in FP8 at 120.9 GB, which fits two RTX PRO 6000 or one H200 NVL for one user by our memory estimate, while its 70.8 GB NVFP4 version fits one RTX PRO 6000 or one DGX Spark with 20 conversations of 32K tokens. Mistral’s release post states a larger “Minimum infrastructure” of 4x NVIDIA HGX H100, 2x NVIDIA HGX H200 or 1x NVIDIA DGX B200, without a context length or load. For 100 concurrent conversations at 32K, the FP8 model needs four H200 NVL or eight RTX PRO 6000 by our estimate.
How much VRAM does Mistral Small 4 need?
The weights take 120.9 GB in FP8 or 70.8 GB in NVFP4, as the Hugging Face file lists show. Its multi-head latent attention adds only about 0.7 GiB of 16-bit KV cache per conversation of 32,768 tokens, so one H200 NVL holds the FP8 model with 15 such conversations. With tensor parallelism vLLM keeps that cache in full on every card.
Can I run Mistral locally on a DGX Spark?
Ministral 3 14B, Devstral Small 2 and the NVFP4 version of Mistral Small 4 fit the 128 GB of one Spark, Small 4 with 20 conversations of 32K in about 80 GiB by our estimate. Mistral Small 4 in FP8, Medium 3.5, Devstral 2 and Large 3 exceed what one Spark holds. The Spark’s 273 GB/s of memory bandwidth limits speed before its memory does.
How much VRAM does Devstral need?
Devstral Small 2 (24B) is 25.8 GB in FP8 and fits one RTX PRO 6000, one H200 NVL or one DGX Spark, with 5 GiB of 16-bit KV cache per 32K conversation. Devstral 2 (123B, dense) is about 128 GB in FP8 and takes 11 GiB of cache per 32K conversation, so it needs two cards for one user and four H200 NVL or eight RTX PRO 6000 for 20. Mistral states that Devstral 2 requires a minimum of four H100-class GPUs for deployment.
What GPUs does Mistral Large 3 need?
Mistral Large 3 has 675B parameters, of which 41B are active, and its FP8 checkpoint of 682 GB runs on eight H200 NVL with room for about 20 conversations of 32K by our estimate; vLLM’s recipe names one 8xH200 node. The 403 GB NVFP4 version fits four H200 NVL, running the 4-bit weights weight-only, or eight RTX PRO 6000. vLLM’s recipe lists the NVFP4 version for contexts under 64K tokens and FP8 beyond, while Mistral’s card reports a recalibration of 9 September 2026 that should bring long-context performance close to the original model.
Can I self-host Mistral models for commercial use?
Mistral Small 4, Mistral Large 3, the Ministral 3 models and Devstral Small 2 are published under Apache 2.0, which permits commercial use. Mistral Medium 3.5 and Devstral 2 use a modified MIT licence that grants no rights to companies whose global consolidated monthly revenue exceeds a threshold set in the licence. Whether that applies to your company is a legal assessment for your legal department.

Send us the Mistral model and format, the context length you will configure, your peak number of concurrent requests and the power feed at the rack position. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna