Mistral hardware requirements: Mistral Small 4, Devstral, Medium 3.5 and Large 3 on-premise
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- By our memory estimate, Mistral Small 4 (119B, 6.5B active) runs for one user on one RTX PRO 6000 or one DGX Spark in its 70.8 GB NVFP4 version, and on two RTX PRO 6000 or one H200 NVL in its 120.9 GB FP8 release; Mistral’s own stated minimum for production is larger
- Ministral 3 14B (15.7 GB) and Devstral Small 2 (25.8 GB) fit one card or one Spark, but their cache takes 5 GiB per 32K conversation, so 20 users need two RTX PRO 6000 by our estimate
- The dense 128B models, Mistral Medium 3.5 and Devstral 2, need 11 GiB of 16-bit cache per 32K conversation: two cards for one user, four H200 NVL or eight RTX PRO 6000 for 20
- Mistral Large 3 (675B) needs eight H200 NVL in FP8 (682 GB) for 20 users at 32K, or four H200 NVL for one user in its 403 GB NVFP4 version, which vLLM’s recipe lists for contexts under 64K tokens
- Small 4, Large 3, Ministral 3 and Devstral Small 2 are Apache 2.0; the modified MIT licence of Medium 3.5 and Devstral 2 grants no rights to companies above a monthly revenue threshold
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Mistral hardware requirements in brief
Mistral Small 4, Mistral AI’s 119B open model of March 2026, runs for one user on one RTX PRO 6000 or one DGX Spark in its 70.8 GB NVFP4 version, and on two RTX PRO 6000 or one H200 NVL in its 120.9 GB FP8 release. Ministral 3 14B and Devstral Small 2 fit one card of either type. The dense 128B models, Mistral Medium 3.5 and Devstral 2, need two cards for one user and four H200 NVL or eight RTX PRO 6000 for 20. Mistral Large 3, with 675B parameters, needs eight H200 NVL in FP8 or four in its NVFP4 version.
These are our memory estimates for conversations of 32,768 tokens, from Mistral’s Hugging Face files as read on 9 October 2026. Mistral’s own stated minimums are larger, as explained below. Mistral AI calls itself “a French company incorporated in Paris” in its privacy policy, and a self-hosted Mistral model runs on hardware under your control. Our LLM hardware requirements by model compares Mistral Small 4 with other model families.
Open-weight Mistral models in October 2026
| MODEL | PARAMETERS | CHECKPOINT | CONTEXT | LICENCE |
|---|---|---|---|---|
| Ministral 3 14B | 13.5B plus 0.4B vision | FP8, 15.7 GB | 256K | Apache 2.0 |
| Devstral Small 2 | 24B | FP8, 25.8 GB | 256K | Apache 2.0 |
| Mistral Small 4 | 119B MoE, 6.5B active | FP8 120.9 GB; NVFP4 70.8 GB | 256K | Apache 2.0 |
| Devstral 2 | 123B, dense | FP8, about 128 GB | 256K | modified MIT |
| Mistral Medium 3.5 | 128B, dense | FP8, 133.6 GB | 256K | modified MIT |
| Mistral Large 3 | 675B MoE, 41B active | FP8 682 GB; NVFP4 403 GB | 256K | Apache 2.0 |
Model cards and file lists on huggingface.co/mistralai, read on 9 October 2026; “dense” from Mistral’s release posts of 9 December 2025 (Devstral 2) and 22 May 2026 (Medium 3.5). Sizes count one copy of the weights, since several repositories hold the same weights twice, as Mistral’s consolidated files and as Hugging Face files.
Mistral Small 4 has 128 experts, four of them active per token. Its card states 6.5B activated parameters, while Mistral’s release post of 16 March 2026 gives “6B active parameters per token (8B including embedding and output layers)”. All six cards state function calling, and all but the Devstral 2 card state image input. Ministral 3 also comes in 3B and 8B sizes.
The FP8 weights carry one scale per tensor, except those of Mistral Large 3, which use 128 × 128 blocks; TensorRT-LLM’s support matrix lists only per-tensor FP8 for the RTX PRO 6000. Mistral publishes NVFP4 versions of Small 4 and Large 3 only. The RTX PRO 6000 and the DGX Spark compute FP4 directly, while the H200 NVL computes FP8 but not FP4. On the A100 and H100, vLLM’s Mistral Large 3 recipe (24 September 2026) says vLLM can “fall back to Marlin FP4”, with a memory gain but no speed-up over FP8, and the H200 NVL is a Hopper card too. Our guide to FP8, NVFP4 and MXFP4 explains the formats.
Mistral announced Mistral Large 4 on 6 October 2026, with 1 trillion parameters and 52B active, and wrote: “We will release the weights by the end of the month.” Mistral states no format, size or context for it, so we do not size it here. Mistral’s documentation lists all Magistral versions under “Deprecated & retired models”.
KV cache per conversation: MLA and grouped-query attention
The cache per token comes from each model’s config.json, and for Large 3 from its params.json. Mistral Small 4 and Mistral Large 3 use multi-head latent attention (MLA), which stores one compressed vector per layer and token. Small 4 has 36 layers with a 256-wide latent and a 64-wide positional part, 22.5 KiB per token in 16-bit or 0.70 GiB per 32K conversation. Large 3 has 61 layers with 512 plus 64, 68.6 KiB per token or 2.14 GiB per conversation.
The dense models use grouped-query attention with eight KV heads of dimension 128. Ministral 3 14B and Devstral Small 2 have 40 layers, which gives 160 KiB per token and 5 GiB per 32K conversation. Medium 3.5 and Devstral 2 have 88 layers, 352 KiB per token and 11 GiB per conversation, about 16 times the cache of Small 4.
For grouped-query attention, vLLM’s blog post of 7 August 2026 states that “TP splits the KV cache by those heads first”, so each of two cards holds half the cache, up to eight cards for these models’ eight KV heads. For MLA it states that “the latent KV cache is replicated in full across every TP rank”, so each card holds the whole cache. Our budget follows the rule in our guide to how much VRAM an LLM needs: 90 per cent of the memory the driver reports, less 3 GiB per card, which leaves 83.0 GiB per RTX PRO 6000 and 123.4 GiB per H200 NVL. For one DGX Spark (128 GB) we take 102 GB, and the cache is counted in 16-bit, vLLM’s default.
GPUs for 1, 20 and 100 users of each Mistral model
| MODEL, FORMAT | ONE DGX SPARK | RTX PRO 6000 | H200 NVL |
|---|---|---|---|
| Ministral 3 14B, FP8 | yes / no / no | 1 / 2 / 8 | 1 / 1 / 8 |
| Devstral Small 2, FP8 | yes / no / no | 1 / 2 / 8 | 1 / 2 / 8 |
| Mistral Small 4, NVFP4 | yes / yes / no | 1 / 1 / 4 | 1 / 1 / 2, weight-only |
| Mistral Small 4, FP8 | no / no / no | 2 / 2 / 8 | 1 / 2 / 4 |
| Devstral 2, FP8 | no / no / no | 2 / 8 / over 8 | 2 / 4 / over 8 |
| Mistral Medium 3.5, FP8 | no / no / no | 2 / 8 / over 8 | 2 / 4 / over 8 |
| Mistral Large 3, NVFP4 | no / no / no | 8 / over 8 / over 8 | 4 / 8 / over 8, weight-only |
| Mistral Large 3, FP8 | no / no / no | 8 / over 8 / over 8 | 8 / 8 / over 8 |
Our estimates, not measurements: cards for 1, 20 and 100 concurrent conversations of 32,768 tokens with a 16-bit cache, in 1, 2, 4 or 8 cards with tensor parallelism and further copies on further cards; MLA cache in full on every card, grouped-query cache split. Every conversation counts at its full length at once, an upper bound. DGX Spark shows a memory fit only. Hugging Face files read on 9 October 2026.
We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the Mistral model, its format, your context length and peak requests through the form below.
Mistral Small 4 in FP8 and NVFP4
The FP8 release leaves 10.8 GiB on one H200 NVL, room for 15 conversations at 32K, so 20 users need a second card. Two RTX PRO 6000 hold half the weights each and the full cache on both, leaving 26.7 GiB per card for 38 conversations. At the model card’s --gpu_memory_utilization 0.8, each keeps about 17 GiB, room for 24.
Mistral’s release post of 16 March 2026 states: “Minimum infrastructure: 4x NVIDIA HGX H100, 2x NVIDIA HGX H200, or 1x NVIDIA DGX B200.” Its next line recommends 4x HGX H100, 4x HGX H200 or 2x DGX B200 “for optimal performance”. Mistral states no context length, load or speed for either. Our figures count memory only, with no allowance for throughput under load, so for production serving Mistral’s minimum is the vendor’s own starting point.
The NVFP4 version is 70.8 GB and leaves 17.1 GiB on one RTX PRO 6000, room for 24 conversations at 32K. One DGX Spark holds it with 20 conversations in about 80 GiB. The Spark’s memory bandwidth is 273 GB/s, against 1,597 GB/s on the RTX PRO 6000 Server Edition, so on a Spark speed limits a team before memory does.
For 100 users of the FP8 model on RTX PRO 6000, the table counts four copies on two cards each, because larger splits keep the full cache on every card and stop below 100. vLLM’s data-parallel deployment guide says that for MoE models with MLA it “can be advantageous to use data parallel for the attention layers”, so each card holds only its own conversations’ cache. Four RTX PRO 6000 would then hold 100 by our estimate, but we found no document that tests this layout on Small 4.
Devstral 2 and Mistral Medium 3.5: two dense 128B models
Neither fits one DGX Spark or, with a 32K conversation, one H200 NVL. Two RTX PRO 6000 hold either one with room for three to four conversations at 32K, and two H200 NVL for 11. Twenty users need four H200 NVL (33 conversations) or eight RTX PRO 6000 (49), since four RTX PRO 6000 hold only 18 or 19. Mistral wrote on 9 December 2025 that Devstral 2 “requires a minimum of 4 H100-class GPUs for deployment”, and on 22 May 2026 that for Medium 3.5 self-hosting is “possible on as few as four GPUs”. Neither statement gives a context or load, while our two-card figure counts memory only, so plan production serving on at least four GPUs.
An FP8 cache, set with --kv-cache-dtype fp8 in vLLM, halves the 11 GiB per conversation. By our estimate eight H200 NVL then hold 100 conversations of 32K, while eight RTX PRO 6000 stop just short of 100. A dense model reads all its weights for every generated token, so the 4.8 TB/s of the H200 NVL sets a higher ceiling per user than the RTX PRO 6000. Our guide to GPUs for a private coding assistant sizes Devstral Small 2 for developer teams.
Mistral Large 3 on four or eight H200 NVL
vLLM’s recipe states that “The Mistral-Large-3-Instruct FP8 format can be used on one 8xH200 node”. The 682 GB FP8 checkpoint leaves 44.0 GiB per card on eight H200 NVL, enough for 20 conversations at 32K. Eight H200 NVL form two NVLink domains of four joined over PCIe, as our guide to H200 NVL card counts for large models explains. Eight RTX PRO 6000 hold the FP8 weights with 3.6 GiB per card to spare, one conversation at 32K, and need an engine that runs block-scaled FP8 on that card.
The NVFP4 version is 403 GB. Four H200 NVL hold it weight-only with room for 13 conversations at 32K, and eight RTX PRO 6000 compute it natively with room for 16. vLLM’s recipe lists it “for < 64k context length” and points to FP8 beyond. Mistral’s card notes a drop in performance above 64K, and an update on it dated 9 September 2026 says the recalibrated weights “should lead to similar long-context performance as original Mistral-Large-3”. Test the NVFP4 version at your own context length before you size for it.
We build servers with four or eight H200 NVL and their NVLink bridges, and we check the rack, power and airflow before we quote. Describe the model, your rack position and its power feed in the form below.
vLLM settings in Mistral’s model cards
The cards’ vLLM commands set --tensor-parallel-size 2 for Small 4 and Devstral Small 2, and 8 for Medium 3.5, Devstral 2 and Large 3 in FP8. The Small 4 and Devstral Small 2 cards set --max-model-len 262144, and vLLM’s recipes note that a lower value saves memory. The Small 4 card selects an MLA attention backend, FLASH_ATTN_MLA for FP8 and TRITON_MLA for NVFP4. The RTX PRO 6000 has no NVLink, so every split runs over PCIe, and our guide to one model on several GPUs explains what tensor parallelism sends per token.
Licences of the open Mistral models
Mistral Small 4, Mistral Large 3, the Ministral 3 models and Devstral Small 2 are published under Apache 2.0. Mistral Medium 3.5 and Devstral 2 come under a “Modified MIT License” whose second section begins: “You are not authorized to exercise any rights under this license if the global consolidated monthly revenue of your company” (or of your employer) exceeds a threshold set in the licence. As of 9 October 2026, Mistral names no licence for Large 4, which its documentation labels “Open”. Whether a clause applies to your company is a legal assessment for your legal department. Language coverage is in our guide to open LLMs for European languages.
What we supply
We supply the DGX Spark Founders Edition, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, and the H200 NVL with its NVLink bridges, as cards or in AI servers built to order. We size the server from the model, its format, your context and peak requests. Everything comes on one EU contract and invoice with manufacturer warranty; our professional GPU range lists every card. The models, RAG and MLOps on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What are the hardware requirements for Mistral Small 4?
How much VRAM does Mistral Small 4 need?
Can I run Mistral locally on a DGX Spark?
How much VRAM does Devstral need?
What GPUs does Mistral Large 3 need?
Can I self-host Mistral models for commercial use?
Send us the Mistral model and format, the context length you will configure, your peak number of concurrent requests and the power feed at the rack position. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day