LLM hardware requirements by model: gpt-oss, Qwen3.8, DeepSeek, Llama 4, GLM-5.3 and Mistral
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- The checkpoint’s weights plus a KV cache per conversation must fit in 90 per cent of the memory the driver reports, less 3 GiB per card; at 32K tokens a conversation needs about 0.12 GiB of FP8 cache on DeepSeek-V4-Flash and about 3 GiB on Llama 4 or GLM-5.3, and under tensor parallelism DeepSeek, GLM-5.3 and Mistral Small 4 keep that cache on every card
- gpt-oss-20b (13.8 GB), gpt-oss-120b (65.3 GB, MXFP4) and Qwen3.8-27B in FP8 (30.9 GB) run on one DGX Spark, one RTX PRO 6000 or one H200 NVL; for 20 conversations at 32K, gpt-oss-120b needs a second RTX PRO 6000 or one H200 NVL
- Llama 4 Scout in FP8 and Mistral Small 4 (120.9 GB) need two RTX PRO 6000 or one to two H200 NVL, and DeepSeek-V4-Flash-0731 (167 GB, FP4 experts) two RTX PRO 6000; Qwen3.8-Flash-Next (172.78 GiB in FP8) needs two H200 NVL or four to eight RTX PRO 6000 by our estimate
- Llama 4 Maverick in FP8 (about 417 GB) needs four H200 NVL or eight RTX PRO 6000; DeepSeek-V3.2 (690 GB) and GLM-5.3 (756 GB) need eight H200 NVL, which in vLLM 0.31.0 hold 20 users at 32K with pipeline parallelism between the NVLink domains or decode context parallelism, and V3.2 also with data-parallel attention, but not with tensor parallelism over all eight; GLM-5.3 does not fit eight RTX PRO 6000 in FP8
- Meta’s Llama 4 Acceptable Use Policy does not grant the Section 1(a) rights to Llama 4’s multimodal models to individuals domiciled or companies with their principal place of business in the EU, with an exception for end users of a product or service that incorporates them; Qwen3.8-Flash-Next and GLM-5.3 have their vendors’ own licences, the others Apache 2.0 or MIT
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
LLM hardware requirements by model
LLM hardware requirements follow from the model’s memory footprint, because the GPUs must hold the checkpoint’s weights plus a KV cache for every conversation in flight. As of October 2026, gpt-oss-20b, gpt-oss-120b and Qwen3.8-27B run on one DGX Spark, one RTX PRO 6000 or one H200 NVL. Llama 4 Scout, Mistral Small 4, DeepSeek-V4-Flash and Qwen3.8-Flash-Next need one or two H200 NVL, or two to eight RTX PRO 6000, depending on the model and the number of users. Llama 4 Maverick needs four H200 NVL or eight RTX PRO 6000, and DeepSeek-V3.2 and GLM-5.3 need eight H200 NVL.
All configurations use hardware we supply. The DGX Spark Founders Edition has 128 GB of unified memory, the RTX PRO 6000 has 96 GB per card and links cards over PCIe only, and the H200 NVL has 141 GB per card, with NVLink bridges for two or four cards. The configurations are our estimates for one user and for 20 concurrent users at 32,768 tokens.
Model sizes, formats and context windows in October 2026
| MODEL | PARAMETERS | RELEASED FORMAT | CHECKPOINT | CONTEXT |
|---|---|---|---|---|
| gpt-oss-20b | 21B, 3.6B active | MXFP4 experts, BF16 rest | 13.8 GB | 131,072 |
| gpt-oss-120b | 117B, 5.1B active | MXFP4 experts, BF16 rest | 65.3 GB | 131,072 |
| Qwen3.8-27B | 27B, dense | BF16; FP8 version | 55.6 GB; FP8 30.9 GB | 262,144, up to 1M |
| Qwen3. | 125B plus 51B n-gram embedding and 4B MTP; 6B active | BF16; FP8 version | FP8 186 GB | 262,144, up to 1M |
| Deep | 284B plus draft module, 13B active | FP4 experts, FP8 rest | 167 GB | 1M |
| DeepSeek-V3. | 685B with MTP, 37B active | FP8 | 690 GB | 163,840 |
| Llama 4 Scout | 109B, 17B active | BF16; FP8 and NVFP4 from NVIDIA | FP8 111.6 GB; NVFP4 65.3 GB | 10M |
| Llama 4 Maverick | 400B, 17B active | BF16 and FP8 | FP8 about 417 GB | 1M |
| GLM-5.3 | 753B, active not stated | FP8 | 756 GB | 1M |
| Mistral Small 4 | 119B, 6.5B active | FP8; NVFP4 version | 120.9 GB; NVFP4 70.8 GB | 256K |
Model cards, file lists and config.json files on Hugging Face, read on 9 October 2026; DeepSeek-V3.2’s active parameters from DeepSeek’s V3 card, V4-Flash’s from its preview card, GLM-5.3’s context from Z.ai’s documentation. Sizes count one copy of the weights; Maverick’s as in our H200 NVL guide.
The RTX PRO 6000 and DGX Spark compute FP4 directly. The H200 NVL computes FP8 but not FP4, so vLLM runs 4-bit weights on it weight-only, with the arithmetic in 16-bit, as our guide to FP8, NVFP4 and MXFP4 explains. On the RTX PRO 6000, TensorRT-LLM’s support matrix lists only per-tensor FP8, while the FP8 checkpoints of both Qwen3.8 models, DeepSeek and GLM-5.3 use 128 × 128 block scales, so check that your engine runs them on that card.
How we estimate GPU memory for each model
The cache comes from each configuration file at 32,768 tokens per conversation: full-attention layers keep every token and sliding-window layers only the window, as vLLM’s design notes describe. It is counted in 16-bit, vLLM’s default, except for Llama 4, Qwen3.8-Flash-Next and DeepSeek-V4-Flash, where vLLM’s recipes set an FP8 cache. We assume vLLM 0.31.0 with tensor parallelism across all cards. Both DeepSeek models, GLM-5.3 and Mistral Small 4 keep one KV head per layer, and vLLM’s blog of 7 August 2026 notes that “once TP exceeds the number of KV heads, the cache starts duplicating across GPUs”, so each card holds the full cache. Usable memory is 90 per cent of what the driver reports, 95.6 GiB per RTX PRO 6000 and 140.4 GiB per H200 NVL, minus 3 GiB per card, the rule in our guide to how much VRAM an LLM needs. For one DGX Spark (128 GB) we take 102 GB, the lower end of the working set in our DGX Spark sizing article.
OpenAI’s card says MXFP4 lets gpt-oss-120b “run on a single 80GB GPU” and gpt-oss-20b “run within 16GB of memory”. Meta’s card says Llama 4 Scout “can fit within a single H100 GPU with on-the-fly int4 quantization” and that Maverick’s FP8 weights “fit on a single H100 DGX host”. Mistral’s vLLM command for Small 4 sets --tensor-parallel-size 2.
Minimum configuration for one user and for 20 users
| MODEL, FORMAT | ONE DGX SPARK | RTX PRO 6000 | H200 NVL |
|---|---|---|---|
| gpt-oss-20b, MXFP4 | yes / yes | 1 / 1 | 1 / 1 |
| gpt-oss-120b, MXFP4 | yes / yes | 1 / 2 | 1 / 1 |
| Qwen3.8-27B, FP8 | yes / yes | 1 / 1 | 1 / 1 |
| Qwen3. | no / no | 4 / 8 | 2 / 2 |
| Deep | no / no | 2 / 2 | 2 / 2, weight-only |
| DeepSeek-V3. | no / no | 8 / over 8 | 8 / 8 with PP, DCP or DP |
| Llama 4 Scout, FP8 | no / no | 2 / 2 | 1 / 2 |
| Llama 4 Maverick, FP8 | no / no | 8 / 8 | 4 / 4 |
| GLM-5.3, FP8 | no / no | over 8 / over 8 | 8 / 8 with PP or DCP |
| Mistral Small 4, FP8 | no / no | 2 / 2 | 1 / 2 |
Our estimates, not measurements, of the cards needed for one conversation and for 20 concurrent conversations of 32,768 tokens, in configurations of 1, 2, 4 or 8 cards, with tensor parallelism in vLLM 0.31.0; DGX Spark shows a memory fit only. PP, DCP and DP: pipeline, decode context and data-parallel attention, as in the section on eight cards. The H200 NVL figure for DeepSeek-V4-Flash assumes its FP4 experts run weight-only. Maverick’s configuration file is gated, so its cache uses Scout’s layout from the Transformers documentation.
The two figures in a cell differ where the cache, not the weights, sets the limit. On one RTX PRO 6000, gpt-oss-120b leaves 22.2 GiB for the cache, room for 19 conversations of 32K at 1.125 GiB each. Twenty users therefore need a second card, while one H200 NVL holds 55. Mistral Small 4 leaves 10.8 GiB on one H200 NVL, enough for one conversation but not for the 14.1 GiB that 20 need.
We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the model, your context length and peak requests in flight through the form below, and we reply with a configuration and quote.
One card or one DGX Spark: gpt-oss and Qwen3.8-27B
gpt-oss-20b, gpt-oss-120b and Qwen3.8-27B in FP8 each fit one card for a single user, and one H200 NVL holds any of them with 20 conversations at 32K. gpt-oss keeps the whole context in only half of its layers, while the other half attend over a 128-token window. A 32K conversation therefore costs 0.75 GiB of 16-bit cache on gpt-oss-20b and 1.125 GiB on gpt-oss-120b. Qwen3.8-27B keeps a full cache in 16 of its 64 layers, with four KV heads of dimension 256, and its Gated DeltaNet layers hold a fixed state per conversation, about 2 GiB per 32K conversation in all.
One DGX Spark holds all three, gpt-oss-120b with our 20 conversations at 32K in about 83 GiB and room for about 30. Its 273 GB/s of memory bandwidth, lower than either card’s, limits a team before memory does.
Llama 4 Scout, Mistral Small 4, DeepSeek-V4-Flash and Qwen3.8-Flash-Next
Llama 4 Scout in NVIDIA’s FP8 checkpoint takes two RTX PRO 6000 or one H200 NVL, and 20 conversations need a second H200 NVL. We counted all 48 layers at the full 32K, since we found no vLLM document on how it stores the chunked-attention layers. NVIDIA’s 65.3 GB NVFP4 checkpoint fits one RTX PRO 6000 or one Spark for a single user, though vLLM’s Scout recipe selects it only on compute capability 10.0, and NVIDIA lists the RTX PRO 6000 at 12.0.
Mistral Small 4 ships in FP8 at 120.9 GB, and two RTX PRO 6000 hold it for 20 users, as many GPUs as Mistral’s own command uses. Its 70.8 GB NVFP4 version fits one RTX PRO 6000 or one Spark with 20 conversations, on the card by a small margin.
DeepSeek-V4-Flash-0731, the official release, needs two RTX PRO 6000 for its 167 GB. Its compressed attention keeps the FP8 cache at about 0.12 GiB per 32K conversation by our estimate, so two cards serve 20 users even with the full cache on each. vLLM’s recipe of 29 September 2026 runs it on eight RTX PRO 6000. The H200 NVL has no native path for its FP4 experts, as our H200 NVL guide notes, so two cards hold the model only if the engine keeps the experts in FP4 and runs them weight-only, which no DeepSeek or vLLM document we found states. For high and max reasoning effort DeepSeek recommends up to 384K output tokens, and at that length the cache decides how many users fit. The newer DeepSeek-V4.1-Flash (10 September 2026, about 511 GB) is sized in our DeepSeek hardware guide.
Qwen3.8-Flash-Next needs two H200 NVL for its FP8 checkpoint, which vLLM’s recipe puts at 172.78 GiB. The recipe notes that “H100’s 80GB/GPU is not enough headroom for the 51B N-gram/PLE embedding table under plain TP4”, so we size it per card. Four RTX PRO 6000 hold it with too little left per card for 20 conversations at 32K, by our estimate. Twenty users need eight cards, or four with the table in host memory, for which the recipe asks at least 51 GB plus runtime headroom.
The recipe uses tensor-expert parallelism (TEP8) on eight H200, because “plain TP8 is incompatible with its 128-wide quantization blocks”, so eight RTX PRO 6000 would need the same layout.
Llama 4 Maverick, DeepSeek-V3.2 and GLM-5.3 on four or eight cards
Llama 4 Maverick in FP8, about 417 GB, fits four H200 NVL joined by a four-way NVLink bridge into one 564 GB domain. That leaves about 105 GiB, enough for 20 conversations at 32K. Four RTX PRO 6000 have 384 GB, too little for the weights, so Maverick needs eight, split over PCIe.
DeepSeek-V3.2 and GLM-5.3 need eight H200 NVL, which form two NVLink domains of four joined over PCIe, as our guide to H200 NVL card counts for large models explains. With tensor parallelism over all eight, every card keeps the full cache of each conversation, room for about 18 conversations of 32K on DeepSeek-V3.2 and 11 on GLM-5.3. Pipeline parallelism (PP) between the domains, the guide’s layout, halves the cache per card, for about 36 and 23. Decode context parallelism (DCP, -dcp 8), which vLLM 0.30.0 lists for sparse-MLA models, splits each cache across the cards by token, for about 144 and 92. With data-parallel attention (DP, -dp 8), which vLLM’s V3.2 recipe prefers, each card keeps only its own conversations but repeats the non-expert weights, about 88 on V3.2 by our count. Eight RTX PRO 6000 leave 2.7 GiB per card after DeepSeek-V3.2’s weights, enough for one conversation at 32K, or about nine with -dcp 8. GLM-5.3’s FP8 weights take 704 GiB, more than the 664 GiB that eight RTX PRO 6000 provide by our rule.
We build servers with four or eight H200 NVL and their NVLink bridges, and we check the rack, power and airflow before we quote. Describe the model, your rack position and its power feed in the form below.
Licence terms that limit use of a model in the EU
gpt-oss, Qwen3.8-27B and Mistral Small 4 are published under Apache 2.0, and both DeepSeek models under the MIT licence, as their model cards state. For Llama 4, Meta’s Llama 4 Acceptable Use Policy, to which the Llama 4 Community License Agreement refers, states: “With respect to any multimodal models included in Llama 4, the rights granted under Section 1(a) of the Llama 4 Community License Agreement are not being granted to you if you are an individual domiciled in, or a company with a principal place of business in, the European Union.” It continues: “This restriction does not apply to end users of a product or service that incorporates any such multimodal models.” Meta’s card calls the Llama 4 models “natively multimodal AI models”.
Qwen3.8-Flash-Next comes under the Qwen Community License 1.0, whose clause on service businesses reads: “If the licensee or any of its affiliates conducts a Model as a Service or AI Work Assistant business, the licensee shall obtain a separate license from Qwen before Using the Software or its derivative works for any commercial purpose.” The clause exempts internal use that makes neither the model, its outputs nor its capabilities available to a third party. The GLM-5.3 License asks Model-as-a-Service operators above a revenue threshold to pass a security review by Z.AI. Whether a clause applies to your company is a legal assessment for your legal department. Our private ChatGPT alternative guide compares the licences of the models it covers.
What we supply
We supply the DGX Spark Founders Edition, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, and the H200 NVL with its two-way and four-way NVLink bridges. They come as cards or in AI servers built to order, on one EU contract and invoice with manufacturer warranty. We return a configuration and a quote within one business day, and our professional GPU range lists every card. Models, RAG and MLOps on top of the hardware are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What hardware do I need to run an LLM locally?
How much VRAM does gpt-oss-120b need?
What are the hardware requirements for DeepSeek?
How much VRAM do Qwen3 and Qwen3.8 models need?
What are the hardware requirements for Llama 4?
What hardware do GLM-4.5 and GLM-5.3 need?
Send us the model, the precision you plan to run, the context length you will declare and your peak number of concurrent requests. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day