BLOG · GUIDE ·

LLM hardware requirements by model: gpt-oss, Qwen3.8, DeepSeek, Llama 4, GLM-5.3 and Mistral

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • The checkpoint’s weights plus a KV cache per conversation must fit in 90 per cent of the memory the driver reports, less 3 GiB per card; at 32K tokens a conversation needs about 0.12 GiB of FP8 cache on DeepSeek-V4-Flash and about 3 GiB on Llama 4 or GLM-5.3, and under tensor parallelism DeepSeek, GLM-5.3 and Mistral Small 4 keep that cache on every card
  • gpt-oss-20b (13.8 GB), gpt-oss-120b (65.3 GB, MXFP4) and Qwen3.8-27B in FP8 (30.9 GB) run on one DGX Spark, one RTX PRO 6000 or one H200 NVL; for 20 conversations at 32K, gpt-oss-120b needs a second RTX PRO 6000 or one H200 NVL
  • Llama 4 Scout in FP8 and Mistral Small 4 (120.9 GB) need two RTX PRO 6000 or one to two H200 NVL, and DeepSeek-V4-Flash-0731 (167 GB, FP4 experts) two RTX PRO 6000; Qwen3.8-Flash-Next (172.78 GiB in FP8) needs two H200 NVL or four to eight RTX PRO 6000 by our estimate
  • Llama 4 Maverick in FP8 (about 417 GB) needs four H200 NVL or eight RTX PRO 6000; DeepSeek-V3.2 (690 GB) and GLM-5.3 (756 GB) need eight H200 NVL, which in vLLM 0.31.0 hold 20 users at 32K with pipeline parallelism between the NVLink domains or decode context parallelism, and V3.2 also with data-parallel attention, but not with tensor parallelism over all eight; GLM-5.3 does not fit eight RTX PRO 6000 in FP8
  • Meta’s Llama 4 Acceptable Use Policy does not grant the Section 1(a) rights to Llama 4’s multimodal models to individuals domiciled or companies with their principal place of business in the EU, with an exception for end users of a product or service that incorporates them; Qwen3.8-Flash-Next and GLM-5.3 have their vendors’ own licences, the others Apache 2.0 or MIT

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

LLM hardware requirements by model

LLM hardware requirements follow from the model’s memory footprint, because the GPUs must hold the checkpoint’s weights plus a KV cache for every conversation in flight. As of October 2026, gpt-oss-20b, gpt-oss-120b and Qwen3.8-27B run on one DGX Spark, one RTX PRO 6000 or one H200 NVL. Llama 4 Scout, Mistral Small 4, DeepSeek-V4-Flash and Qwen3.8-Flash-Next need one or two H200 NVL, or two to eight RTX PRO 6000, depending on the model and the number of users. Llama 4 Maverick needs four H200 NVL or eight RTX PRO 6000, and DeepSeek-V3.2 and GLM-5.3 need eight H200 NVL.

All configurations use hardware we supply. The DGX Spark Founders Edition has 128 GB of unified memory, the RTX PRO 6000 has 96 GB per card and links cards over PCIe only, and the H200 NVL has 141 GB per card, with NVLink bridges for two or four cards. The configurations are our estimates for one user and for 20 concurrent users at 32,768 tokens.

Model sizes, formats and context windows in October 2026

MODELPARAMETERSRELEASED FORMATCHECKPOINTCONTEXT
gpt-oss-20b21B, 3.6B activeMXFP4 experts, BF16 rest13.8 GB131,072
gpt-oss-120b117B, 5.1B activeMXFP4 experts, BF16 rest65.3 GB131,072
Qwen3.8-27B27B, denseBF16; FP8 version55.6 GB; FP8 30.9 GB262,144, up to 1M
Qwen3.8-Flash-Next125B plus 51B n-gram embedding and 4B MTP; 6B activeBF16; FP8 versionFP8 186 GB262,144, up to 1M
DeepSeek-V4-Flash-0731284B plus draft module, 13B activeFP4 experts, FP8 rest167 GB1M
DeepSeek-V3.2685B with MTP, 37B activeFP8690 GB163,840
Llama 4 Scout109B, 17B activeBF16; FP8 and NVFP4 from NVIDIAFP8 111.6 GB; NVFP4 65.3 GB10M
Llama 4 Maverick400B, 17B activeBF16 and FP8FP8 about 417 GB1M
GLM-5.3753B, active not statedFP8756 GB1M
Mistral Small 4119B, 6.5B activeFP8; NVFP4 version120.9 GB; NVFP4 70.8 GB256K

Model cards, file lists and config.json files on Hugging Face, read on 9 October 2026; DeepSeek-V3.2’s active parameters from DeepSeek’s V3 card, V4-Flash’s from its preview card, GLM-5.3’s context from Z.ai’s documentation. Sizes count one copy of the weights; Maverick’s as in our H200 NVL guide.

The RTX PRO 6000 and DGX Spark compute FP4 directly. The H200 NVL computes FP8 but not FP4, so vLLM runs 4-bit weights on it weight-only, with the arithmetic in 16-bit, as our guide to FP8, NVFP4 and MXFP4 explains. On the RTX PRO 6000, TensorRT-LLM’s support matrix lists only per-tensor FP8, while the FP8 checkpoints of both Qwen3.8 models, DeepSeek and GLM-5.3 use 128 × 128 block scales, so check that your engine runs them on that card.

How we estimate GPU memory for each model

The cache comes from each configuration file at 32,768 tokens per conversation: full-attention layers keep every token and sliding-window layers only the window, as vLLM’s design notes describe. It is counted in 16-bit, vLLM’s default, except for Llama 4, Qwen3.8-Flash-Next and DeepSeek-V4-Flash, where vLLM’s recipes set an FP8 cache. We assume vLLM 0.31.0 with tensor parallelism across all cards. Both DeepSeek models, GLM-5.3 and Mistral Small 4 keep one KV head per layer, and vLLM’s blog of 7 August 2026 notes that “once TP exceeds the number of KV heads, the cache starts duplicating across GPUs”, so each card holds the full cache. Usable memory is 90 per cent of what the driver reports, 95.6 GiB per RTX PRO 6000 and 140.4 GiB per H200 NVL, minus 3 GiB per card, the rule in our guide to how much VRAM an LLM needs. For one DGX Spark (128 GB) we take 102 GB, the lower end of the working set in our DGX Spark sizing article.

OpenAI’s card says MXFP4 lets gpt-oss-120b “run on a single 80GB GPU” and gpt-oss-20b “run within 16GB of memory”. Meta’s card says Llama 4 Scout “can fit within a single H100 GPU with on-the-fly int4 quantization” and that Maverick’s FP8 weights “fit on a single H100 DGX host”. Mistral’s vLLM command for Small 4 sets --tensor-parallel-size 2.

Minimum configuration for one user and for 20 users

MODEL, FORMATONE DGX SPARKRTX PRO 6000H200 NVL
gpt-oss-20b, MXFP4yes / yes1 / 11 / 1
gpt-oss-120b, MXFP4yes / yes1 / 21 / 1
Qwen3.8-27B, FP8yes / yes1 / 11 / 1
Qwen3.8-Flash-Next, FP8no / no4 / 82 / 2
DeepSeek-V4-Flash, FP4+FP8no / no2 / 22 / 2, weight-only
DeepSeek-V3.2, FP8no / no8 / over 88 / 8 with PP, DCP or DP
Llama 4 Scout, FP8no / no2 / 21 / 2
Llama 4 Maverick, FP8no / no8 / 84 / 4
GLM-5.3, FP8no / noover 8 / over 88 / 8 with PP or DCP
Mistral Small 4, FP8no / no2 / 21 / 2

Our estimates, not measurements, of the cards needed for one conversation and for 20 concurrent conversations of 32,768 tokens, in configurations of 1, 2, 4 or 8 cards, with tensor parallelism in vLLM 0.31.0; DGX Spark shows a memory fit only. PP, DCP and DP: pipeline, decode context and data-parallel attention, as in the section on eight cards. The H200 NVL figure for DeepSeek-V4-Flash assumes its FP4 experts run weight-only. Maverick’s configuration file is gated, so its cache uses Scout’s layout from the Transformers documentation.

The two figures in a cell differ where the cache, not the weights, sets the limit. On one RTX PRO 6000, gpt-oss-120b leaves 22.2 GiB for the cache, room for 19 conversations of 32K at 1.125 GiB each. Twenty users therefore need a second card, while one H200 NVL holds 55. Mistral Small 4 leaves 10.8 GiB on one H200 NVL, enough for one conversation but not for the 14.1 GiB that 20 need.

We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the model, your context length and peak requests in flight through the form below, and we reply with a configuration and quote.

One card or one DGX Spark: gpt-oss and Qwen3.8-27B

gpt-oss-20b, gpt-oss-120b and Qwen3.8-27B in FP8 each fit one card for a single user, and one H200 NVL holds any of them with 20 conversations at 32K. gpt-oss keeps the whole context in only half of its layers, while the other half attend over a 128-token window. A 32K conversation therefore costs 0.75 GiB of 16-bit cache on gpt-oss-20b and 1.125 GiB on gpt-oss-120b. Qwen3.8-27B keeps a full cache in 16 of its 64 layers, with four KV heads of dimension 256, and its Gated DeltaNet layers hold a fixed state per conversation, about 2 GiB per 32K conversation in all.

One DGX Spark holds all three, gpt-oss-120b with our 20 conversations at 32K in about 83 GiB and room for about 30. Its 273 GB/s of memory bandwidth, lower than either card’s, limits a team before memory does.

Llama 4 Scout, Mistral Small 4, DeepSeek-V4-Flash and Qwen3.8-Flash-Next

Llama 4 Scout in NVIDIA’s FP8 checkpoint takes two RTX PRO 6000 or one H200 NVL, and 20 conversations need a second H200 NVL. We counted all 48 layers at the full 32K, since we found no vLLM document on how it stores the chunked-attention layers. NVIDIA’s 65.3 GB NVFP4 checkpoint fits one RTX PRO 6000 or one Spark for a single user, though vLLM’s Scout recipe selects it only on compute capability 10.0, and NVIDIA lists the RTX PRO 6000 at 12.0.

Mistral Small 4 ships in FP8 at 120.9 GB, and two RTX PRO 6000 hold it for 20 users, as many GPUs as Mistral’s own command uses. Its 70.8 GB NVFP4 version fits one RTX PRO 6000 or one Spark with 20 conversations, on the card by a small margin.

DeepSeek-V4-Flash-0731, the official release, needs two RTX PRO 6000 for its 167 GB. Its compressed attention keeps the FP8 cache at about 0.12 GiB per 32K conversation by our estimate, so two cards serve 20 users even with the full cache on each. vLLM’s recipe of 29 September 2026 runs it on eight RTX PRO 6000. The H200 NVL has no native path for its FP4 experts, as our H200 NVL guide notes, so two cards hold the model only if the engine keeps the experts in FP4 and runs them weight-only, which no DeepSeek or vLLM document we found states. For high and max reasoning effort DeepSeek recommends up to 384K output tokens, and at that length the cache decides how many users fit. The newer DeepSeek-V4.1-Flash (10 September 2026, about 511 GB) is sized in our DeepSeek hardware guide.

Qwen3.8-Flash-Next needs two H200 NVL for its FP8 checkpoint, which vLLM’s recipe puts at 172.78 GiB. The recipe notes that “H100’s 80GB/GPU is not enough headroom for the 51B N-gram/PLE embedding table under plain TP4”, so we size it per card. Four RTX PRO 6000 hold it with too little left per card for 20 conversations at 32K, by our estimate. Twenty users need eight cards, or four with the table in host memory, for which the recipe asks at least 51 GB plus runtime headroom.

The recipe uses tensor-expert parallelism (TEP8) on eight H200, because “plain TP8 is incompatible with its 128-wide quantization blocks”, so eight RTX PRO 6000 would need the same layout.

Llama 4 Maverick, DeepSeek-V3.2 and GLM-5.3 on four or eight cards

Llama 4 Maverick in FP8, about 417 GB, fits four H200 NVL joined by a four-way NVLink bridge into one 564 GB domain. That leaves about 105 GiB, enough for 20 conversations at 32K. Four RTX PRO 6000 have 384 GB, too little for the weights, so Maverick needs eight, split over PCIe.

DeepSeek-V3.2 and GLM-5.3 need eight H200 NVL, which form two NVLink domains of four joined over PCIe, as our guide to H200 NVL card counts for large models explains. With tensor parallelism over all eight, every card keeps the full cache of each conversation, room for about 18 conversations of 32K on DeepSeek-V3.2 and 11 on GLM-5.3. Pipeline parallelism (PP) between the domains, the guide’s layout, halves the cache per card, for about 36 and 23. Decode context parallelism (DCP, -dcp 8), which vLLM 0.30.0 lists for sparse-MLA models, splits each cache across the cards by token, for about 144 and 92. With data-parallel attention (DP, -dp 8), which vLLM’s V3.2 recipe prefers, each card keeps only its own conversations but repeats the non-expert weights, about 88 on V3.2 by our count. Eight RTX PRO 6000 leave 2.7 GiB per card after DeepSeek-V3.2’s weights, enough for one conversation at 32K, or about nine with -dcp 8. GLM-5.3’s FP8 weights take 704 GiB, more than the 664 GiB that eight RTX PRO 6000 provide by our rule.

We build servers with four or eight H200 NVL and their NVLink bridges, and we check the rack, power and airflow before we quote. Describe the model, your rack position and its power feed in the form below.

Licence terms that limit use of a model in the EU

gpt-oss, Qwen3.8-27B and Mistral Small 4 are published under Apache 2.0, and both DeepSeek models under the MIT licence, as their model cards state. For Llama 4, Meta’s Llama 4 Acceptable Use Policy, to which the Llama 4 Community License Agreement refers, states: “With respect to any multimodal models included in Llama 4, the rights granted under Section 1(a) of the Llama 4 Community License Agreement are not being granted to you if you are an individual domiciled in, or a company with a principal place of business in, the European Union.” It continues: “This restriction does not apply to end users of a product or service that incorporates any such multimodal models.” Meta’s card calls the Llama 4 models “natively multimodal AI models”.

Qwen3.8-Flash-Next comes under the Qwen Community License 1.0, whose clause on service businesses reads: “If the licensee or any of its affiliates conducts a Model as a Service or AI Work Assistant business, the licensee shall obtain a separate license from Qwen before Using the Software or its derivative works for any commercial purpose.” The clause exempts internal use that makes neither the model, its outputs nor its capabilities available to a third party. The GLM-5.3 License asks Model-as-a-Service operators above a revenue threshold to pass a security review by Z.AI. Whether a clause applies to your company is a legal assessment for your legal department. Our private ChatGPT alternative guide compares the licences of the models it covers.

What we supply

We supply the DGX Spark Founders Edition, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, and the H200 NVL with its two-way and four-way NVLink bridges. They come as cards or in AI servers built to order, on one EU contract and invoice with manufacturer warranty. We return a configuration and a quote within one business day, and our professional GPU range lists every card. Models, RAG and MLOps on top of the hardware are our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What hardware do I need to run an LLM locally?
That follows from the model’s memory footprint, because the GPUs must hold the weights plus a KV cache for every conversation in flight. As of October 2026, gpt-oss-20b, gpt-oss-120b and Qwen3.8-27B in FP8 run on one DGX Spark, one 96 GB RTX PRO 6000 or one 141 GB H200 NVL, while models from 109B to 753B parameters need one to eight H200 NVL or two to eight RTX PRO 6000. Our estimates count 32,768 tokens per conversation and 90 per cent of the memory the driver reports, less 3 GiB per card.
How much VRAM does gpt-oss-120b need?
Its checkpoint, with the expert weights in MXFP4, is 65.3 GB, or 60.8 GiB, and OpenAI’s model card says MXFP4 makes it “run on a single 80GB GPU”. Each conversation at 32K adds 1.125 GiB of 16-bit cache, so one 96 GB RTX PRO 6000 holds about 19 such conversations and one H200 NVL about 55. One DGX Spark (128 GB) holds it as well, with 20 conversations at 32K in about 83 GiB and room for about 30.
What are the hardware requirements for DeepSeek?
DeepSeek-V4-Flash, with 284B parameters of which 13B are active, is a 167 GB checkpoint with FP4 expert weights in its 0731 release, which two RTX PRO 6000 hold and compute natively. The H200 NVL has no FP4 arithmetic, so two of them hold it only if the engine keeps the experts in FP4 and runs them weight-only, although NVIDIA lists the H100 and H200 as supported hardware. DeepSeek-V3.2 in FP8 is 690 GB and needs eight H200 NVL, which serve 20 users at 32K in vLLM with pipeline parallelism, decode context parallelism or data-parallel attention, while eight RTX PRO 6000 under tensor parallelism hold its weights with room for one conversation.
How much VRAM do Qwen3 and Qwen3.8 models need?
The current dense model, Qwen3.8-27B, is 30.9 GB in Qwen’s FP8 version and serves 20 conversations at 32K on one RTX PRO 6000, one H200 NVL or one DGX Spark by our estimate. Qwen3.8-Flash-Next is 172.78 GiB in FP8, including a 51B n-gram embedding table, and needs two H200 NVL or four to eight RTX PRO 6000, because vLLM’s recipe finds 80 GB per GPU too little headroom for that table. Earlier Qwen3 sizes follow the same arithmetic, with the weights from the checkpoint and the cache from its configuration file.
What are the hardware requirements for Llama 4?
Llama 4 Scout, with 109B parameters and 17B active, is 111.6 GB in NVIDIA’s FP8 checkpoint and fits two RTX PRO 6000 or one H200 NVL for one user, while NVIDIA’s 65.3 GB NVFP4 version for Blackwell GPUs fits one RTX PRO 6000. Llama 4 Maverick, with 400B parameters, is about 417 GB in FP8 and needs four H200 NVL or eight RTX PRO 6000. Meta’s Llama 4 Acceptable Use Policy does not grant the Section 1(a) rights to its multimodal models to individuals domiciled or companies with their principal place of business in the EU, except for end users of a product or service that incorporates them; whether that applies is for the company’s legal department to assess.
What hardware do GLM-4.5 and GLM-5.3 need?
GLM-4.5, with 355B parameters, is 361 GB in FP8 and needs four H200 NVL for its weights, as our H200 NVL card count shows. As of October 2026 the current model is GLM-5.3, with 753B parameters and a 756 GB FP8 checkpoint that needs eight H200 NVL and exceeds what eight RTX PRO 6000 can hold. vLLM’s recipe for the earlier GLM-5 and GLM-5.1 also serves the FP8 model on eight H200 with 141 GB each.

Send us the model, the precision you plan to run, the context length you will declare and your peak number of concurrent requests. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna