BLOG · GUIDE ·

Qwen hardware requirements: Qwen3.8-27B and Flash-Next on DGX Spark, RTX PRO 6000 and H200 NVL

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Qwen3.8-27B in Qwen’s FP8 version (30.9 GB) runs on one DGX Spark, one RTX PRO 6000 or one H200 NVL, and one card of either kind holds 20 conversations of 32,768 tokens by our estimate
  • Qwen3.8-27B keeps a full KV cache in 16 of its 64 layers: 64 KiB per token in 16-bit plus a fixed 144 MiB of linear-attention state, about 2.14 GiB per 32K conversation
  • Qwen3.8-Flash-Next (172.78 GiB in FP8, including a 51B n-gram embedding table) needs two H200 NVL, or four RTX PRO 6000 for one user and eight for 20, by our reading of vLLM’s recipe
  • For 100 conversations at 32K, our estimate is four RTX PRO 6000 or two H200 NVL for Qwen3.8-27B, and four H200 NVL for Flash-Next, or two with its n-gram table offloaded to host memory
  • Qwen3.8-27B is Apache 2.0; Flash-Next comes under the Qwen Community License 1.0 and the 2.4T model, whose 2.5 TB FP8 checkpoint exceeds eight H200 NVL, under the Qwen3.8-Max License

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Qwen hardware requirements for 1, 20 and 100 users

Qwen hardware requirements depend on the model size and the checkpoint format. Qwen3.8-27B, the dense model of Qwen’s current generation, runs on one DGX Spark, one RTX PRO 6000 or one H200 NVL in Qwen’s 30.9 GB FP8 version, and one card of either kind serves 20 conversations of 32,768 tokens. Qwen3.8-Flash-Next, a mixture-of-experts model with a 172.78 GiB FP8 checkpoint, needs two H200 NVL, or four RTX PRO 6000 for one user and eight for 20. For 100 concurrent conversations at 32K, our estimate is four RTX PRO 6000 or two H200 NVL for Qwen3.8-27B, and four H200 NVL for Flash-Next.

These are planning estimates, not measurements, with model facts from Qwen’s Hugging Face pages and the official vLLM recipes as read on 9 October 2026. Our LLM hardware requirements by model compares Qwen3.8 with other model families.

Qwen3.8 models, checkpoints and context in October 2026

Qwen’s Hugging Face organisation lists three Qwen3.8 language models, each in BF16 and in Qwen’s own FP8 version, published in August 2026.

MODELPARAMETERSLAYOUTCHECKPOINTSCONTEXT
Qwen3.8-27B27B, dense, vision encoderfull attention in 16 of 64 layersBF16 55.6 GB; FP8 30.9 GB262,144, up to 1M
Qwen3.8-Flash-Next125B plus 51B n-gram embedding and 4B MTP; 6B active; vision encodersparse attention in 12 of 48 layers; 512 expertsFP8 172.78 GiB; BF16 335.28 GiB262,144, up to 1M
Qwen3.8-2.4T-A95B2.4T, 95B active, text onlyfull attention in 23 of 92 layers; 512 expertsFP8 repository 2.5 TB262,144, up to 1,010,000

Model cards, file lists and config.json files on Hugging Face (Qwen), and vLLM’s Qwen3.8-Flash-Next recipe (updated 30 September 2026) for the Flash-Next checkpoint sizes, read on 9 October 2026.

The FP8 versions of the 27B and Flash-Next models scale weights in blocks of 128 × 128 and keep the token embeddings and the output layer in 16-bit, as their config.json files state. On Hugging Face we found no Qwen3.8 checkpoint from NVIDIA and no NVFP4 version from Qwen. vLLM’s 27B recipe runs third-party NVFP4 checkpoints and states that the model “Fits one Blackwell GPU in every precision: NVFP4 in 24.6 GiB”.

The 27B card states that “Qwen3.8 models operate in thinking mode by default”, and a request turns it off with enable_thinking set to false, as on Flash-Next. The 2.4T model, in its card’s words, “requires thinking mode for all interactions”.

KV cache per conversation in Qwen3.8’s hybrid layout

We size every configuration with one rule: the weights plus a KV cache per conversation must fit in 90 per cent of the memory the driver reports, less 3 GiB per card, which leaves 83.04 GiB on an RTX PRO 6000 and 123.36 GiB on an H200 NVL. For one DGX Spark (128 GB), the version we supply, we take 102 GB. Our guide to how much VRAM an LLM needs explains the rule.

In Qwen3.8, three of every four layers use Gated DeltaNet, a linear attention that keeps a fixed state per conversation instead of a per-token cache. Qwen3.8-27B’s config.json gives 16 full-attention layers with four KV heads of dimension 256, which makes 64 KiB per token in 16-bit. Its 48 DeltaNet layers each hold 48 value heads of 128 × 128 values in float32, the mamba_ssm_dtype the file sets, so 144 MiB per conversation whatever its length. Flash-Next has 12 sparse-attention layers with two KV heads of 256, 12 KiB per token in FP8. Its indexer adds one 128-wide key per layer, which we count at up to 1.5 KiB per token, and its 36 DeltaNet layers hold 108 MiB.

MODEL, CACHEPER TOKEN32K TOKENS262,144 TOKENS
Qwen3.8-27B, 16-bit64 KiB2.14 GiB16.1 GiB
Qwen3.8-27B, FP832 KiB1.14 GiB8.1 GiB
Flash-Next, FP812 KiB, plus indexerabout 0.5 GiBabout 3.5 GiB

Our arithmetic from the config.json files of Qwen3.8-27B and Qwen3.8-Flash-Next-FP8, for one conversation and the whole model, including the linear-attention state. For several cards we assume vLLM splits the KV and DeltaNet heads between the GPUs, each keeping at least one KV head, so Flash-Next’s two KV heads cost more per card than an even share on four or eight; we found no vLLM document that states this.

A 1M-token conversation on Qwen3.8-27B would need about 61 GiB of 16-bit cache. The cards recommend an output limit of 262,144 tokens for reasoning content, and preserve_thinking, on by default, keeps the thinking blocks of earlier turns in the conversation. The context you declare to vLLM decides how many full-length conversations fit.

GPUs for Qwen3.8 at 1, 20 and 100 users

MODEL, FORMATONE DGX SPARKRTX PRO 6000H200 NVL
Qwen3.8-27B, FP8yes / yes / no1 / 1 / 41 / 1 / 2
Qwen3.8-27B, BF16yes / yes / no1 / 2 / 41 / 1 / 4
Flash-Next, FP8no / no / no4 / 8 / over 82 / 2 / 4
Flash-Next, PLE offloadno / no / no4 / 4 / 42 / 2 / 2

Our estimates, not measurements, of the cards needed for 1, 20 and 100 concurrent conversations of 32,768 tokens, in configurations of 1, 2, 4 or 8 cards; 16-bit cache for Qwen3.8-27B and FP8 cache for Flash-Next, as vLLM’s Flash-Next recipe sets. DGX Spark shows a memory fit only, for the 128 GB version. PLE offload keeps the n-gram table in host memory, which vLLM’s recipe validates on H100 only.

Qwen3.8-27B in FP8 leaves 54.3 GiB for the cache on one RTX PRO 6000, room for 25 conversations at 32K, and 94.6 GiB, room for 44, on one H200 NVL. Split over four RTX PRO 6000 with tensor parallelism, it holds about 141 conversations, and two H200 NVL hold about 101. An FP8 cache, which most CUDA commands in vLLM’s 27B recipe set, halves the cost per conversation: 47 fit one RTX PRO 6000, so two cards serve 100 users.

Since the 27B weights fit one card, four RTX PRO 6000 can also run one copy each, 25 conversations per copy with no traffic over PCIe, a trade-off our article on splitting one model over PCIe or NVLink works through.

We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the Qwen variant, your context length and peak conversations through the form below, and we reply with a configuration and a quote.

Qwen3.8-27B on one card or one DGX Spark

The BF16 checkpoint fits one card as well, but it leaves 31.3 GiB on an RTX PRO 6000, enough for 14 conversations at 32K, so Qwen’s FP8 version is the practical choice for a team. vLLM’s 27B recipe, updated on 8 October 2026, puts the FP8 weights at 14.28 GiB per GPU on two GPUs, in line with our arithmetic.

One 128 GB DGX Spark holds either checkpoint, the FP8 one with room for about 30 conversations at 32K. On a Spark, speed limits a team before memory does. Its memory bandwidth is 273 GB/s, against 1,597 GB/s on the RTX PRO 6000 Server Edition and 4.8 TB/s on the H200 NVL. A dense model reads all its weights for every token, so one Spark suits a developer or a small pilot.

On the RTX PRO 6000, check the engine before you buy. The 27B recipe reports vLLM runs on two consumer Blackwell cards of the same sm120 chip family and states for them that “The block-scaled FP8 checkpoint needs no workaround.” TensorRT-LLM’s support matrix (version 1.3.0rc29) lists only per-tensor FP8 for that chip family, not the 128 × 128 block scales of Qwen’s FP8 version.

Qwen3.8-Flash-Next: the n-gram table sets the card count

Flash-Next’s 51B n-gram embedding table, about 47.5 GiB in FP8, raises its card count. vLLM’s recipe states that “H100’s 80GB/GPU is not enough headroom for the 51B N-gram/PLE embedding table” under plain four-way tensor parallelism. We read that as one copy of the table on every GPU, plus each GPU’s share of the rest.

On that reading, two H200 NVL carry 110.1 GiB each and keep 13.2 GiB for the cache, enough for about 46 conversations at 32K. Four RTX PRO 6000 carry 78.8 GiB each and keep 4.2 GiB, about 16 conversations. Eight RTX PRO 6000 keep 19.9 GiB per card, about 80 conversations, which serves 20 users but not 100. Four H200 NVL keep 44.5 GiB per card, room for about 170.

The recipe also offers an alternative. With VLLM_PLE_CPU_OFFLOAD=1 the table moves to host memory, for which it asks “at least 51 GB plus runtime headroom”. Four RTX PRO 6000 then keep 51.7 GiB per card, and two H200 NVL keep 60.7 GiB, either enough for 100 conversations by our estimate. The recipe validates the offload on four H100 only, so test it on your platform first.

On eight GPUs the recipe uses tensor parallelism with expert parallelism (TEP8), because “plain TP8 is incompatible with its 128-wide quantization blocks”. Its command for eight H200 uses TEP8, and its KV-offloading test ran four-way tensor parallelism on four H200. Two-way tensor parallelism is validated only on GB300, and the body has no RTX PRO 6000 command, so our layouts for two H200 NVL and for RTX PRO 6000 are estimates. On RTX PRO 6000 cards, all-reduce and expert traffic run over PCIe, while two or four H200 NVL can share an NVLink bridge.

We build servers with four or eight RTX PRO 6000 or H200 NVL, and we check the rack, power and airflow before we quote. Describe your rack position, its power feed and the host memory you plan in the form below.

vLLM settings for Qwen3.8 from the official recipes

vLLM’s Flash-Next recipe, updated on 30 September 2026, serves the FP8 checkpoint on eight H200 with --tensor-parallel-size 8, --enable-expert-parallel, --moe-backend triton and --gpu-memory-utilization 0.85. The settings that change memory are --kv-cache-dtype fp8 and --attention-config.indexer_kv_dtype fp8, which store both caches in FP8. It adds --enable-prefix-caching and multi-token prediction with three speculative tokens; the MTP layer keeps a small cache of its own.

Both recipes use --reasoning-parser qwen3, and their commands with tool calling add --tool-call-parser qwen3_coder. Without --max-model-len, vLLM takes the native 262,144 tokens, and for Mamba cache errors at startup the Flash-Next recipe says to keep --max-num-seqs 256.

The 1M-token context needs a YaRN setting. The 27B card warns that “All the notable open-source frameworks implement static YaRN”, which can affect quality on shorter texts, and advises enabling it only when long contexts are needed.

Qwen3.8-2.4T-A95B and earlier Qwen3 models

Qwen3.8-2.4T-A95B is a 2.5 TB repository in FP8, more than the 1,128 GB of eight H200 NVL, so by weights alone it needs at least three eight-card servers, a multi-node cluster beyond the configurations sized here.

Earlier Qwen models follow the same arithmetic. Our guide to H200 NVL card counts for large models puts Qwen3-235B-A22B-2507, by weights alone, at two H200 NVL in FP8 (236 GB) and four in BF16 (470 GB), and Qwen3.5-397B in FP8 (406 GB) at four. Qwen3-235B’s 134 GB NVFP4 version fits two RTX PRO 6000, which run FP4 natively, while the H200 NVL runs it only weight-only. For the Qwen3-Coder models, our private coding assistant guide gives the card count.

Licence terms of the Qwen3.8 models

Qwen3.8-27B is published under Apache 2.0, as its model card states. Qwen3.8-Flash-Next comes under the Qwen Community License 1.0, whose clause 2 reads: “If the licensee or any of its affiliates conducts a Model as a Service or AI Work Assistant business, the licensee shall obtain a separate license from Qwen before Using the Software or its derivative works for any commercial purpose.” The clause exempts internal use that makes neither the software, its outputs nor its model capabilities available to any third party.

The 2.4T model has its own Qwen3.8-Max License with clauses of the same structure, but its clause 2 applies only above a revenue threshold. Both licences define Model as a Service as “giving a third party access to language model inference or fine-tuning” in a way that lets the third party “exercise meaningful control over the inputs, parameters, or training data”. Neither text contains a territorial restriction as read on 9 October 2026. Whether a clause applies to your company is a legal assessment for your legal department.

What we supply

We supply the DGX Spark Founders Edition, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, and the H200 NVL with its two-way and four-way NVLink bridges, as cards or in AI servers built to order. They come on one EU contract and invoice with manufacturer warranty, and our professional GPU range lists every card. With the Qwen variant, the context length and the number of concurrent users, we return a configuration and a quote within one business day. Running the model with RAG and MLOps on top is our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What are the hardware requirements for Qwen?
As of October 2026, Qwen3.8-27B in Qwen’s 30.9 GB FP8 version runs on one DGX Spark, one 96 GB RTX PRO 6000 or one 141 GB H200 NVL, and one card serves 20 conversations of 32K tokens by our estimate. Qwen3.8-Flash-Next, 172.78 GiB in FP8, needs two H200 NVL, or four to eight RTX PRO 6000. The 2.4T model, 2.5 TB in FP8, exceeds an eight-card H200 NVL server.
How much VRAM does Qwen3.8-27B need?
The weights take 30.9 GB in Qwen’s FP8 version and 55.6 GB in BF16. Each 32K-token conversation adds about 2.14 GiB of 16-bit cache and linear-attention state, or 1.14 GiB with an FP8 cache. One RTX PRO 6000 with the FP8 weights holds about 25 such conversations with a 16-bit cache and 47 with an FP8 cache, by our estimate.
How much VRAM does Qwen3.8-Flash-Next need?
vLLM’s recipe puts the FP8 checkpoint at 172.78 GiB, including a 51B n-gram embedding table, and states that 80 GB per GPU is not enough headroom for that table under four-way tensor parallelism. By our reading two H200 NVL hold it with about 46 conversations of 32K, while four RTX PRO 6000 hold it with about 16. With the table offloaded to host memory, four RTX PRO 6000 or two H200 NVL serve 100 conversations by our estimate.
Can I run Qwen locally on a DGX Spark?
Qwen3.8-27B fits one 128 GB DGX Spark in FP8 or BF16, with room for about 30 conversations of 32K in FP8 by our estimate. The Spark’s 273 GB/s of memory bandwidth limits speed for a dense model, so it suits a developer or a small pilot. Qwen3.8-Flash-Next does not fit one Spark.
What hardware does Qwen3-235B need?
Qwen3-235B-A22B-2507 is 236 GB in FP8 and needs two H200 NVL for its weights, or four in BF16 at 470 GB. Its 134 GB NVFP4 version fits two RTX PRO 6000, which run FP4 natively, while the H200 NVL runs NVFP4 only weight-only. Qwen states that a one-million-token context needs about 1,000 GB of GPU memory in total.
Can Qwen3.8 be used commercially on-premise?
Qwen3.8-27B is published under Apache 2.0. Qwen3.8-Flash-Next comes under the Qwen Community License 1.0, which requires a separate licence for companies running a Model as a Service or AI Work Assistant business, with an exception for internal use that gives no third party access. Whether a clause applies to a company is a legal assessment for its legal department.

Send us the Qwen variant, the precision you plan to run, the context length you will declare and your peak number of concurrent conversations. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna