DeepSeek hardware requirements: V4-Flash, V4.1-Flash and V3.2 on RTX PRO 6000 and H200 NVL
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- DeepSeek-V4-Flash (284B parameters, 13B active, FP4 experts and FP8 for the rest) is 160 GB, its official -0731 release 167 GB; by our estimate two RTX PRO 6000 hold it for 20 conversations at 32K with an FP8 cache and four for 100, while two H200 NVL hold it only if the engine runs the FP4 experts weight-only
- vLLM’s DeepSeek-V4-Flash recipe, updated 29 September 2026, states that serving the -0731 checkpoint without speculative decoding was verified on eight RTX PRO 6000 over PCIe, and runs it on two DGX Spark because one Spark’s 128 GB is too small
- DeepSeek-V3.2 (685B with MTP, 37B active) is 690 GB in FP8; eight H200 NVL hold it with cache for about 88 conversations at 32K in vLLM’s data-parallel mode by our estimate, while eight RTX PRO 6000 leave room for about one
- DeepSeek-V4.1-Flash, released on 10 September 2026, is about 511 GB including 196.6B parameters of Engram tables; vLLM runs it on four H200 with those tables in host memory, and eight RTX PRO 6000 hold it by our memory estimate
- vLLM states that tensor parallelism keeps the full latent KV cache on every card, so a 32K conversation takes about 0.12 GiB of FP8 cache per card on V4-Flash and 2.4 GiB of 16-bit cache on V3.2; all these checkpoints are published under the MIT licence
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
DeepSeek hardware requirements in October 2026
DeepSeek’s hardware requirements depend on the variant, and the smallest server configuration for DeepSeek-V4-Flash is two RTX PRO 6000 cards, which hold its 160 GB checkpoint, or 167 GB for the official -0731 release, with room for 20 conversations of 32,768 tokens in the FP8 cache that vLLM’s recipe sets. A hundred such conversations need four RTX PRO 6000, or two H200 NVL if the serving engine keeps the FP4 expert weights in FP4. DeepSeek-V3.2 needs eight H200 NVL for its 690 GB of FP8 weights, with cache for about 88 conversations at 32K with data-parallel attention, so 100 such users need more than eight cards in that mode. DeepSeek-V4.1-Flash takes four H200 NVL or, by memory, eight RTX PRO 6000.
These are our estimates for hardware we supply: the DGX Spark Founders Edition with 128 GB of unified memory, the RTX PRO 6000 with 96 GB on PCIe, and the H200 NVL with 141 GB and NVLink bridges for two or four cards. Our LLM hardware requirements by model places DeepSeek beside other model families.
DeepSeek models on Hugging Face in October 2026
DeepSeek’s API change log dates DeepSeek-V3.2 to 1 December 2025, DeepSeek-V4 to 24 April 2026, the DeepSeek-V4-Flash update to 31 July 2026 and DeepSeek-V4.1-Flash to 10 September 2026. The model card of the July update calls it “the official release of DeepSeek-V4-Flash, superseding the preview version”.
| MODEL | PARAMETERS | WEIGHTS | CHECKPOINT | CONTEXT |
|---|---|---|---|---|
| Deep | 284B, 13B active | FP4 experts, FP8 rest | 160 GB | 1M |
| Deep | 304B by NVIDIA’s count, 13B active | FP4 experts, FP8 rest | 167 GB | 1M |
| DeepSeek-V4. | 552B plus 196.6B Engram; 8B or 16B active | FP4 experts, FP8 rest | about 511 GB | 1M |
| Deep | 1.6T, 49B active | FP4 experts, FP8 rest | 893 GB | 1M |
| DeepSeek-V3. | 685B with MTP, 37B active | FP8 | 690 GB | 163,840 |
| DeepSeek-V3. | as V3.2 | NVFP4 linear layers | 415 GB | as V3.2 |
Model cards, file lists and config.json files on Hugging Face and the vLLM recipes for V4-Flash and V4.1-Flash, read on 9 October 2026; V4-Flash-0731 parameters from NVIDIA’s NVFP4 card, V4-Pro parameters from DeepSeek’s V4 card, V3.2’s active parameters from DeepSeek’s V3 card.
The 7 GB between the preview and the -0731 checkpoint is a draft module for speculative decoding, which vLLM’s recipe names as “the difference”. For V4.1-Flash, DeepSeek’s release note gives “8B active parameters for input, 16B for output”, and vLLM’s recipe adds that the two Engram tables “alone are 196.6B parameters (~183 GiB)”.
NVIDIA’s NVFP4 version of V3.2 quantises only the linear operators of the transformer blocks and reduces disk size and GPU memory “by approximately 1.66x”. NVIDIA’s NVFP4 version of V4-Flash-0731 is “slightly larger than its source”, so it saves no memory.
KV cache per conversation: MLA and compressed attention
Our sizing rule, explained in our guide to how much VRAM an LLM needs, gives the weights and the cache 90 per cent of the memory the driver reports, less 3 GiB per card: 83.0 GiB per RTX PRO 6000 and 123.4 GiB per H200 NVL. For a DGX Spark (128 GB) we take 102 GB per system. The cache is counted from each config.json at 32,768 tokens per conversation, in 16-bit for V3.2 and in FP8 for the V4 models, whose vLLM recipe sets --kv-cache-dtype fp8 in its RTX PRO 6000 command.
DeepSeek-V3.2 uses multi-head latent attention (MLA). Its config.json sets kv_lora_rank to 512 and qk_rope_head_dim to 64, so each of its 61 layers stores 576 values per token instead of keys and values for all 128 heads. DeepSeek Sparse Attention adds an indexer with keys of 128 dimensions per layer, which we count in FP8 with a scale. That makes about 76.5 KiB per token and 2.39 GiB per 32K conversation. By our reading, the sparse attention reduces computation while the cache still holds every token.
DeepSeek-V4-Flash combines “Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA)”, with one KV head of 512 dimensions. The config.json of the -0731 release compresses 20 layers by a ratio of 4 and 19 by 128 and leaves four uncompressed, which by our reading of vLLM’s blog of 24 April 2026 “use purely a sliding window for local information without compression” and keep only 128 tokens. By our count that is 6.4 KiB per token in 16-bit and at most 3.9 KiB in FP8, about 0.12 GiB per 32K conversation and 1.5 GiB at 384K. For V4.1-Flash, vLLM’s recipe cites DeepSeek’s figure for the global KV of “890 bytes per token”, about 28 MiB per 32K conversation.
vLLM’s blog post of 7 August 2026 states for MLA models: “Under normal TP there is nothing to split by head, meaning the latent KV cache is replicated in full across every TP rank.” For GQA models the cache “begins duplicating” once the tensor-parallel size exceeds the number of KV heads; V4-Flash has one, so by our reading every card holds its whole cache too. Under tensor parallelism we therefore compare the cache per conversation with the spare memory of one card. With data-parallel attention, which vLLM’s documentation calls advantageous for MoE models with MLA, “Each DP engine has an independent KV cache”, but every card also holds the non-expert weights in full.
Cards for 1, 20 and 100 concurrent users
| MODEL, FORMAT | DGX SPARK | RTX PRO 6000 | H200 NVL |
|---|---|---|---|
| V4-Flash or -0731 | 2 / 2 / 2 | 2 / 2 / 4 | 2 / 2 / 2, weight-only |
| V4.1-Flash | no / no / no | 8 / 8 / 8 | 4 / 4 / 4 |
| V3.2, FP8 | no / no / no | 8, tight / over 8 / over 8 | 8 / 8 / over 8 |
| V3.2, NVIDIA NVFP4 | no / no / no | 8 / over 8 / over 8 | use FP8 |
Our estimates, not measurements, for 1, 20 and 100 conversations of 32,768 tokens on 1, 2, 4 or 8 cards or up to four DGX Spark (memory fit only), in vLLM with tensor parallelism and the full cache on every card; V3.2 on the H200 NVL in vLLM’s -dp 8 --enable-expert-parallel. Cache in FP8 for V4 models, as vLLM’s recipe sets it, in 16-bit for V3.2. H200 NVL figures for V4 models assume FP4 experts run weight-only.
Two RTX PRO 6000 leave 5.3 GiB per card after the -0731 weights, room for about 43 conversations at 32K, and 8.5 GiB after the preview, about 66. A hundred conversations need 12.3 GiB on each card, so 100 users need four cards.
V3.2 leaves 2.7 GiB per card on eight RTX PRO 6000, about one conversation at 32K. Eight H200 NVL in tensor parallelism leave 43 GiB per card, about 18 conversations. In vLLM’s data-parallel mode each card holds about 18.7 GiB of non-expert weights by our count from config.json, plus an eighth of the experts, which leaves about 26.7 GiB per card, room for 11 conversations each and 88 across eight cards. vLLM’s post of 29 September 2025 describes an FP8 cache for this model of 656 bytes per token and layer, which would allow about 144, but the V3.2 recipe sets no cache type. Decode context parallelism (-dcp 8), listed for sparse-MLA models in vLLM 0.30.0, splits each cache across the cards by token, for about 144 on eight H200 NVL and nine on eight RTX PRO 6000; no recipe we found runs V3.2 that way.
We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Tell us which DeepSeek variant you plan to run, at what context and for how many users at peak, and we reply with a configuration and quote.
DeepSeek-V4-Flash on RTX PRO 6000, H200 NVL and two DGX Spark
The RTX PRO 6000 computes FP4 on its Blackwell Tensor Cores, so DeepSeek-V4-Flash runs on it as released. vLLM’s recipe has a profile headed “RTX PRO 6000 8× (8×96 GB, sm_120)” for the -0731 checkpoint and states that “Serving without speculative decoding was verified on 8× RTX PRO 6000 (PCIe, no NVLink).” Two cards are our memory estimate, not a layout the recipe verifies. The recipe says Think Max “requires --max-model-len >= 393216 (384K tokens) to avoid truncation”. At 384K, a conversation needs about 1.5 GiB of FP8 cache on every card. Two RTX PRO 6000 hold about three such conversations next to the -0731 weights, four hold about 29 and eight about 43.
The H200 NVL has FP8 but no FP4 arithmetic, as our guide to H200 NVL card counts for large models explains. vLLM’s recipe lists H200 layouts and recommends data parallelism of four with expert parallelism, which “uses 4 of 8 GPUs per replica on H200/B200/B300”. It cites reports that “establish Hopper compatibility” but does not say how Hopper computes the FP4 experts. Two H200 NVL hold the weights with about 46 GiB per card to spare only if the experts stay in FP4.
The recipe states that “A single GB10’s 128 GB unified memory is below the FP8/NVFP4 checkpoint footprint” and runs two Spark systems with tensor parallelism of two. By our rule, two DGX Spark leave about 17 GiB per system for the cache, which each system holds in full.
DeepSeek-V3.2 and V4.1-Flash on eight or four cards
DeepSeek-V3.2 in FP8 takes 642.6 GiB, more than four cards of either type can hold. Eight H200 NVL form two NVLink domains of four, joined over PCIe. We found no vLLM or DeepSeek document that runs V3.2 on the RTX PRO 6000, and its FP8 weights use 128 × 128 block scales, so test the engine on that card first.
NVIDIA’s NVFP4 version of V3.2, at 415 GB, fits eight RTX PRO 6000 with about 34.7 GiB per card to spare, about 14 conversations at 32K under tensor parallelism. NVIDIA states that “you need 8xB200 GPU and TensorRT LLM version 1.2.0rc8 or above”, so on the RTX PRO 6000 it is an untested path.
DeepSeek-V4.1-Flash weighs about 476 GiB. vLLM’s recipe, updated 3 October 2026, states that “On H200, Tensor Parallel uses the TP4 default with Engram CPU offload”, and that the tables “move to pinned host DRAM and are read through UVA”. The offload moves 23.6 GiB per GPU off the card, leaving “81.2 GiB of resident weights and 38.5 GiB of KV per GPU”, so four H200 NVL on a four-way bridge hold the cache for 100 conversations at 32K. The recipe gives no host memory figure for the H200 layout and no RTX PRO 6000 profile; eight of those cards hold the weights with about 23.5 GiB per card to spare by our rule.
DeepSeek-V4-Pro-0813, at 893 GB, exceeds eight RTX PRO 6000 by our rule, and eight H200 NVL hold its weights only with the experts kept in FP4.
We build servers with four or eight H200 NVL and their NVLink bridges and we check the rack, power and airflow before we quote. Describe your rack position and its power feed in the form below.
vLLM settings for DeepSeek over PCIe and NVLink
vLLM’s DeepSeek-V3.2 recipe, updated 24 September 2026, starts with --tensor-parallel-size 8 and then advises to “avoid using -tp=8 for DeepSeek-V3.2 with FlashMLA-Sparse”, because “TP=8 yields only 16 heads (128/8) per rank but is padded to 64 heads, incurring overhead and hurting performance.” It recommends “TP=1~2 + DP/EP”, two on Hopper, and calls -dp 8 --enable-expert-parallel “the recommended serving mode”. On eight H200 NVL, expert traffic between the two NVLink domains crosses PCIe.
For V4-Flash on the RTX PRO 6000, the recipe sets --tensor-parallel-size 8 with --enable-expert-parallel, --kv-cache-dtype fp8 and --block-size 256, and keeps the indexer cache in FP8 because “sm_120 cannot use the SM100 FP4 indexer cache or deep_gemm_mega_moe.” Speculative decoding with the -0731 draft module needs a nightly vLLM image on that card. Expert parallelism sends tokens to the cards that hold their experts, over PCIe on the RTX PRO 6000; our article on tensor, pipeline and expert parallelism over PCIe and NVLink covers the link and the peer-to-peer settings.
MIT licence and data handling
The DeepSeek checkpoints in this article are published under the MIT licence, and each card states: “This repository and the model weights are licensed under the MIT License.” The licence permits use, modification and distribution on condition that the copyright and permission notice are included in copies, and it excludes any warranty. A model run from downloaded weights on your own servers sends no prompts or outputs to DeepSeek. Whether further terms apply to your use is a legal assessment for your legal department.
Our article on a private ChatGPT server by company size relates these configurations to staff numbers.
What we supply
We supply the DGX Spark Founders Edition, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, and the H200 NVL with its two-way and four-way NVLink bridges. They come as cards for a server you already run or in AI servers built to order, on one EU contract and invoice with manufacturer warranty. We return a configuration and a quote within one business day, and our professional GPU range lists every card. Deploying DeepSeek with vLLM on these servers is part of our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What are the hardware requirements for DeepSeek?
How much VRAM does DeepSeek-V4-Flash need?
How much VRAM does DeepSeek 671B need?
Can I run DeepSeek locally on a DGX Spark?
Does DeepSeek-V4-Flash run on the H200 NVL?
Can DeepSeek be used commercially on premise?
Send us the DeepSeek variant, the context length you will configure and the number of conversations in flight at peak. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day