Kimi K2 hardware requirements: K2.6, K2-Thinking and K2-Instruct on H200 NVL and RTX PRO 6000
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Kimi-K2.6 and Kimi-K2-Thinking (1T parameters, 32B active, native INT4) are about 595 GB on Hugging Face; vLLM’s Kimi-K2-Thinking recipe reports KV cache for about 21 conversations of 32K on eight H200 with tensor parallelism, and our rule gives eight H200 NVL about 25
- Decode context parallelism splits each conversation’s cache across the eight cards: vLLM’s recipe reports 5,721,088 tokens of cache on eight H200 with it, about 174 conversations of 32K, so 100 users fit one eight-card server
- Kimi K2 uses DeepSeek-V3’s multi-head latent attention: 68.6 KiB of 16-bit cache per token, 2.14 GiB per 32K conversation and 17.2 GiB at the full 262,144 tokens, which vLLM keeps in full on every card under tensor parallelism
- Eight RTX PRO 6000 hold the INT4 weights with at most about 13.8 GiB per card to spare, six conversations at 32K or about 51 with decode context parallelism by our rule; we found no Moonshot AI or vLLM document that runs Kimi K2 on this card, and four DGX Spark are too small for the published checkpoints
- Kimi-K2-Instruct-0905 in block FP8 is 1,029 GB, which on eight H200 NVL leaves about 20 GiB per card of what the driver reports but 3.6 GiB within our rule, room for one conversation of 32K, and Moonshot AI names 16 H200 or H20 GPUs as its smallest unit at 128K; the Kimi K2 weights are published under a Modified MIT License with a display clause for very large services
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Kimi K2 hardware requirements in October 2026
The hardware requirements of the current Kimi K2 models start at one server with eight H200 NVL cards. Kimi-K2.6 and Kimi-K2-Thinking ship with native INT4 weights of about 595 GB. vLLM’s Kimi-K2-Thinking recipe, run on eight H200 with 141 GB each, reports KV cache for about 21 conversations of 32,768 tokens under tensor parallelism and about 174 with decode context parallelism, so 20 and 100 concurrent users both fit. Our own rule gives eight H200 NVL about 25 and 200; we plan with the recipe’s lower figures.
Eight RTX PRO 6000 also hold the INT4 weights, but no Moonshot AI or vLLM document we found runs Kimi K2 on that card. The older Kimi-K2-Instruct-0905 in FP8 is 1,029 GB, which leaves eight H200 NVL room for one conversation at most. Four DGX Spark (128 GB) are too small for the checkpoints that Moonshot AI and NVIDIA publish.
These are our estimates for hardware we supply: the DGX Spark Founders Edition (128 GB of unified memory), the RTX PRO 6000 (96 GB, PCIe) and the H200 NVL (141 GB, NVLink bridges for two or four cards). Our LLM hardware requirements by model applies the same method to other model families.
Kimi K2 models on Hugging Face: K2.6, K2-Thinking and K2-Instruct
As of 9 October 2026, moonshotai’s Hugging Face organisation lists Kimi-K2-Instruct with its 0905 update, Kimi-K2-Thinking, Kimi-K2.5, Kimi-K2.6 and Kimi-K2.7-Code. All have 1T total and 32B activated parameters, 61 layers and 384 experts, of which 8 are selected per token. The K2.6 card states that it “has the same architecture as Kimi-K2.5”, adds the 400M-parameter MoonViT vision encoder and gives a context length of 256K. The K2.7-Code card describes “a coding-focused agentic model built upon Kimi K2.6” with the same architecture, so the K2.6 figures below apply to it by our reading.
| MODEL | PARAMETERS | WEIGHTS | CHECKPOINT | CONTEXT |
|---|---|---|---|---|
| Kimi-K2.6 | 1T, 32B active | INT4 experts, BF16 rest | 595.2 GB | 256K |
| Kimi-K2-Thinking | 1T, 32B active | INT4 MoE weights (QAT) | 594.2 GB | 256K |
| Kimi-K2-Instruct-0905 | 1T, 32B active | block FP8 | 1,029.2 GB | 256K |
| Kimi-K2.6, NVIDIA NVFP4 | as K2.6 | NVFP4 MoE linear layers | 595.2 GB | 256K |
Model cards, file lists (sum of the safetensors shards) and config.json files on Hugging Face, read on 9 October 2026; 256K is 262,144 tokens in K2.6’s config.json. K2.5 and K2.7-Code share the K2.6 architecture; we did not sum their files.
The K2-Thinking card explains the INT4 weights: “we adopt Quantization-Aware Training (QAT) during the post-training phase, applying INT4 weight-only quantization to the MoE components”. K2.6 “adopts the same native int4 quantization method”, and its config.json quantises the linear layers to 4-bit integers in groups of 32 and keeps attention, the shared expert, the dense first layer and the output head in 16-bit. The 0905 card states that its checkpoints “are stored in the block-fp8 format”, and that its context window “has been increased from 128k to 256k tokens”.
NVIDIA’s NVFP4 version of K2.6 quantises “only the weights and activations of the linear operators within transformer blocks in MoE”. Its card lists NVIDIA Blackwell, names the B200 as test hardware and starts vLLM with --tensor-parallel-size 4. Moonshot AI’s newer Kimi-K3, with 2.8T parameters and 104B active, MXFP4 weights and a 1M context, is a different architecture under its own Kimi K3 License and outside this article.
KV cache per conversation: latent attention in Kimi K2
Our sizing rule, explained in our guide to how much VRAM an LLM needs, gives the weights and the cache 90 per cent of the memory the driver reports, less 3 GiB per card: 83.0 GiB per RTX PRO 6000 and 123.4 GiB per H200 NVL. For a DGX Spark (128 GB) we take 102 GB per system. The cache is counted in 16-bit, vLLM’s default, because vLLM’s Kimi-K2-Thinking commands for the H200 set no cache type.
Kimi K2 uses the multi-head latent attention (MLA) of DeepSeek-V3: the text section of its config.json names the DeepseekV3ForCausalLM architecture with a kv_lora_rank of 512 and a qk_rope_head_dim of 64. Each of the 61 layers stores 576 values per token, which makes 68.6 KiB per token in 16-bit, 2.14 GiB per 32K conversation, 8.6 GiB at 128K and 17.2 GiB at the full 262,144 tokens. An FP8 cache, which vLLM’s K2.6 recipe sets on its B300 path, brings these figures down to about half.
vLLM’s blog post of 7 August 2026 on decode context parallelism states: “Under normal TP there is nothing to split by head, meaning the latent KV cache is replicated in full across every TP rank.” Under tensor parallelism we therefore compare one conversation’s cache with one card’s spare memory. Three layouts divide the cache instead. Pipeline parallelism (PP) gives each card only its own layers’ cache. Decode context parallelism (DCP), set with --decode-context-parallel-size up to the tensor-parallel size, “is able to split KV cache across GPUs by sequence (context) dimension”. With data-parallel attention (DP), “Each DP engine has an independent KV cache” in vLLM’s documentation, but every card repeats the attention and other non-expert weights; we have not sized DP for Kimi K2. Our long-context LLM hardware article compares them with other models.
GPUs for 1, 20 and 100 concurrent users
| MODEL, FORMAT | DGX SPARK | RTX PRO 6000 | H200 NVL |
|---|---|---|---|
| K2.6 or K2-Thinking, INT4 | no / no / no | 8 / 8 with DCP / over 8 | 8 / 8 / 8 with DCP |
| K2.6, NVIDIA NVFP4 | no / no / no | 8 / 8 with DCP / over 8 | use INT4 |
| K2-Instruct-0905, FP8 | no / no / no | over 8 / over 8 / over 8 | 8 / over 8 / over 8 |
Our estimates, not measurements, for 1, 20 and 100 conversations of 32,768 tokens with a 16-bit cache, on 1, 2, 4 or 8 cards or up to four DGX Spark (128 GB), memory fit only; tensor parallelism over all cards, or decode context parallelism (DCP) where the cell says so. No Moonshot AI or vLLM document runs Kimi K2 on the RTX PRO 6000; FP8 on eight H200 NVL is tight, see the text.
The INT4 weights take 554.3 GiB, or 69.3 GiB per card on eight cards. Four H200 NVL provide 493 GiB by our rule and four RTX PRO 6000 332 GiB, so both card types need eight. Eight RTX PRO 6000 leave about 13.8 GiB per card, room for six conversations at 32K under tensor parallelism, 12 with two pipeline stages and about 51 with -dcp 8, so a hundred users need at least two such servers with DCP. These RTX figures are upper bounds, because on eight H200 vLLM’s recipe leaves about 7.4 GiB per card less cache than our rule, and the same shortfall would halve the spare memory of an RTX PRO 6000. vLLM’s DCP blog also names “extending support to a wider variety of backends” as further work, so test DCP on this card before you size for it. Eight H200 NVL leave 54.1 GiB per card by our rule, about 25 conversations at 32K with tensor parallelism, 50 with two pipeline stages and about 200 with DCP.
Kimi-K2-Instruct-0905 takes 958.5 GiB in FP8, 119.8 GiB per card on eight H200 NVL. That is about 20.6 GiB below the 140.4 GiB the driver reports, but within our rule it leaves 3.6 GiB per card, one conversation at 32K, or about 13 with DCP. With the 7.4 GiB shortfall seen in the K2-Thinking recipe, no cache would remain at vLLM’s default memory setting. Moonshot AI’s deployment guide states that “The smallest deployment unit for Kimi-K2 FP8 weights with 128k seqlen on mainstream H200 or H20 platform is a cluster with 16 GPUs”. For a team on the FP8 model, two eight-card servers are the minimum; K2.6 in INT4 serves the same numbers on one.
We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Tell us which Kimi K2 variant you plan to run, with the context and the users at peak, and we reply with a configuration and quote.
Kimi K2 on eight H200 NVL: what the vLLM recipes set
vLLM’s Kimi-K2-Thinking recipe, updated 17 April 2026, names “8x H200 or 8x H20 GPUs” as hardware. Its low-latency command uses --tensor-parallel-size 8, and its high-throughput command adds --decode-context-parallel-size 8. The recipe states that “DCP multiplies the GPU KV cache size by dcp_world_size” and reports, from its benchmark on eight H200, a cache of 715,072 tokens with tensor parallelism and 5,721,088 tokens with DCP. That is about 21 and 174 conversations of 32K, from 46.8 GiB of cache per card, about 7.4 GiB less than our rule leaves. By our reading the difference is runtime memory that vLLM measures at start-up and our flat 3 GiB reserve does not cover, so we plan with the recipe’s figures where it gives them. At the full 262,144 tokens, they mean two conversations with tensor parallelism and 21 with DCP.
The K2.6 recipe, updated 29 July 2026, lists “8× H200 GPUs (verified), or equivalent aggregate VRAM (~640 GB)” for the INT4 checkpoint. Its CPU offload option extends the prefix cache into host DRAM with SimpleCPUOffloadConnector and “uses 220 GiB per rank by default”. By our reading it is a cache for reuse, which saves prefill time but does not raise the longest context, and eight ranks at the default would take 1,760 GiB of server memory by our arithmetic.
The recipes do not state the form factor of their H200. The H200 NVL has the same 141 GB, and in an eight-card server its bridges form two NVLink domains of four, joined over PCIe, as our guide to H200 NVL card counts for large models explains. Tensor parallelism over eight cards then crosses PCIe between the domains. Four-way tensor parallelism in each domain with two pipeline stages keeps the per-layer traffic on NVLink and halves the cache per card, which doubles the recipe’s tensor-parallel figure to about 43 conversations at 32K by our estimate. Eight H200 NVL at up to 600 W each draw 4.8 kW before processors, memory and fans.
We build servers with eight H200 NVL and their NVLink bridges and check the rack, power and airflow before we quote. Describe your rack position and its power feed in the form below.
Kimi K2 on RTX PRO 6000 and DGX Spark
Eight RTX PRO 6000 have 768 GB, more than the “~640 GB” of aggregate memory the K2.6 recipe names. The INT4 checkpoint uses the compressed-tensors “pack-quantized” format, so test the engine and its kernels on the RTX PRO 6000 before you size a server on it. The RTX PRO 6000 computes FP4 on its Blackwell Tensor Cores, which suits NVIDIA’s NVFP4 version, but NVIDIA tested that version on the B200, a different Blackwell GPU.
The cards connect over PCIe only, so tensor parallelism over eight sends every layer’s exchange across PCIe. Our article on tensor, pipeline and expert parallelism over PCIe and NVLink covers the link and the peer-to-peer settings. At 262,144 tokens a conversation needs 17.2 GiB on every card under tensor parallelism, more than the 13.8 GiB left, so a full-length conversation on eight RTX PRO 6000 needs two pipeline stages or DCP.
The DGX Spark does not hold the published Kimi K2 checkpoints. Four of them provide 408 GB by our rule, against about 595 GB of INT4 weights. The K2.6 card also lists the KTransformers engine, and community quantisations with fewer bits per weight are smaller; we have not sized either.
Modified MIT License and data handling
Moonshot AI releases the code and the weights of the Kimi K2 models in this article under a Modified MIT License. Its only modification applies when the software or a derivative is used in commercial products or services “that have more than 100 million monthly active users”, or above a monthly revenue threshold stated in the licence. Such products must “prominently display” the model name, “Kimi K2.6” in the K2.6 licence, on their user interface. In all other respects the MIT terms apply. NVIDIA’s NVFP4 version is governed by the NVIDIA Open Model License, with the Modified MIT License named as additional information. A model run from downloaded weights on your own servers sends no prompts or outputs to Moonshot AI. Whether a clause applies to your company is a legal assessment for your legal department.
What we supply
We supply the H200 NVL with its two-way and four-way NVLink bridges, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, and the DGX Spark Founders Edition. They come as cards for a server you already run or in AI servers built to order, on one EU contract and invoice with manufacturer warranty. We return a configuration and a quote within one business day, and our professional GPU range lists every card. Deploying Kimi K2 with vLLM on these servers is part of our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What are the hardware requirements for Kimi K2?
How much VRAM does Kimi K2 need?
Can I run Kimi K2 locally?
What hardware does Kimi K2 Thinking need?
How many GPUs does Kimi K2.6 need for 100 users?
Can Kimi K2 be used commercially on premise?
Send us the Kimi K2 variant, the context length you will configure and the number of conversations in flight at peak. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day