Long-context LLM hardware: KV cache and VRAM for 128K to 1M token conversations
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- At 128K tokens the KV cache, not the weights, sets the memory: one 128K conversation takes 40 GiB of 16-bit cache on Llama 3.3 70B, 9.6 GiB on DeepSeek-V3.2, 4.5 GiB on gpt-oss-120b and about 0.8 GiB on DeepSeek-V4-Flash
- For full attention the cache per token is 2 × layers × KV heads × head dimension × bytes per value; sliding-window and linear-attention layers hold a fixed amount, and latent attention (MLA) stores one compressed latent per layer
- As of October 2026, DeepSeek-V4-Flash, GLM-5.3, Llama 4 Scout (10M) and the Qwen3.8 models with YaRN are stated for 1M tokens or more; one 1M conversation needs about 6.4 GiB of cache on V4-Flash, 64 GiB on Qwen3.8-27B and up to 98 GiB on GLM-5.3
- By our estimate one H200 NVL or two RTX PRO 6000 hold one 128K conversation of Llama 3.3 70B in FP8, ten need four H200 NVL or eight RTX PRO 6000, and one DGX Spark holds Qwen3.8-27B in FP8 with one 1M conversation
- Under tensor parallelism a latent-attention cache is copied to every card; offloading through vLLM, LMCache or NVIDIA Dynamo keeps blocks for reuse, and vLLM’s HiSparse moves part of GLM-5.3’s cache to CPU memory while decoding
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Long-context LLM hardware: the KV cache sets the memory
At 128K tokens and beyond, the KV cache decides how much GPU memory a model needs, more than the weights do, and it grows with every token of every conversation in flight. On Llama 3.3 70B one 128K conversation takes 40 GiB of 16-bit cache, more than half the model’s FP8 weights. On gpt-oss-120b it takes 4.5 GiB, and on DeepSeek-V4-Flash about 0.8 GiB. The difference comes from the attention design in each model’s config.json, so long-context LLM hardware is sized model by model.
Here 128K means 131,072 tokens and 1M means 1,048,576. Configurations are our estimates with the rule of our guide to how much VRAM an LLM needs, 90 per cent of the memory the driver reports less 3 GiB per card, and weights as in our LLM hardware requirements by model.
KV cache per token: formula and figures by model
For a model with full attention in every layer, the cache per token is 2 × layers × KV heads × head dimension × bytes per value, where the 2 counts keys and values. Llama 3.3 70B has 80 layers, 8 KV heads and a head dimension of 128, which gives 320 KiB per token in 16-bit. Models with multi-head latent attention (MLA) store one compressed latent per layer instead, of kv_lora_rank plus the RoPE dimension: 512 + 64 values for DeepSeek-V3.2 and GLM-5.3, 256 + 64 for Mistral Small 4. Layers that keep only a window or a fixed state do not grow with the context, so they stay out of the per-token figure.
| MODEL | CONTEXT, AS STATED | KV PER TOKEN, 16-BIT | 128K CONVERSATION | 1M CONVERSATION |
|---|---|---|---|---|
| Llama 3.3 70B | 128K | 320 KiB | 40 GiB | above stated limit |
| Llama 4 Scout | 10M | 192 KiB, upper | 24 GiB | 192 GiB |
| GLM-5.3 | 1M | 97.8 KiB, upper | 12.2 GiB | 97.8 GiB |
| DeepSeek-V3. | 163,840 | 76.5 KiB | 9.6 GiB | above stated limit |
| Qwen3.8-27B | 262,144, up to 1M | 64 KiB + state | 8.1 GiB | 64.1 GiB |
| gpt-oss-120b | 131,072 | 36 KiB | 4.5 GiB | above stated limit |
| Qwen3. | 262,144, up to 1M | 24 KiB + state | 3.1 GiB | 24.1 GiB |
| Mistral Small 4 | 256K | 22.5 KiB | 2.8 GiB | above stated limit |
| Deep | 1M | 6.4 KiB | 0.8 GiB | 6.4 GiB |
Our arithmetic from config.json files on Hugging Face, read 9 October 2026, Scout from the Transformers defaults; context from the model cards and, for GLM-5.3, Z.ai’s documentation. Qwen states 1,000,000 tokens, about 5 per cent below the 1M column. An FP8 cache roughly halves the attention figures.
How sliding-window, linear and latent attention shrink the cache
gpt-oss-120b has 36 layers with 8 KV heads of dimension 64: 18 attend over the whole context and 18 over a 128-token sliding window. vLLM’s design notes for its hybrid KV cache manager reserve slots “only for the most recent sliding_window_size tokens” in sliding-window layers, so only the 18 full layers count, 36 KiB per token.
Qwen3.8-27B keeps a full cache in 16 of its 64 layers, with 4 KV heads of dimension 256; the other 48 are Gated DeltaNet layers with a fixed state of about 0.14 GiB in 32-bit per conversation, by our estimate. Qwen3.8-Flash-Next has 12 full layers of 2 KV heads, plus a Qwen Sparse Attention indexer whose cache we found undocumented and have not counted. Both cards state “262,144 natively and extensible up to 1,000,000 tokens” with YaRN, which the 27B card advises only when long contexts are required, as static YaRN is “potentially impacting performance on shorter texts”.
DeepSeek-V3.2 and GLM-5.3 add a sparse-attention indexer to MLA, whose FP8 keys and scales we count at 132 bytes per token and layer. GLM-5.3 marks 57 of its 78 indexer layers “shared”; counting keys in all 78 gives an upper bound, 90.5 KiB if the shared ones keep none.
The configuration of DeepSeek-V4-Flash-0731, the official release, lists 43 layers, 20 compressed by a ratio of 4, 19 by 128 and four uncompressed, with one KV head of dimension 512 and a 128-token sliding window. vLLM’s blog of 24 April 2026 says some layers “use purely a sliding window for local information without compression”, by our reading the uncompressed ones, and that with a BF16 cache “DeepSeek V4 only has 9.62 GiB KV cache per sequence at 1M context”. The same arithmetic gives V4-Flash-0731 6.4 KiB per token and about 6.4 GiB at 1M, and at most 3.9 GiB with the FP8 cache that vLLM’s recipe sets.
Llama 4 Scout’s configuration file is gated, so we use the Transformers defaults for Scout: 48 layers, 8 KV heads of dimension 128 and an attention_chunk_size of 8,192. vLLM’s design notes give Llama 4 “3 local : 1 full” layers without stating how many tokens the local layers keep, so we count all 48 at full length; the 12 full layers alone take 48 KiB per token.
Which configuration holds 128K, 256K or 1M tokens
| MODEL, CONTEXT | CACHE PER CONV. | 1 CONVERSATION | 10 CONVERSATIONS |
|---|---|---|---|
| gpt-oss-120b, 128K | 4.5 GiB | one DGX Spark, RTX PRO 6000 or H200 NVL | two RTX PRO 6000 or one H200 NVL |
| Qwen3.8-27B FP8, 256K | 16.1 GiB | one DGX Spark, RTX PRO 6000 or H200 NVL | four RTX PRO 6000 or two H200 NVL |
| Qwen3.8-27B FP8, 1M | 64.1 GiB | one DGX Spark or H200 NVL, or two RTX PRO 6000 | eight H200 NVL as two copies of four |
| Llama 3.3 70B FP8, 128K | 40 GiB | one H200 NVL or two RTX PRO 6000 | four H200 NVL or eight RTX PRO 6000 |
| Deep | 6.4 GiB | four RTX PRO 6000 or two H200 NVL | eight RTX PRO 6000 as two copies of four, or four H200 NVL |
| Llama 4 Scout FP8, 1M | 192 GiB | four RTX PRO 6000 or four H200 NVL | more than eight H200 NVL |
| GLM-5.3 FP8, 128K | 12.2 GiB | eight H200 NVL | eight H200 NVL with DCP or HiSparse |
| GLM-5.3 FP8, 1M | 97.8 GiB | eight H200 NVL with DCP or HiSparse | more than eight H200 NVL without HiSparse |
Our estimates, not measurements, with a 16-bit cache and the upper counts for Scout and GLM-5.3, from 83.0 GiB per RTX PRO 6000, 123.4 GiB per H200 NVL and 102 GB per DGX Spark (128 GB). DeepSeek-V4-Flash and GLM-5.3 keep their full cache on every card under tensor parallelism; DCP is decode context parallelism.
One DGX Spark holds gpt-oss-120b with seven 128K conversations, and Qwen3.8-27B in FP8 with one 1M conversation by a margin of about 2 GiB; our guide to what fits in 128 GB on a DGX Spark explains the 102 GB working set.
On the H200 NVL, the DeepSeek-V4-Flash rows assume its FP4 experts run weight-only, which our hub article could not confirm from DeepSeek or vLLM documents. Llama 4 Scout at 1M needs two cards instead of four if its local layers keep only their chunk.
We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the document sizes and the context you need, with the number of conversations at peak, and we size the server.
Long conversations on several GPUs
Tensor parallelism splits a model’s KV heads across the cards: eight for Llama 3.3 70B, gpt-oss-120b and Llama 4 Scout, four for Qwen3.8-27B. vLLM’s blog of 7 August 2026 on decode context parallelism states that “once TP exceeds the number of KV heads, the cache starts duplicating across GPUs”, so ten 1M conversations of Qwen3.8-27B run as two copies on four H200 NVL each.
Latent attention has no heads to split. The same post states that under tensor parallelism “the latent KV cache is replicated in full across every TP rank”. SGLang’s guide to data-parallel attention describes the same duplication for DeepSeek’s MLA, while each data-parallel replica “Maintains its own KV cache (no duplication)”. In both layouts every card holds the whole cache of each conversation it serves, so one card’s spare memory limits the longest context; DeepSeek-V4-Flash, with one KV head, behaves the same way under tensor parallelism.
Decode context parallelism (DCP) splits one request’s cache across a tensor-parallel set by token position. vLLM sets it with --decode-context-parallel-size for MLA and GQA models, and release 0.30.0 lists “PCP+DCP on sparse-MLA models”; SGLang sets it with --dcp-size and shows it for DeepSeek-V3.1. Pipeline parallelism also divides the cache, as each card keeps only its own layers’ cache. A 1M conversation of GLM-5.3 on eight H200 NVL, with about 35 GiB per card left after the weights, needs DCP, HiSparse offloading or, with an FP8 cache, two pipeline stages; we found no recipe that runs GLM-5.3 or DeepSeek-V4 with DCP. With 123 GiB per card for weights and cache against 83 GiB, the H200 NVL holds long MLA conversations on fewer cards.
Prefix caching, FP8 KV cache and chunked prefill
The features below follow vLLM’s “latest” documentation, a developer preview, as read on 9 October 2026; the current release is 0.31.0 of 5 October. Automatic prefix caching “caches the KV cache of existing queries, so that a new query can directly reuse the KV cache” when the prefix matches, as with repeated questions about one long document or a multi-round conversation. It “only reduces the time of processing the queries (the prefilling phase)”.
An FP8 cache, set with --kv-cache-dtype fp8, “can significantly reduce its memory footprint” according to vLLM and roughly halves the attention figures in both tables. For DeepSeek-V3.2, vLLM’s post of 29 September 2025 describes an FP8 format of 656 bytes per token and layer, 57 per cent of the 1,152 bytes in BF16.
A long prompt is computed in a prefill before the first token. In vLLM V1 “chunked prefill is enabled by default whenever possible”: large prefills run in chunks batched with decode requests, so a long prompt holds up other users’ answers less, and max_num_batched_tokens trades inter-token latency against time to first token. We found no vendor figure for the first-token time of a 128K or 1M prompt on these cards, so test with your own prompt lengths. Our comparison of vLLM, SGLang, TensorRT-LLM and Ollama covers the engines.
KV cache offloading to CPU memory and NVMe
vLLM’s blog of 10 September 2026 says its tiered offloading “preserves evicted KV data across host memory, storage, and remote peers”, and on a hit “vLLM reloads the data from a lower tier” instead of recomputing. Its development documentation lists --kv-offloading-size, the CPU buffer in GiB, with the backends native and lmcache. NVIDIA Dynamo’s documentation for release 1.5.1 describes the native path: “vLLM copies sealed GPU KV blocks to pinned CPU memory.”
LMCache 0.5.5 of 12 September 2026 moves KV caches into CPU RAM, local SSD and remote back ends such as Redis/Valkey or S3-compatible storage; its PyPI page says it “reduces TTFT” and improves throughput for long-context, multi-turn and RAG workloads. Dynamo’s KV Block Manager (KVBM), in its development documentation, spans GPU memory, pinned host memory, remote memory and SSDs for vLLM and TensorRT-LLM, and NVIDIA writes that offloading “is most effective when KV Cache exceeds GPU memory and cache reuse outweighs the overhead of transferring data.”
These documents treat offloaded blocks as a cache for reuse by a returning session or a repeated document. By our reading, the conversation being generated keeps its cache in GPU memory, so this offloading saves prefill time rather than raising the longest context a configuration holds.
vLLM’s HiSparse for sparse attention goes further. vLLM’s blog of 8 September 2026 reports GLM-5.3 at its full 1M context on one node of eight H200 with Hybrid HiSparse, which moves the cache pages the indexer does not select to pinned CPU memory when GPU memory runs short. HiSparse came with release 0.30.0, 0.31.0 lists “HiSparse hardening”, and the post says it is “currently implemented only for NVIDIA GPUs”.
Our AI servers have ECC memory sized to the GPU pool and NVMe tiers for models and indexes. Write to us with the context lengths and how often prompts repeat, and we size memory and storage for the offloaded cache with the GPUs.
Long context or RAG
Long context puts the whole document set into every request, with its own cache and prefill per conversation, while retrieval-augmented generation puts in only the retrieved passages. Long context suits one large document or code base queried repeatedly, where prefix caching saves the repeated prefill. RAG suits a corpus larger than any context and documents with different access rights, which our guide to RAG on company data filters at retrieval.
What we supply
We supply the DGX Spark Founders Edition, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, and the H200 NVL with two-way and four-way NVLink bridges, as cards or in AI servers built to order. We size and source the GPUs, system memory and NVMe storage together, on one EU contract and invoice with manufacturer warranty, and check the rack, power and airflow before we quote. Our professional GPU range lists every card. Serving engines, RAG and MLOps are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
How much VRAM do I need for a 128K context?
How much memory does a 1M token context need?
How do I calculate the KV cache size per token?
Which GPU is suited to long-context LLM inference?
Does KV cache offloading allow longer contexts?
Does an FP8 KV cache reduce memory for long context?
Send us the models, the context length per conversation, the number of conversations at peak and how often prompts repeat. We reply within one business day with the memory budget, a configuration with GPUs, system memory and NVMe storage, and a written quote.
Talk to an expertWe reply within one business day