BLOG · GUIDE ·

Long-context LLM hardware: KV cache and VRAM for 128K to 1M token conversations

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • At 128K tokens the KV cache, not the weights, sets the memory: one 128K conversation takes 40 GiB of 16-bit cache on Llama 3.3 70B, 9.6 GiB on DeepSeek-V3.2, 4.5 GiB on gpt-oss-120b and about 0.8 GiB on DeepSeek-V4-Flash
  • For full attention the cache per token is 2 × layers × KV heads × head dimension × bytes per value; sliding-window and linear-attention layers hold a fixed amount, and latent attention (MLA) stores one compressed latent per layer
  • As of October 2026, DeepSeek-V4-Flash, GLM-5.3, Llama 4 Scout (10M) and the Qwen3.8 models with YaRN are stated for 1M tokens or more; one 1M conversation needs about 6.4 GiB of cache on V4-Flash, 64 GiB on Qwen3.8-27B and up to 98 GiB on GLM-5.3
  • By our estimate one H200 NVL or two RTX PRO 6000 hold one 128K conversation of Llama 3.3 70B in FP8, ten need four H200 NVL or eight RTX PRO 6000, and one DGX Spark holds Qwen3.8-27B in FP8 with one 1M conversation
  • Under tensor parallelism a latent-attention cache is copied to every card; offloading through vLLM, LMCache or NVIDIA Dynamo keeps blocks for reuse, and vLLM’s HiSparse moves part of GLM-5.3’s cache to CPU memory while decoding

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Long-context LLM hardware: the KV cache sets the memory

At 128K tokens and beyond, the KV cache decides how much GPU memory a model needs, more than the weights do, and it grows with every token of every conversation in flight. On Llama 3.3 70B one 128K conversation takes 40 GiB of 16-bit cache, more than half the model’s FP8 weights. On gpt-oss-120b it takes 4.5 GiB, and on DeepSeek-V4-Flash about 0.8 GiB. The difference comes from the attention design in each model’s config.json, so long-context LLM hardware is sized model by model.

Here 128K means 131,072 tokens and 1M means 1,048,576. Configurations are our estimates with the rule of our guide to how much VRAM an LLM needs, 90 per cent of the memory the driver reports less 3 GiB per card, and weights as in our LLM hardware requirements by model.

KV cache per token: formula and figures by model

For a model with full attention in every layer, the cache per token is 2 × layers × KV heads × head dimension × bytes per value, where the 2 counts keys and values. Llama 3.3 70B has 80 layers, 8 KV heads and a head dimension of 128, which gives 320 KiB per token in 16-bit. Models with multi-head latent attention (MLA) store one compressed latent per layer instead, of kv_lora_rank plus the RoPE dimension: 512 + 64 values for DeepSeek-V3.2 and GLM-5.3, 256 + 64 for Mistral Small 4. Layers that keep only a window or a fixed state do not grow with the context, so they stay out of the per-token figure.

MODELCONTEXT, AS STATEDKV PER TOKEN, 16-BIT128K CONVERSATION1M CONVERSATION
Llama 3.3 70B128K320 KiB40 GiBabove stated limit
Llama 4 Scout10M192 KiB, upper24 GiB192 GiB
GLM-5.31M97.8 KiB, upper12.2 GiB97.8 GiB
DeepSeek-V3.2163,84076.5 KiB9.6 GiBabove stated limit
Qwen3.8-27B262,144, up to 1M64 KiB + state8.1 GiB64.1 GiB
gpt-oss-120b131,07236 KiB4.5 GiBabove stated limit
Qwen3.8-Flash-Next262,144, up to 1M24 KiB + state3.1 GiB24.1 GiB
Mistral Small 4256K22.5 KiB2.8 GiBabove stated limit
DeepSeek-V4-Flash-07311M6.4 KiB0.8 GiB6.4 GiB

Our arithmetic from config.json files on Hugging Face, read 9 October 2026, Scout from the Transformers defaults; context from the model cards and, for GLM-5.3, Z.ai’s documentation. Qwen states 1,000,000 tokens, about 5 per cent below the 1M column. An FP8 cache roughly halves the attention figures.

How sliding-window, linear and latent attention shrink the cache

gpt-oss-120b has 36 layers with 8 KV heads of dimension 64: 18 attend over the whole context and 18 over a 128-token sliding window. vLLM’s design notes for its hybrid KV cache manager reserve slots “only for the most recent sliding_window_size tokens” in sliding-window layers, so only the 18 full layers count, 36 KiB per token.

Qwen3.8-27B keeps a full cache in 16 of its 64 layers, with 4 KV heads of dimension 256; the other 48 are Gated DeltaNet layers with a fixed state of about 0.14 GiB in 32-bit per conversation, by our estimate. Qwen3.8-Flash-Next has 12 full layers of 2 KV heads, plus a Qwen Sparse Attention indexer whose cache we found undocumented and have not counted. Both cards state “262,144 natively and extensible up to 1,000,000 tokens” with YaRN, which the 27B card advises only when long contexts are required, as static YaRN is “potentially impacting performance on shorter texts”.

DeepSeek-V3.2 and GLM-5.3 add a sparse-attention indexer to MLA, whose FP8 keys and scales we count at 132 bytes per token and layer. GLM-5.3 marks 57 of its 78 indexer layers “shared”; counting keys in all 78 gives an upper bound, 90.5 KiB if the shared ones keep none.

The configuration of DeepSeek-V4-Flash-0731, the official release, lists 43 layers, 20 compressed by a ratio of 4, 19 by 128 and four uncompressed, with one KV head of dimension 512 and a 128-token sliding window. vLLM’s blog of 24 April 2026 says some layers “use purely a sliding window for local information without compression”, by our reading the uncompressed ones, and that with a BF16 cache “DeepSeek V4 only has 9.62 GiB KV cache per sequence at 1M context”. The same arithmetic gives V4-Flash-0731 6.4 KiB per token and about 6.4 GiB at 1M, and at most 3.9 GiB with the FP8 cache that vLLM’s recipe sets.

Llama 4 Scout’s configuration file is gated, so we use the Transformers defaults for Scout: 48 layers, 8 KV heads of dimension 128 and an attention_chunk_size of 8,192. vLLM’s design notes give Llama 4 “3 local : 1 full” layers without stating how many tokens the local layers keep, so we count all 48 at full length; the 12 full layers alone take 48 KiB per token.

Which configuration holds 128K, 256K or 1M tokens

MODEL, CONTEXTCACHE PER CONV.1 CONVERSATION10 CONVERSATIONS
gpt-oss-120b, 128K4.5 GiBone DGX Spark, RTX PRO 6000 or H200 NVLtwo RTX PRO 6000 or one H200 NVL
Qwen3.8-27B FP8, 256K16.1 GiBone DGX Spark, RTX PRO 6000 or H200 NVLfour RTX PRO 6000 or two H200 NVL
Qwen3.8-27B FP8, 1M64.1 GiBone DGX Spark or H200 NVL, or two RTX PRO 6000eight H200 NVL as two copies of four
Llama 3.3 70B FP8, 128K40 GiBone H200 NVL or two RTX PRO 6000four H200 NVL or eight RTX PRO 6000
DeepSeek-V4-Flash-0731, 1M6.4 GiBfour RTX PRO 6000 or two H200 NVLeight RTX PRO 6000 as two copies of four, or four H200 NVL
Llama 4 Scout FP8, 1M192 GiBfour RTX PRO 6000 or four H200 NVLmore than eight H200 NVL
GLM-5.3 FP8, 128K12.2 GiBeight H200 NVLeight H200 NVL with DCP or HiSparse
GLM-5.3 FP8, 1M97.8 GiBeight H200 NVL with DCP or HiSparsemore than eight H200 NVL without HiSparse

Our estimates, not measurements, with a 16-bit cache and the upper counts for Scout and GLM-5.3, from 83.0 GiB per RTX PRO 6000, 123.4 GiB per H200 NVL and 102 GB per DGX Spark (128 GB). DeepSeek-V4-Flash and GLM-5.3 keep their full cache on every card under tensor parallelism; DCP is decode context parallelism.

One DGX Spark holds gpt-oss-120b with seven 128K conversations, and Qwen3.8-27B in FP8 with one 1M conversation by a margin of about 2 GiB; our guide to what fits in 128 GB on a DGX Spark explains the 102 GB working set.

On the H200 NVL, the DeepSeek-V4-Flash rows assume its FP4 experts run weight-only, which our hub article could not confirm from DeepSeek or vLLM documents. Llama 4 Scout at 1M needs two cards instead of four if its local layers keep only their chunk.

We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the document sizes and the context you need, with the number of conversations at peak, and we size the server.

Long conversations on several GPUs

Tensor parallelism splits a model’s KV heads across the cards: eight for Llama 3.3 70B, gpt-oss-120b and Llama 4 Scout, four for Qwen3.8-27B. vLLM’s blog of 7 August 2026 on decode context parallelism states that “once TP exceeds the number of KV heads, the cache starts duplicating across GPUs”, so ten 1M conversations of Qwen3.8-27B run as two copies on four H200 NVL each.

Latent attention has no heads to split. The same post states that under tensor parallelism “the latent KV cache is replicated in full across every TP rank”. SGLang’s guide to data-parallel attention describes the same duplication for DeepSeek’s MLA, while each data-parallel replica “Maintains its own KV cache (no duplication)”. In both layouts every card holds the whole cache of each conversation it serves, so one card’s spare memory limits the longest context; DeepSeek-V4-Flash, with one KV head, behaves the same way under tensor parallelism.

Decode context parallelism (DCP) splits one request’s cache across a tensor-parallel set by token position. vLLM sets it with --decode-context-parallel-size for MLA and GQA models, and release 0.30.0 lists “PCP+DCP on sparse-MLA models”; SGLang sets it with --dcp-size and shows it for DeepSeek-V3.1. Pipeline parallelism also divides the cache, as each card keeps only its own layers’ cache. A 1M conversation of GLM-5.3 on eight H200 NVL, with about 35 GiB per card left after the weights, needs DCP, HiSparse offloading or, with an FP8 cache, two pipeline stages; we found no recipe that runs GLM-5.3 or DeepSeek-V4 with DCP. With 123 GiB per card for weights and cache against 83 GiB, the H200 NVL holds long MLA conversations on fewer cards.

Prefix caching, FP8 KV cache and chunked prefill

The features below follow vLLM’s “latest” documentation, a developer preview, as read on 9 October 2026; the current release is 0.31.0 of 5 October. Automatic prefix caching “caches the KV cache of existing queries, so that a new query can directly reuse the KV cache” when the prefix matches, as with repeated questions about one long document or a multi-round conversation. It “only reduces the time of processing the queries (the prefilling phase)”.

An FP8 cache, set with --kv-cache-dtype fp8, “can significantly reduce its memory footprint” according to vLLM and roughly halves the attention figures in both tables. For DeepSeek-V3.2, vLLM’s post of 29 September 2025 describes an FP8 format of 656 bytes per token and layer, 57 per cent of the 1,152 bytes in BF16.

A long prompt is computed in a prefill before the first token. In vLLM V1 “chunked prefill is enabled by default whenever possible”: large prefills run in chunks batched with decode requests, so a long prompt holds up other users’ answers less, and max_num_batched_tokens trades inter-token latency against time to first token. We found no vendor figure for the first-token time of a 128K or 1M prompt on these cards, so test with your own prompt lengths. Our comparison of vLLM, SGLang, TensorRT-LLM and Ollama covers the engines.

KV cache offloading to CPU memory and NVMe

vLLM’s blog of 10 September 2026 says its tiered offloading “preserves evicted KV data across host memory, storage, and remote peers”, and on a hit “vLLM reloads the data from a lower tier” instead of recomputing. Its development documentation lists --kv-offloading-size, the CPU buffer in GiB, with the backends native and lmcache. NVIDIA Dynamo’s documentation for release 1.5.1 describes the native path: “vLLM copies sealed GPU KV blocks to pinned CPU memory.”

LMCache 0.5.5 of 12 September 2026 moves KV caches into CPU RAM, local SSD and remote back ends such as Redis/Valkey or S3-compatible storage; its PyPI page says it “reduces TTFT” and improves throughput for long-context, multi-turn and RAG workloads. Dynamo’s KV Block Manager (KVBM), in its development documentation, spans GPU memory, pinned host memory, remote memory and SSDs for vLLM and TensorRT-LLM, and NVIDIA writes that offloading “is most effective when KV Cache exceeds GPU memory and cache reuse outweighs the overhead of transferring data.”

These documents treat offloaded blocks as a cache for reuse by a returning session or a repeated document. By our reading, the conversation being generated keeps its cache in GPU memory, so this offloading saves prefill time rather than raising the longest context a configuration holds.

vLLM’s HiSparse for sparse attention goes further. vLLM’s blog of 8 September 2026 reports GLM-5.3 at its full 1M context on one node of eight H200 with Hybrid HiSparse, which moves the cache pages the indexer does not select to pinned CPU memory when GPU memory runs short. HiSparse came with release 0.30.0, 0.31.0 lists “HiSparse hardening”, and the post says it is “currently implemented only for NVIDIA GPUs”.

Our AI servers have ECC memory sized to the GPU pool and NVMe tiers for models and indexes. Write to us with the context lengths and how often prompts repeat, and we size memory and storage for the offloaded cache with the GPUs.

Long context or RAG

Long context puts the whole document set into every request, with its own cache and prefill per conversation, while retrieval-augmented generation puts in only the retrieved passages. Long context suits one large document or code base queried repeatedly, where prefix caching saves the repeated prefill. RAG suits a corpus larger than any context and documents with different access rights, which our guide to RAG on company data filters at retrieval.

What we supply

We supply the DGX Spark Founders Edition, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, and the H200 NVL with two-way and four-way NVLink bridges, as cards or in AI servers built to order. We size and source the GPUs, system memory and NVMe storage together, on one EU contract and invoice with manufacturer warranty, and check the rack, power and airflow before we quote. Our professional GPU range lists every card. Serving engines, RAG and MLOps are our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

How much VRAM do I need for a 128K context?
It depends on the model’s attention design more than on its size. One 128K conversation takes 40 GiB of 16-bit KV cache on Llama 3.3 70B, 9.6 GiB on DeepSeek-V3.2, 4.5 GiB on gpt-oss-120b and about 0.8 GiB on DeepSeek-V4-Flash, on top of the weights. By our estimate one H200 NVL or two RTX PRO 6000 hold Llama 3.3 70B in FP8 with one 128K conversation, and one RTX PRO 6000 holds gpt-oss-120b with four.
How much memory does a 1M token context need?
With a 16-bit cache one 1M-token conversation needs about 6.4 GiB on DeepSeek-V4-Flash, 64 GiB on Qwen3.8-27B, up to 98 GiB on GLM-5.3 and up to 192 GiB on Llama 4 Scout, by our arithmetic from their configuration files. An FP8 cache roughly halves these figures. Under tensor parallelism a latent-attention conversation is copied to every card, so a long one needs decode context parallelism, pipeline parallelism or, for GLM-5.3 in vLLM, HiSparse offloading.
How do I calculate the KV cache size per token?
For full attention multiply 2 × layers × KV heads × head dimension × bytes per value, all read from the model’s config.json; Llama 3.3 70B gives 2 × 80 × 8 × 128 × 2 bytes = 320 KiB. For latent attention take layers × (kv_lora_rank + RoPE dimension) × bytes. Sliding-window and linear-attention layers hold a fixed amount and are left out of the per-token figure.
Which GPU is suited to long-context LLM inference?
A card with more memory per GPU, because under tensor parallelism a latent-attention conversation is copied to every card and a GQA model splits only across as many cards as it has KV heads. The H200 NVL leaves about 123 GiB per card for weights and cache against 83 GiB on the RTX PRO 6000, and bridges 2 or 4 cards over NVLink. One DGX Spark holds a mid-size model such as Qwen3.8-27B with one 1M conversation, by our estimate.
Does KV cache offloading allow longer contexts?
By our reading of the vLLM, LMCache and NVIDIA Dynamo documentation, offloading keeps cache blocks in CPU memory or on SSD for reuse, so a returning session or repeated document skips the prefill, while the conversation being generated keeps its cache in GPU memory. The exception we found is vLLM’s HiSparse for GLM-5.3, which moves unselected cache pages to CPU memory during decoding; vLLM’s blog of 8 September 2026 reports the full 1M context on eight H200 with it.
Does an FP8 KV cache reduce memory for long context?
Yes, it stores keys and values in one byte instead of two and roughly halves the cache of every conversation; vLLM’s FP8 format for DeepSeek-V3.2 takes 57 per cent of the BF16 size. In vLLM it is set with --kv-cache-dtype fp8, and without calibration all scales are set to 1.0, so vLLM recommends calibrated scales for accuracy. Check your model’s accuracy with the FP8 cache on your own evaluation set before you size the server on it.

Send us the models, the context length per conversation, the number of conversations at peak and how often prompts repeat. We reply within one business day with the memory budget, a configuration with GPUs, system memory and NVMe storage, and a written quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna