NVIDIA Dynamo and disaggregated serving: when splitting prefill and decode pays across servers
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Disaggregated serving runs prefill (the prompt) and decode (the answer) in separate GPU pools and moves each request’s KV cache between them; vLLM’s documentation states that its disaggregated prefill “DOES NOT improve throughput” and lists separate tuning of time to first token and inter-token latency as its purpose
- NVIDIA Dynamo, at release 1.5.1 of 6 October 2026 in its documentation, puts a frontend, a KV-aware router, a planner and NIXL transfers in front of vLLM 0.28.0, SGLang 0.5.18 or TensorRT-LLM 1.3.0rc25; llm-d, documented as v0.10, does the same on Kubernetes and underlies KServe’s LLMInferenceService
- KV-aware routing sends each request to the copy that already holds its prompt prefix; it works with ordinary replicas and needs no RDMA, so it is the first step for two to four servers
- Between servers the KV transfer needs RDMA: llm-d says disaggregated serving “requires high performance (RDMA) interconnects between nodes” and keeps TCP for testing; a 4,000-token prompt of Llama 3.3 70B carries about 0.66 GB of FP8 cache, about 26 ms at 200 Gb/s by our arithmetic
- Each pool holds its own copy of the weights, so on eight RTX PRO 6000 with Llama 3.3 70B and an FP8 cache, two prefill and six decode cards hold 72 conversations at 8K where eight replicas hold 96
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
NVIDIA Dynamo and disaggregated serving: what they do and when they pay
Disaggregated serving runs the two phases of an LLM request on separate GPUs: prefill workers process the prompt and build its KV cache, decode workers generate the answer, and the cache moves from one to the other for every request. NVIDIA Dynamo and llm-d are open-source frameworks that add this split, plus routing by KV cache hits, in front of vLLM, SGLang or TensorRT-LLM. At large scale the split lets each pool run its own parallelism and keeps long prompts from stalling answers in progress.
With two to four servers the gain depends on prompt lengths, cache reuse and traffic. vLLM’s documentation on disaggregated prefill, dated 3 October 2026, states: “Disaggregated prefill DOES NOT improve throughput.” Dynamo’s documentation adds: “At low concurrency, an aggregated worker often avoids transfer and pool-fragmentation overhead.” KV-aware routing, by contrast, helps whenever several copies of a model see repeated prefixes, and it needs no special network.
Why prefill and decode are split
Prefill computes the whole prompt in one pass and is limited by compute; decode produces one token per step for every request in the batch and is limited by memory bandwidth. Our guide to LLM latency targets shows how prefill sets the time to first token (TTFT) and decode the inter-token latency (ITL). Dynamo’s documentation describes the two phases as having “different computation characteristics and memory footprints”.
On a shared GPU, a long prompt competes with the decode steps of every other request. llm-d’s architecture documentation puts it this way: “For long context requests, prefills can slow down processing of existing requests in the decode phase.” Split pools also allow different layouts, “using a larger TP for the memory-bound decoding phase while a smaller TP for the computation-bound prefill phase.” vLLM names two reasons for the feature: “Tuning time-to-first-token (TTFT) and inter-token-latency (ITL) separately” and “Controlling tail ITL”.
The idea was formalised in the DistServe paper on arXiv (first version January 2024), which measures goodput, the request rate served within both TTFT and per-token limits per GPU. Its abstract reports that DistServe “can serve 7.4x more requests or 12.6x tighter SLO” than the systems it was compared with. vLLM marks its own feature as experimental, and a single vLLM instance already softens the conflict with chunked prefill, which mixes prompt chunks with decode steps.
NVIDIA Dynamo and llm-d: the components
| COMPONENT | ROLE | PROJECT |
|---|---|---|
| Frontend | OpenAI-compatible API on port 8000; applies the chat template and tokenises | Dynamo |
| KV router | picks the copy with the lowest projected cost from cache overlap and load | Dynamo; llm-d Router (EPP) |
| Prefill router, sidecar | sends the prompt to a prefill worker, then the request to decode | Dynamo Prefill |
| Workers | prefill or decode instances of vLLM, SGLang or TensorRT-LLM | both |
| Planner | sets the number of prefill and decode workers from live metrics | Dynamo; llm-d autoscaling |
| NIXL | moves KV blocks between GPUs, hosts and storage | NVIDIA, used by both |
| KV offloading | keeps cache blocks in CPU memory or on SSD for reuse | Dynamo (LMCache, FlexKV); llm-d KV offloader |
NVIDIA Dynamo documentation (marked latest, v1.5.1) and GitHub repository, llm-d documentation (marked v0.10, latest) and NIXL repository, read on 10 October 2026.
NVIDIA calls Dynamo “an open source inference framework for serving generative AI”. It is licensed under Apache 2.0 and, in its repository’s words, “Built in Rust for performance, Python for extensibility.” Its compatibility page, read on 10 October 2026, lists release 1.5.1 of 6 October 2026, a patch to 1.5.0 of 18 September, with vLLM 0.28.0, SGLang 0.5.18 and TensorRT-LLM 1.3.0rc25. It names CUDA 13.0 or 13.1, a 580 or later driver, Ubuntu 24.04 and the Blackwell, Hopper, Ada Lovelace and Ampere architectures. The GitHub releases page, read the same day, still marks 1.4.2 as the latest release and shows 1.5.0 only as model-specific pre-releases, so check the version of the image you deploy. The vLLM version Dynamo pairs with is three minor releases behind vLLM’s own 0.31.0 of 5 October, so a feature new in vLLM reaches Dynamo with a later release. Our comparison of vLLM, SGLang and TensorRT-LLM covers the engines themselves.
Dynamo runs on bare metal, under Slurm or on Kubernetes through its operator. Workers find each other through a discovery backend, etcd by default outside Kubernetes, and KV events travel over ZMQ by default, with NATS as an option. llm-d is a CNCF sandbox project founded by Red Hat, Google Cloud, IBM Research, CoreWeave and NVIDIA. Its documentation, marked v0.10 (latest), describes a proxy that conforms to the Gateway API Inference Extension; the project’s GitHub README shows a version 0.8 badge and announces v0.7 in its news of May 2026. KServe’s LLMInferenceService, at API version v1alpha1 in the KServe 0.20 documentation, is “Built on the foundation of llm-d”, and our guide to a private LLM platform on Kubernetes places its router in that stack.
How a disaggregated request flows
In Dynamo, as its documentation describes the flow with vLLM:
- The frontend accepts the request, applies the chat template and tokenises it.
- The PrefillRouter selects a prefill worker, by KV-aware routing or by load.
- The prefill worker computes the KV cache and returns transfer metadata with the result.
- The router adds the metadata to the request and routes it to a decode worker.
- The decode worker pulls the KV blocks from the prefill worker through NIXL and generates the answer.
Dynamo states that NIXL transfers the cache “directly from the VRAM of the prefill engine to the VRAM of the decode engine”. With vLLM and TensorRT-LLM, prefill runs first and decode waits for it, while with SGLang decode starts while the transfer proceeds. In vLLM itself, a prefill instance runs with kv_role set to kv_producer and a decode instance with kv_consumer, through the NixlConnector.
llm-d turns the order around. Its endpoint picker selects a decode worker first and asks a decider “given how much of the prompt is cached on D, should this request run disagg?” Only a request with a large uncached part goes to a remote prefill worker; for the others the decode worker computes the prompt itself. The routing proxy runs as a sidecar in each decode pod.
KV-aware routing without disaggregation
A common layout with two to four servers runs complete copies of a model behind a load balancer. KV-aware routing improves exactly that layout. Dynamo’s router “chooses the worker that can serve a request at the lowest projected cost”: “A large cache overlap lowers the prefill part of the score”, while active prefill and decode blocks raise it. Workers publish KV creation and release events, and Dynamo notes that “Round-robin and load-only policies do not account for that difference.”
llm-d offers two modes. The approximate mode estimates tokens from characters and assumes the chosen pod now holds the prefix; the precise mode reads KVEvents that vLLM emits over ZMQ and, in llm-d’s words, “provides 100% accuracy by leveraging actual token data.” Both depend on prefix caching in the engine, such as vLLM’s automatic prefix caching.
The gain comes from shared prefixes: a long system prompt, multi-turn chats, agents that resend their history and RAG prompts over the same documents. Our guide to long-context LLM hardware covers prefix caching and offloading of cache blocks to CPU memory and SSD, which Dynamo and llm-d both integrate.
What the KV transfer needs from the network
Disaggregation across servers moves the cache of every request over the network. llm-d states: “Disaggregated Serving requires high performance (RDMA) interconnects between nodes for efficient KV transfer.” Without RDMA, NIXL falls back to TCP, which llm-d calls “not efficient” and says “should only be used for testing and development”. Its prefill and decode guide adds that “high-bandwidth networking (IB, RoCE, EFA) is highly recommended for production usage.” NIXL’s default transport in vLLM is UCX, and Dynamo lists NVLink, InfiniBand through UCX and PCIe as paths.
The size of one transfer follows from the model’s configuration. Llama 3.3 70B has 80 layers, 8 KV heads and a head dimension of 128, so its cache takes 320 KiB per token in 16-bit and half that in FP8. A 4,000-token RAG prompt carries about 0.66 GB of FP8 cache. At line rate that is about 0.2 s at 25 Gb/s, 26 ms at 200 Gb/s and 13 ms at 400 Gb/s by our arithmetic, against an example TTFT target of 2.5 s for RAG in our latency guide. Inside one server the same transfer runs over PCIe Gen5, about 10 ms at the 64 GB/s an x16 link carries in each direction, or over the NVLink bridge between two or four H200 NVL, which NVIDIA rates at 900 GB/s per GPU. The RTX PRO 6000 has no NVLink.
Our guide to a GPU inference cluster without InfiniBand covers the fabric itself, RoCE or InfiniBand, and NVIDIA’s reference architecture.
We build AI servers with 2 to 8 GPUs per node and fit 25 to 400G NICs to the serving pattern. Describe your models, prompt lengths and peak requests in flight in the form below.
Published results and their setups
| RESULT | SETUP | SOURCE, DATE |
|---|---|---|
| Over 2x throughput | Llama 70B on Hopper, vLLM, FP8, 3K in / 50 out; TP8DP2 against prefill TP2DP4 plus decode TP8 | NVIDIA blog, 18 March 2025 |
| Up to 70% more tokens/s | GPT-OSS on AWS p6-b200 cloud instances, P/D against standard vLLM; reported by AWS | llm-d README, read 10 October 2026 |
| 10 to 30% more throughput | GPT-OSS-120B and Llama 3.3 70B on AMD MI300X, same infrastructure; reported by Oracle | llm-d README, read 10 October 2026 |
| 3x output, 2x faster TTFT | prefix-cache-aware routing against round-robin, Llama 3.1 70B on 4 MI300X | llm-d README, read 10 October 2026 |
Vendor- and partner-reported figures as each source states them; the llm-d README dates none of them; none of these setups uses the RTX PRO 6000 or the H200 NVL.
The Hopper result used 3,000-token prompts and 50-token answers, a prefill-heavy mix that favours the split. The NVIDIA blog also tested its router on two HGX H100 nodes with eight copies of DeepSeek-R1-Distill-Llama-70B at tensor parallel 2, on 100,000 R1 requests averaging 4K input and 800 output tokens, and shows the TTFT gain only as a chart. In the sources we read, no result covers the cards in this article, so treat these figures as direction, not as a sizing basis.
Aggregated or disaggregated with two to four servers
Each pool keeps its own copy of the weights, and only decode workers hold the caches of conversations in progress. On a server with eight RTX PRO 6000 running Llama 3.3 70B in FP8 with an FP8 cache, the sizing rule of our AI server specification guide gives 12 conversations at a full 8K context per card. Eight aggregated replicas hold 96; two prefill and six decode cards hold 72. Dynamo’s documentation warns that “When decode KV capacity is the bottleneck, adding prefill replicas can reduce total serving capacity”.
| CONDITION | AGGREGATED | DISAGGREGATED |
|---|---|---|
| Short prompts, long answers | replicas plus KV-aware routing | little to gain |
| RAG: long prompts | chunked prefill first | candidate if ITL misses its target |
| Shared prefixes, agents | KV-aware routing | no gain beyond routing |
| Decode memory is the limit | keep every card for decode | loses conversations |
| Model fits one card | replicas, plain Ethernet | inside one server over PCIe or NVLink |
| Large MoE with DP/EP | pipeline bubbles, per llm-d | “essential” per llm-d; RDMA |
Our reading of the Dynamo (documentation v1.5.1), llm-d (documentation v0.10) and vLLM documentation as read on 10 October 2026; card counts from our specification guide.
For a company of 500 to 2,000 staff, a sensible order is to run replicas behind KV-aware routing first and then measure TTFT and ITL at the 95th percentile with the prompt lengths of production traffic. Disaggregation becomes worth a test when long prompts push ITL over its target while decode memory has room. Dynamo warns that the boundary “should not be encoded as a universal token threshold”. For DP/EP deployments of large mixture-of-experts models llm-d calls disaggregated serving “essential to avoid pipeline bubbles”, and that layout spans servers on an RDMA fabric.
We check the rack, power and airflow before we quote, and the configuration and quote follow within one business day. Send us your prompt and answer lengths, the models and the number of servers through the form below.
What we supply
We build AI servers to order with 2 to 8 GPUs per node, sized by model size and concurrent users, assembled and burn-in tested, with manufacturer warranty on every component. The cards include the RTX PRO 6000 Server Edition, the H200 NVL with NVLink bridges, the L40S and the L4, with 25, 100, 200 or 400G NICs for the network the serving pattern needs. NVIDIA AI Enterprise and vGPU licences come on the same EU contract and invoice. Model deployment with vLLM and a Kubernetes platform with KServe are part of our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What is NVIDIA Dynamo?
What is disaggregated serving?
Does prefill and decode disaggregation increase throughput?
What is the difference between Dynamo and vLLM?
What is llm-d?
What is KV cache aware routing?
Send us the models you serve, typical prompt and answer lengths, the peak number of requests in flight and the number of servers you plan. We reply within one business day with a configuration per server, cards and NICs included, and a written quote.
Talk to an expertWe reply within one business day