AI agents hardware: why agentic AI needs more GPU than chat, and how to size an on-premise server
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- An agent task is a chain of model calls: each step sends the system prompt, the tool definitions, the history and every tool result again, so one task processes many times the tokens of a chat turn
- Anthropic’s engineering blog of 13 June 2025 states that in its data “agents typically use about 4× more tokens than chat interactions”, and multi-agent systems about 15× more; NVIDIA’s Dynamo documentation says coding agents make hundreds of API calls per session
- In an illustrative example with assumed values, not measurements, an agent task of 8 calls sends 102,400 input tokens against 3,500 for one chat turn; prefix caching cuts the tokens to compute to about 19,100 if the earlier steps stay cached
- vLLM has enabled automatic prefix caching by default since V1; it shortens prefill only, and cached blocks are evicted least recently used first, so a long tool pause can cost the agent its prefix
- Size an agent server by tasks in the busiest hour × model calls per task × tokens per call, then check memory for open sessions and test latency per step on the candidate GPU
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Why AI agents need more GPU than chat
AI agents need more GPU than chat because one agent task is a chain of model calls. An agent calls the model, runs a tool, adds the result to its context and calls the model again, and every call sends the system prompt, the tool definitions and the whole history so far. A task therefore processes many times the tokens of a chat turn, holds a longer KV cache while it runs, and makes the user wait for the sum of all steps rather than for one answer.
Anthropic’s engineering blog of 13 June 2025 states that in Anthropic’s data “agents typically use about 4× more tokens than chat interactions”. The same post puts multi-agent systems at about 15 times the tokens of chats. NVIDIA’s Dynamo documentation on agentic inference, updated on 2 October 2026, says coding agents make “hundreds of API calls per coding session, each carrying the full conversation history”.
These figures come from cloud services, so the next section uses an illustrative example with assumed values. Connecting an agent to vLLM or Ollama is covered in our guide to AI agents with n8n on a private LLM.
Model calls and tokens per agent task: an illustrative example
The example compares one chat turn with one helpdesk agent task that reads a ticket, queries two systems and drafts a reply. Every value in it is an illustrative assumption, not a measurement. The chat turn sends a 1,000-token system prompt, 2,000 tokens of earlier conversation and a 500-token question, and the model answers in 500 tokens.
The agent starts from 6,000 tokens of system prompt and tool definitions plus a 500-token request. It makes 8 model calls, within the default of 10 that n8n’s AI Agent allows per prompt. Each call produces 300 tokens of reasoning and a tool call, and each of the 7 tool results adds 1,500 tokens. The first call sends 6,500 tokens and the eighth sends 19,100, so the task sends 102,400 input tokens in all.
| ITEM | CHAT TURN | AGENT TASK | AGENT VS CHAT |
|---|---|---|---|
| Model calls | 1 | 8 | 8 times |
| Input tokens sent | 3,500 | 102,400 | about 29 times |
| Prefill with prefix cache | about 1,000 | about 19,100 | about 19 times |
| Output tokens | 500 | 2,400 | about 5 times |
| Context at the end | 4,000 | 19,400 | about 5 times |
| KV cache, gpt-oss-120b | 0.14 GiB | 0.67 GiB | about 5 times |
Illustrative example: all values are assumptions, not measurements; ratios by our arithmetic. Max Iterations default from n8n’s documentation; KV cache at 36 KiB per token in 16-bit for gpt-oss-120b, as in our private ChatGPT sizing guide.
The example’s ratio of about 29 is far above Anthropic’s 4 because the two figures count different things. Anthropic does not say whether a chat interaction in its data is one turn or a whole conversation, or how cached tokens count. The example sets one chat turn against a whole task and counts every token sent, including the 6,000 tokens of system prompt and tool definitions resent on all 8 calls. After prefix caching the ratio falls to about 19, and a chat of several turns would lower it further. Neither ratio transfers to your workflow, so count calls and tokens per task in a pilot’s gateway log.
Prefill load and prefix caching for agents
Input tokens are processed in the prefill, which is compute work and sets the time to the first token, while generation is limited by memory bandwidth, so an agent shifts the load towards prefill. In the example it sends 43 input tokens for every token it generates, against 7 in the chat turn.
Prefix caching recovers most of that work. vLLM’s documentation says automatic prefix caching “caches the KV cache of existing queries, so that a new query can directly reuse the KV cache if it shares the same prefix”. vLLM’s blog of 27 January 2025 announced “we now enable prefix caching by default in V1”. The limit is stated in the same documentation: APC “does not reduce the time of generating new tokens (the decoding phase)”.
If every earlier step is still cached, each call of the example computes only the new tool result and the previous output, 1,800 tokens, after a first call of 6,500. That gives about 19,100 tokens to compute instead of 102,400, so about 81 per cent of the input comes from the cache. For a hosted coding agent on a managed cloud API, not a Dynamo test, NVIDIA’s Dynamo documentation states that after the first call “every subsequent call to the same worker hits 85-97% cache”.
The cache is not reserved for a waiting agent. vLLM’s design notes describe how a finished request’s blocks stay reusable only until vLLM evicts them, least recently used first, to make room for new work, so a waiting agent’s blocks compete with every other request. NVIDIA’s Dynamo documentation warns that “A 2-30 second tool call pause can age out an agent’s entire prefix, forcing full recomputation when it resumes.”
Latency adds up across agent steps
A chat user sees text after one time to first token and reads while the answer streams. An agent’s result arrives only after the last step, and each step adds queueing, prefill, generation and the tool’s run time.
With example values of 6 seconds per model call and 2 seconds per tool, the 8 calls and 7 tools of the example take 62 seconds. Our guide to a private coding assistant reports NVIDIA’s measured agent tasks with 32K-token prompts on one DGX Spark.
In NVIDIA’s blog of 8 May 2026, a 52K-token prompt with a stable prefix reached its first token in 168 ms on one B200 GPU. A per-session header that varied within the prefix of the same prompt raised that to 912 ms, because it prevented reuse of the cached prefix.
Tool definitions, MCP and the prompt prefix
Every tool the agent may call is described in its prompt with a name, a description and a JSON Schema for its arguments, as the Model Context Protocol defines them. Anthropic’s engineering blog of 4 November 2025 notes that “Most MCP clients load all tool definitions upfront directly into context”, and that tool descriptions take context space, “increasing response time and costs”. In its example workflow, presenting the tools as code that the agent loads only when needed cut the token usage from 150,000 to 2,000.
Revision 2026-07-28 of the MCP specification asks servers to return tools in a deterministic order, which “improves LLM prompt cache hit rates when tools are included in model context”. Give each agent only the tools its task needs, which n8n’s MCP Client Tool allows per server, keep their order and descriptions fixed, and put anything that changes per session after the stable part of the prompt. Securing the servers is covered in our guide to MCP server security.
Levers that reduce the GPU load of agents
| LEVER | EFFECT ON THE GPU | DOCUMENTED IN |
|---|---|---|
| Automatic prefix caching | skips prefill of a shared prefix; decoding unchanged | vLLM docs; on by default since V1 |
| Stable prompt start | keeps cache hits across steps and sessions | MCP 2026-07-28 tools page; NVIDIA blog, 8 May 2026 |
| Cache-aware routing | the next call reaches the copy that holds its prefix | NVIDIA Dynamo docs, updated 2 October 2026 |
| KV cache offloading | an evicted prefix reloads from CPU memory | vLLM and LMCache docs; off by default in vLLM |
| Fewer tool definitions | shorter prompt on every call | Anthropic, 4 November 2025; n8n MCP Client Tool |
| Structured outputs | JSON that parses; tool calls only with strict | vLLM structured outputs and tool calling docs |
| Small model for routing | simple steps off the large model | position paper, arXiv 2506.02153, v3 of 22 September 2026 |
| Iteration limit | caps model calls per task | n8n AI Agent, Max Iterations |
Sources as read on 10 October 2026; vLLM “latest” documentation, current release 0.31.0. Effects as each source describes them, without measured gains for your workload.
With several copies of a model behind a load balancer, the Dynamo documentation warns that “Without cache-aware routing, turn 2 of a conversation has a ~1/N chance of landing on the same worker as turn 1”, N being the number of workers. Session affinity at the load balancer, or a KV-aware router such as Dynamo’s, keeps an agent on one copy. Offloading keeps evicted blocks in host memory for a returning request, such as an agent back from a tool; our guide to long-context LLM hardware describes vLLM’s and LMCache’s options.
vLLM’s structured outputs, with the xgrammar or guidance backends, constrain a reply to a JSON schema, a regex, a choice list or a grammar. For tool calls in auto mode, vLLM’s tool calling documentation of 8 October 2026 applies such constraints only when a tool sets strict: true or the server raises --tool-strict-level; otherwise “vLLM extracts tool calls from raw text, so arguments may occasionally be malformed”. A call that parses still leaves the choice of tool to the model. A position paper on arXiv, first submitted on 2 June 2025, argues that small language models are “sufficiently powerful, inherently more suitable, and necessarily more economical for many invocations in agentic systems”. Routing, classification and extraction steps can run on a small model.
Sizing a GPU server for agentic workloads
Agent load is counted in calls and tokens, not in users. Our guide to private ChatGPT server sizing by company size uses Little’s law to turn requests into requests in flight, and the same method applies to agent tasks.
- Tasks: count the agent tasks in the busiest hour, per workflow, from a pilot or as a marked estimate.
- Calls and tokens: take model calls per task, input and output tokens per call and the context at the last step from the gateway log.
- Token load: multiply tasks by tokens for the hourly input, the input after prefix cache hits and the output.
- Sessions: multiply tasks per second by the task duration for the average number of open agent sessions, apply a peak factor and multiply by the KV cache at the last step, the memory that keeps their prefixes cached between calls.
- Test: load-test the candidate GPU with recorded tasks, prefix caching on, and check time to first token per step and the task duration against your target.
With the example task, 120 tasks in the busiest hour mean 960 model calls, 12.3 million input tokens sent, about 2.3 million tokens to compute after cache hits and 288,000 output tokens. At 62 seconds per task, about 2 sessions are open on average and about 4 at a peak factor of 2. Their cache of about 2.8 GiB for gpt-oss-120b is a small share of one card, so here the time per step decides the hardware, not memory.
We build inference servers sized by model size and concurrent users. Send us the tasks per busy hour, the calls per task and the model through the form below, and we return a configuration and quote within one business day.
Which GPUs suit on-premise agents
For a pilot, one DGX Spark Founders Edition with 128 GB runs a model such as gpt-oss-120b and the workflow on a desk. For production, a server with RTX PRO 6000 Server Edition cards carries one copy of that model per card. By the estimate in our private ChatGPT sizing guide, one card keeps about 22.2 GiB of cache beside that model, room for about 33 sessions at the example’s 19,400 tokens.
The H200 NVL leaves about 62.5 GiB beside the same model, about 93 such sessions. Its 4.8 TB/s of memory bandwidth, against 1,597 GB/s on the RTX PRO 6000 Server Edition in NVIDIA’s specifications, shortens the generation part of every step. NVLink bridges join two or four H200 NVL when a larger model must be split.
Small models for routing, embeddings and reranking fit a 24 GB MIG instance of an RTX PRO 6000 or an L4, outside the main model’s batch. Where agents run business processes, two servers that each carry the peak remove the single point of failure, with session affinity so that each agent keeps its cache.
Our Private AI/ML service includes agent workflows and ITOps automation with n8n and Ansible. Describe the first agent workflow you want to run and the tools it will call in the form below.
What we supply
We supply the DGX Spark Founders Edition, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, the H200 NVL with NVLink bridges, and the L4, L40S and smaller RTX PRO cards for routing and RAG models. They come as cards or in AI servers built to order with 2 to 8 GPUs per node, burn-in tested, with manufacturer warranty, on one EU contract and invoice. We check the rack, power and airflow before we quote. Models, agent workflows and the platform on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What hardware do AI agents need?
Why do agentic AI workloads need more GPU than chat?
How many model calls does an AI agent make per task?
Does prefix caching help AI agents?
Can a local LLM server handle tool calling for agents?
How many AI agents can one GPU run?
Send us the agent workflows you plan, the tasks in the busiest hour, the model, the tools each agent calls and your response targets. We reply within one business day with a configuration and a quote, and we check the rack, power and airflow before we quote.
Talk to an expertWe reply within one business day