BLOG · GUIDE ·

AI agents hardware: why agentic AI needs more GPU than chat, and how to size an on-premise server

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • An agent task is a chain of model calls: each step sends the system prompt, the tool definitions, the history and every tool result again, so one task processes many times the tokens of a chat turn
  • Anthropic’s engineering blog of 13 June 2025 states that in its data “agents typically use about 4× more tokens than chat interactions”, and multi-agent systems about 15× more; NVIDIA’s Dynamo documentation says coding agents make hundreds of API calls per session
  • In an illustrative example with assumed values, not measurements, an agent task of 8 calls sends 102,400 input tokens against 3,500 for one chat turn; prefix caching cuts the tokens to compute to about 19,100 if the earlier steps stay cached
  • vLLM has enabled automatic prefix caching by default since V1; it shortens prefill only, and cached blocks are evicted least recently used first, so a long tool pause can cost the agent its prefix
  • Size an agent server by tasks in the busiest hour × model calls per task × tokens per call, then check memory for open sessions and test latency per step on the candidate GPU

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Why AI agents need more GPU than chat

AI agents need more GPU than chat because one agent task is a chain of model calls. An agent calls the model, runs a tool, adds the result to its context and calls the model again, and every call sends the system prompt, the tool definitions and the whole history so far. A task therefore processes many times the tokens of a chat turn, holds a longer KV cache while it runs, and makes the user wait for the sum of all steps rather than for one answer.

Anthropic’s engineering blog of 13 June 2025 states that in Anthropic’s data “agents typically use about 4× more tokens than chat interactions”. The same post puts multi-agent systems at about 15 times the tokens of chats. NVIDIA’s Dynamo documentation on agentic inference, updated on 2 October 2026, says coding agents make “hundreds of API calls per coding session, each carrying the full conversation history”.

These figures come from cloud services, so the next section uses an illustrative example with assumed values. Connecting an agent to vLLM or Ollama is covered in our guide to AI agents with n8n on a private LLM.

Model calls and tokens per agent task: an illustrative example

The example compares one chat turn with one helpdesk agent task that reads a ticket, queries two systems and drafts a reply. Every value in it is an illustrative assumption, not a measurement. The chat turn sends a 1,000-token system prompt, 2,000 tokens of earlier conversation and a 500-token question, and the model answers in 500 tokens.

The agent starts from 6,000 tokens of system prompt and tool definitions plus a 500-token request. It makes 8 model calls, within the default of 10 that n8n’s AI Agent allows per prompt. Each call produces 300 tokens of reasoning and a tool call, and each of the 7 tool results adds 1,500 tokens. The first call sends 6,500 tokens and the eighth sends 19,100, so the task sends 102,400 input tokens in all.

ITEMCHAT TURNAGENT TASKAGENT VS CHAT
Model calls188 times
Input tokens sent3,500102,400about 29 times
Prefill with prefix cacheabout 1,000about 19,100about 19 times
Output tokens5002,400about 5 times
Context at the end4,00019,400about 5 times
KV cache, gpt-oss-120b0.14 GiB0.67 GiBabout 5 times

Illustrative example: all values are assumptions, not measurements; ratios by our arithmetic. Max Iterations default from n8n’s documentation; KV cache at 36 KiB per token in 16-bit for gpt-oss-120b, as in our private ChatGPT sizing guide.

The example’s ratio of about 29 is far above Anthropic’s 4 because the two figures count different things. Anthropic does not say whether a chat interaction in its data is one turn or a whole conversation, or how cached tokens count. The example sets one chat turn against a whole task and counts every token sent, including the 6,000 tokens of system prompt and tool definitions resent on all 8 calls. After prefix caching the ratio falls to about 19, and a chat of several turns would lower it further. Neither ratio transfers to your workflow, so count calls and tokens per task in a pilot’s gateway log.

Prefill load and prefix caching for agents

Input tokens are processed in the prefill, which is compute work and sets the time to the first token, while generation is limited by memory bandwidth, so an agent shifts the load towards prefill. In the example it sends 43 input tokens for every token it generates, against 7 in the chat turn.

Prefix caching recovers most of that work. vLLM’s documentation says automatic prefix caching “caches the KV cache of existing queries, so that a new query can directly reuse the KV cache if it shares the same prefix”. vLLM’s blog of 27 January 2025 announced “we now enable prefix caching by default in V1”. The limit is stated in the same documentation: APC “does not reduce the time of generating new tokens (the decoding phase)”.

If every earlier step is still cached, each call of the example computes only the new tool result and the previous output, 1,800 tokens, after a first call of 6,500. That gives about 19,100 tokens to compute instead of 102,400, so about 81 per cent of the input comes from the cache. For a hosted coding agent on a managed cloud API, not a Dynamo test, NVIDIA’s Dynamo documentation states that after the first call “every subsequent call to the same worker hits 85-97% cache”.

The cache is not reserved for a waiting agent. vLLM’s design notes describe how a finished request’s blocks stay reusable only until vLLM evicts them, least recently used first, to make room for new work, so a waiting agent’s blocks compete with every other request. NVIDIA’s Dynamo documentation warns that “A 2-30 second tool call pause can age out an agent’s entire prefix, forcing full recomputation when it resumes.”

Latency adds up across agent steps

A chat user sees text after one time to first token and reads while the answer streams. An agent’s result arrives only after the last step, and each step adds queueing, prefill, generation and the tool’s run time.

With example values of 6 seconds per model call and 2 seconds per tool, the 8 calls and 7 tools of the example take 62 seconds. Our guide to a private coding assistant reports NVIDIA’s measured agent tasks with 32K-token prompts on one DGX Spark.

In NVIDIA’s blog of 8 May 2026, a 52K-token prompt with a stable prefix reached its first token in 168 ms on one B200 GPU. A per-session header that varied within the prefix of the same prompt raised that to 912 ms, because it prevented reuse of the cached prefix.

Tool definitions, MCP and the prompt prefix

Every tool the agent may call is described in its prompt with a name, a description and a JSON Schema for its arguments, as the Model Context Protocol defines them. Anthropic’s engineering blog of 4 November 2025 notes that “Most MCP clients load all tool definitions upfront directly into context”, and that tool descriptions take context space, “increasing response time and costs”. In its example workflow, presenting the tools as code that the agent loads only when needed cut the token usage from 150,000 to 2,000.

Revision 2026-07-28 of the MCP specification asks servers to return tools in a deterministic order, which “improves LLM prompt cache hit rates when tools are included in model context”. Give each agent only the tools its task needs, which n8n’s MCP Client Tool allows per server, keep their order and descriptions fixed, and put anything that changes per session after the stable part of the prompt. Securing the servers is covered in our guide to MCP server security.

Levers that reduce the GPU load of agents

LEVEREFFECT ON THE GPUDOCUMENTED IN
Automatic prefix cachingskips prefill of a shared prefix; decoding unchangedvLLM docs; on by default since V1
Stable prompt startkeeps cache hits across steps and sessionsMCP 2026-07-28 tools page; NVIDIA blog, 8 May 2026
Cache-aware routingthe next call reaches the copy that holds its prefixNVIDIA Dynamo docs, updated 2 October 2026
KV cache offloadingan evicted prefix reloads from CPU memoryvLLM and LMCache docs; off by default in vLLM
Fewer tool definitionsshorter prompt on every callAnthropic, 4 November 2025; n8n MCP Client Tool
Structured outputsJSON that parses; tool calls only with strictvLLM structured outputs and tool calling docs
Small model for routingsimple steps off the large modelposition paper, arXiv 2506.02153, v3 of 22 September 2026
Iteration limitcaps model calls per taskn8n AI Agent, Max Iterations

Sources as read on 10 October 2026; vLLM “latest” documentation, current release 0.31.0. Effects as each source describes them, without measured gains for your workload.

With several copies of a model behind a load balancer, the Dynamo documentation warns that “Without cache-aware routing, turn 2 of a conversation has a ~1/N chance of landing on the same worker as turn 1”, N being the number of workers. Session affinity at the load balancer, or a KV-aware router such as Dynamo’s, keeps an agent on one copy. Offloading keeps evicted blocks in host memory for a returning request, such as an agent back from a tool; our guide to long-context LLM hardware describes vLLM’s and LMCache’s options.

vLLM’s structured outputs, with the xgrammar or guidance backends, constrain a reply to a JSON schema, a regex, a choice list or a grammar. For tool calls in auto mode, vLLM’s tool calling documentation of 8 October 2026 applies such constraints only when a tool sets strict: true or the server raises --tool-strict-level; otherwise “vLLM extracts tool calls from raw text, so arguments may occasionally be malformed”. A call that parses still leaves the choice of tool to the model. A position paper on arXiv, first submitted on 2 June 2025, argues that small language models are “sufficiently powerful, inherently more suitable, and necessarily more economical for many invocations in agentic systems”. Routing, classification and extraction steps can run on a small model.

Sizing a GPU server for agentic workloads

Agent load is counted in calls and tokens, not in users. Our guide to private ChatGPT server sizing by company size uses Little’s law to turn requests into requests in flight, and the same method applies to agent tasks.

  1. Tasks: count the agent tasks in the busiest hour, per workflow, from a pilot or as a marked estimate.
  2. Calls and tokens: take model calls per task, input and output tokens per call and the context at the last step from the gateway log.
  3. Token load: multiply tasks by tokens for the hourly input, the input after prefix cache hits and the output.
  4. Sessions: multiply tasks per second by the task duration for the average number of open agent sessions, apply a peak factor and multiply by the KV cache at the last step, the memory that keeps their prefixes cached between calls.
  5. Test: load-test the candidate GPU with recorded tasks, prefix caching on, and check time to first token per step and the task duration against your target.

With the example task, 120 tasks in the busiest hour mean 960 model calls, 12.3 million input tokens sent, about 2.3 million tokens to compute after cache hits and 288,000 output tokens. At 62 seconds per task, about 2 sessions are open on average and about 4 at a peak factor of 2. Their cache of about 2.8 GiB for gpt-oss-120b is a small share of one card, so here the time per step decides the hardware, not memory.

We build inference servers sized by model size and concurrent users. Send us the tasks per busy hour, the calls per task and the model through the form below, and we return a configuration and quote within one business day.

Which GPUs suit on-premise agents

For a pilot, one DGX Spark Founders Edition with 128 GB runs a model such as gpt-oss-120b and the workflow on a desk. For production, a server with RTX PRO 6000 Server Edition cards carries one copy of that model per card. By the estimate in our private ChatGPT sizing guide, one card keeps about 22.2 GiB of cache beside that model, room for about 33 sessions at the example’s 19,400 tokens.

The H200 NVL leaves about 62.5 GiB beside the same model, about 93 such sessions. Its 4.8 TB/s of memory bandwidth, against 1,597 GB/s on the RTX PRO 6000 Server Edition in NVIDIA’s specifications, shortens the generation part of every step. NVLink bridges join two or four H200 NVL when a larger model must be split.

Small models for routing, embeddings and reranking fit a 24 GB MIG instance of an RTX PRO 6000 or an L4, outside the main model’s batch. Where agents run business processes, two servers that each carry the peak remove the single point of failure, with session affinity so that each agent keeps its cache.

Our Private AI/ML service includes agent workflows and ITOps automation with n8n and Ansible. Describe the first agent workflow you want to run and the tools it will call in the form below.

What we supply

We supply the DGX Spark Founders Edition, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, the H200 NVL with NVLink bridges, and the L4, L40S and smaller RTX PRO cards for routing and RAG models. They come as cards or in AI servers built to order with 2 to 8 GPUs per node, burn-in tested, with manufacturer warranty, on one EU contract and invoice. We check the rack, power and airflow before we quote. Models, agent workflows and the platform on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What hardware do AI agents need?
AI agents need the same GPU server as a chat assistant for the same model, sized for more tokens and longer sessions. Each agent task makes several model calls that resend the growing context, so the server must handle the prefill of those prompts, hold a KV cache per open session and keep each step fast enough that the whole task finishes in time. Size it from tasks per busy hour, calls per task and tokens per call rather than from the number of users.
Why do agentic AI workloads need more GPU than chat?
A chat turn is one model call, while an agent task is a chain of calls in which every call sends the system prompt, the tool definitions, the history and all tool results again. Anthropic’s engineering blog of 13 June 2025 states that in its data “agents typically use about 4× more tokens than chat interactions”, and multi-agent systems about 15× more than chats. In our illustrative example with assumed values, one task of 8 calls sends 102,400 input tokens against 3,500 for one chat turn, a wider gap because it sets a whole task against a single turn.
How many model calls does an AI agent make per task?
It depends on the workflow and the iteration limit. NVIDIA’s blog of 8 May 2026 reports 21.0 or 41.7 tool calls per task for a coding agent on a 50-task subset of SWE-Bench Verified, and NVIDIA’s Dynamo documentation says coding agents make hundreds of API calls per session. n8n’s AI Agent node allows 10 model runs per prompt by default, so a helpdesk agent in n8n stays far below coding-agent counts unless the limit is raised.
Does prefix caching help AI agents?
Yes, because each agent call begins with the same tokens as the previous one. vLLM has enabled automatic prefix caching by default since V1, and NVIDIA’s Dynamo documentation cites 85 to 97 per cent cache hits per call after the first for a hosted coding agent on a managed cloud API. It shortens only the prefill, not generation, and cached blocks can be evicted while the agent waits for a tool.
Can a local LLM server handle tool calling for agents?
Yes, vLLM parses tool calls when started with --enable-auto-tool-choice and the tool call parser of the model family, and Ollama supports tool calling for models tagged for tools. In auto mode vLLM constrains tool calls only for tools that set strict: true or when the server raises --tool-strict-level, and otherwise reads them from raw text, where arguments can be malformed; the choice of tool always stays with the model. Test candidate models with your own tools and count wrong tool choices and malformed arguments.
How many AI agents can one GPU run?
In our illustrative example with assumed values, memory beside gpt-oss-120b allows about 33 agent sessions of 19,400 tokens, 0.67 GiB of cache each, on one RTX PRO 6000 and about 93 on one H200 NVL. Latency per step usually sets a lower limit, because every call in a task waits in the same queue. Load-test the card with recorded agent tasks and prefix caching on before you fix the number.

Send us the agent workflows you plan, the tasks in the busiest hour, the model, the tools each agent calls and your response targets. We reply within one business day with a configuration and a quote, and we check the rack, power and airflow before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna