BLOG · GUIDE ·

n8n with a local LLM: on-premise AI agents on vLLM or Ollama, tool calls and approval steps

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • n8n’s AI Agent node uses a chat model sub-node: the OpenAI Chat Model with an OpenAI credential whose Base URL points to vLLM’s /v1 address and whose API Key is vLLM’s key, or the Ollama Chat Model with an Ollama credential, whose Base URL defaults to http://localhost:11434
  • vLLM needs --enable-auto-tool-choice, which its documentation marks mandatory for auto tool choice, and the --tool-call-parser of the model family, such as openai for gpt-oss and hermes for Qwen3 and Qwen2.5; in n8n, switch off Use Responses API, on by default in version 1.3 of the OpenAI Chat Model
  • A Human review step in the AI Agent’s Tools Panel pauses the workflow until a person approves or denies a tool call in n8n’s chat, Slack, Telegram, Gmail or another of nine channels; approval runs the tool with the model’s arguments, and a denial is reported back to the agent
  • By default n8n keeps executions, with each node’s input and output, for 336 hours and at most 10,000 of them; log streaming and external secrets are Enterprise features, and from n8n 2.0 the Execute Command node is disabled by default
  • The Sustainable Use License allows use and modification “only for your own internal business purposes or for non-commercial or personal use”; files marked .ee. fall under the n8n Enterprise License, and n8n does not call itself open source

Eurokommerz × Vixen.UNO: Private AI/ML  Talk to an expert →

Connecting n8n’s AI Agent to a local LLM

n8n connects to a self-hosted model through a chat model sub-node of its AI Agent node: the Ollama Chat Model for an Ollama server, with an Ollama credential whose Base URL defaults to http://localhost:11434, or the OpenAI Chat Model for vLLM and other servers with an OpenAI-compatible API. Since n8n 1.82.0 every AI Agent node works as a Tools Agent, which hands its tools to the model through LangChain’s tool calling interface, so the model must produce tool calls the server can parse. Names and defaults are those of n8n 2.42.3, the latest release on 6 October 2026.

The OpenAI credential has a Base URL field, which its documentation omits and n8n’s source describes as “Override the default base URL for the API”: enter vLLM’s /v1 address there, such as http://gpu01:8000/v1, and the key vLLM was started with as the API Key. Switch off Use Responses API, on by default in version 1.3 of the OpenAI Chat Model, so that n8n uses the chat completion API for which vLLM documents its tool options, and enter the model name vLLM serves.

In Docker, “by default, each container has its own localhost”, so for Ollama in a separate container n8n’s documentation puts the container’s name in the Base URL, as in http://my-ollama:11434, which resolves on a user-defined Docker network. Ollama’s library also lists cloud models, which run “in Ollama’s cloud”, and OLLAMA_NO_CLOUD=1 switches them off together with web search, according to Ollama’s FAQ. An n8n instance shares “selected, anonymous telemetry” with n8n until N8N_DIAGNOSTICS_ENABLED is set to false.

Tool calling on vLLM and Ollama: flags, parsers and models

An agent leaves the choice of tool to the model, which vLLM calls auto tool choice. vLLM’s documentation marks --enable-auto-tool-choice as mandatory for it, with a --tool-call-parser that matches the model family and is “used to parse the model-generated tool call into OpenAI API format”.

MODEL FAMILYVLLM PARSERDOCUMENTED CAVEAT
gpt-oss-20b, gpt-oss-120bopenaiharmony format only: the models “will not work correctly otherwise”
Qwen3, Qwen2.5, QwQ-32Bhermesfor Qwen3, Qwen’s guide adds a reasoning parser
Qwen3-Coder, Coder-Nextqwen3_xml or qwen3_codertwo names for one parser in vLLM 0.31.0
Llama 3.1, 3.2llama3_json with a chat templateno parallel tool calls; arrays can arrive as strings
Llama 4llama4_pythonic with a chat templateparallel tool calls supported
Mistral 7B Instruct v0.3mistral“struggles to generate parallel tool calls correctly”

vLLM 0.31.0 tool calling documentation and parser registry, Qwen’s function calling guide and the gpt-oss-120b and Qwen3-Coder-Next model cards, read on 6 October 2026.

The parser only converts the model’s output; the choice of tool and arguments stays with the model. With a named tool or required tool choice, vLLM constrains the output with structured outputs, which yield “a validly-parsable function call - not a high-quality one”. vLLM 0.31.0 adds --tool-strict-level, which at function constrains the markup and function name of every tool call and at parameter also its arguments. Qwen’s guide warns that generation may not “always follow the protocol even with proper prompting or templates”. Test candidates in a pilot with your own tools and count wrong tool choices and malformed arguments.

On Ollama, tool calling needs no server flag; it depends on the model, and Ollama’s library tags models that support tools. Ollama’s defaults for parallel requests and context differ from vLLM’s, as our comparison of vLLM, SGLang, TensorRT-LLM and Ollama shows.

Our Private AI/ML service deploys open and commercial models on-premise with vLLM, Ollama or NVIDIA AI Enterprise. Tell us which model your workflows should call and which tools it will need.

Context length, memory and iteration limits

Each agent step sends the system message, the tool definitions, the conversation so far and every tool result to the model, so long tool outputs fill a context quickly. On vLLM the limit is --max-model-len; on Ollama it is the server’s context length, which Ollama’s documentation recommends setting to at least 64,000 tokens for agents. n8n’s Ollama Chat Model has its own Context Length option, 2,048 tokens by default once added, so give it the server’s value or leave it out. Our guide to a private coding assistant covers how far agent prompts grow and the GPU memory per session. The AI Agent’s Max Iterations option, 10 by default, caps how often the model runs for one prompt.

Simple Memory passes the last five interactions by default and holds them in the n8n process, where n8n’s source drops a session after an hour without access; in queue mode it “doesn’t work in an active production workflow”. Postgres and Redis Chat Memory keep the history in a database, keyed for example by ticket number.

Human approval before an agent acts

In the AI Agent’s Tools Panel, a Human review section takes an approval channel and the tools that need approval, and “When a tool requires human review, the workflow pauses and waits for a person” to approve or deny. Approval runs the tool “with the input specified by the AI”; on a denial “the action is canceled and the AI is informed of the rejection”. The request goes to n8n’s built-in chat, Slack, Telegram, Gmail or another of nine channels, and the message can show $tool.name and $tool.parameters. n8n advises naming these tools, and the response to a denial, in the system message.

In the workflow itself, nodes such as Gmail have a Send and Wait for Approval operation, whose Type of Approval is Approve Only or Approve and Disapprove. Its approval buttons are links, and n8n’s Slack documentation notes of such buttons that “anyone who has the link can respond, and n8n can’t tell who clicked”. In Slack’s Send and Wait for Response, Capture Who Responded switches to Slack’s interactive buttons, records the approver and makes Restrict Who Can Approve available, but needs the n8n instance “reachable from Slack over public HTTPS”. Send-and-wait operations and the Wait node, which n8n suggests for “more complex approvals”, have a Limit Wait Time option that resumes a waiting execution after a limit, so route that case to the refusal branch. An approval sent through Slack or Gmail leaves your network with the arguments it shows; n8n’s built-in chat stays on your instance.

A ticket or an email can carry instructions aimed at the model, so tools that change systems belong behind a review step, as our article on prompt injection and LLM security explains.

Credentials and what each tool may do

Give every tool its own credential with only the rights its task needs: a read-only token for monitoring and ticket queries, a separate token for ticket comments, an SSH key for the Ansible control node. n8n encrypts credentials “before they get saved to the database” with a key it creates on first launch in ~/.n8n; set it yourself with N8N_ENCRYPTION_KEY, as queue mode requires for every worker.

Workflow sharing, on the Business and Enterprise plans of self-hosted n8n, “allows editors to use all credentials used in the workflow”, including credentials not shared with them, so limit the editors of a workflow that holds a write credential. External secrets, loaded from a store such as HashiCorp Vault, are an Enterprise feature.

From n8n 2.0 the Execute Command node is disabled by default, and under Docker it would run commands “in the n8n container and not the Docker host”. Run Ansible on a control node through the SSH node instead, with a key that OpenSSH’s restrict and command= options tie to one wrapper script: sshd ignores the command n8n sends and passes it in SSH_ORIGINAL_COMMAND, where the script checks playbook and host against an allow-list.

ITOps example: an Ansible runbook behind an approval

Ticket triage needs no tool that writes: the AI Agent reads the ticket, calls read-only tools such as an HTTP Request Tool and returns category, priority and a draft reply through a structured output parser, and an ordinary HTTP Request node writes them to the ticket. A runbook that changes a server adds a check and an approval:

  1. A Webhook node with Header auth or JWT auth and an IP allowlist receives the alert.
  2. The AI Agent queries read-only tools and proposes a playbook from the allow-list in its system message, the target host and its reasoning.
  3. The workflow checks the proposal against the allow-list, and the SSH node runs the playbook on the control node with --check --diff and --limit for that host.
  4. A send-and-wait message gives the system’s owner the proposal and the diff, with Approve and Decline buttons.
  5. On approval the SSH node runs the playbook without --check; a refusal or an expired wait ends with a ticket comment.
  6. An HTTP Request node writes the result and the n8n execution ID to the ticket.

In check mode, Ansible “runs without making any changes on remote systems”, but “Modules that do not support check mode report nothing and do nothing”, so the diff can understate the change. A task set to check_mode: false changes the system even under --check, so review allow-listed playbooks for it, and diff: false keeps secrets in a task’s diff out of the approval message. Where the agent itself must act during a run, put that tool behind a Human review step.

WORKFLOW STEPIN N8NOUTSIDE N8N
TriggerWebhook: Header or JWT auth, IP(s) Allowlistonly the fields the agent needs
Model callvLLM’s key in the credential; Max IterationsvLLM on an internal address
Reading dataa read-only credential per tooltokens with read scope only
Changecheck mode, then apply after approvalforced command and allow-list
ApprovalRestrict Who Can Approve in Slack; Limit Wait Timenamed approvers; a timeout counts as a refusal
Recordexecutions kept 336 hours by defaultresult and execution ID in the ticket

n8n documentation, Ansible’s check mode page (updated 5 October 2026) and OpenSSH’s sshd manual, read on 6 October 2026.

Agent workflows and ITOps automation with n8n and Ansible are part of our Private AI/ML service. Describe in the form below the first routine task you would hand to an agent and who approves it today.

Logging: what n8n and vLLM keep

n8n saves every execution, successful or failed, with each node’s input and output, the agent’s included. By default, executions older than 336 hours are deleted (EXECUTIONS_DATA_MAX_AGE) and at most 10,000 are kept (EXECUTIONS_DATA_PRUNE_MAX_COUNT), so set the age to your retention period, with a count that covers it, or copy what must be kept into the ticket. Log streaming, an Enterprise feature, sends events such as Tool called and LLM generated to a syslog server, a webhook or Sentry.

vLLM writes no prompts to its log by default, because --enable-log-requests is off; switched on, it logs request IDs and parameters at INFO level and prompts only at DEBUG. A gateway in front of vLLM gives each workflow or front end its own key and request log, as our guide to a private LLM platform on Kubernetes describes.

The n8n licence: Sustainable Use License and internal use

n8n “uses the Sustainable Use License and n8n Enterprise License”, and the first reads: “You may use or modify the software only for your own internal business purposes or for non-commercial or personal use.” n8n’s licence FAQ adds that “If workflows are only created or modified by you or people in your organization, you can use n8n under the Community license …”, and its licence page allows consulting and support services “(e.g. building n8n workflows) without the need for a separate license agreement”.

The FAQ rules out hosting n8n for clients to build workflows, letting external end users build or configure workflows through your product, “whether via a custom UI, our API, MCP, or an AI agent acting on the user’s behalf”, and white-labelling. Files with .ee. in their name fall under the n8n Enterprise License, and SSO, log streaming, external secrets, projects and sharing are paid features. n8n does not call itself open source and explains that, according to the Open Source Initiative, “open source licenses can’t include limitations on use”. Whether a planned use counts as internal business use is a legal assessment for your legal department.

What we do

Our Private AI/ML service includes agent workflows and ITOps automation with n8n and Ansible, so that routine operations run under rules you control. Models run on-premise with vLLM, Ollama or NVIDIA AI Enterprise, or on dedicated hardware in a Tier-3 data centre in Lithuania; nothing goes to public APIs unless you enable hybrid mode for a specific task, and queries and answers are logged. We start with a pilot on one process with clear metrics and scale only what has proved its value. Eurokommerz holds the contract and supplies the hardware, with engineering by our partner Vixen.UNO; the first call is free of charge, and the price of the technical assessment is fixed before work begins. Data handling during the project is set out on our security and compliance page.

FAQ

How do I connect n8n to a local LLM?
Attach a chat model sub-node to n8n’s AI Agent node: the Ollama Chat Model with an Ollama credential for an Ollama server, or the OpenAI Chat Model with an OpenAI credential whose Base URL points to an OpenAI-compatible server such as vLLM, for example http://gpu01:8000/v1. The AI Agent passes its tools to the model, so choose a model for which your server documents tool calling. Prompts and tool results then stay on your network unless a tool, an approval channel or a cloud model sends them elsewhere.
How do I use Ollama with the n8n AI Agent?
Add an Ollama Chat Model sub-node to the AI Agent and create an Ollama credential, whose Base URL defaults to http://localhost:11434 and whose optional API Key is meant for an authenticating proxy in front of Ollama; if n8n and Ollama run in separate containers, replace localhost with the Ollama container’s name, as in http://my-ollama:11434. Pick a local model that Ollama’s library tags for tools and set the server’s context to at least 64,000 tokens, as Ollama’s documentation recommends for agents. Add the node’s Context Length option only with a value of your own, because it defaults to 2,048 tokens.
How do I connect n8n to vLLM?
Start vLLM with --enable-auto-tool-choice, the --tool-call-parser for your model family and an --api-key, then create an OpenAI credential in n8n with that key and a Base URL such as http://gpu01:8000/v1. In the OpenAI Chat Model node, switch off Use Responses API, which version 1.3 of the node turns on by default, and enter the model name that vLLM serves. vLLM’s documentation marks --enable-auto-tool-choice as mandatory for auto tool choice, the mode an agent relies on.
Which local LLMs support tool calling in n8n?
Use a model family for which your server documents a tool parser: vLLM 0.31.0 lists openai for gpt-oss, hermes for Qwen2.5, qwen3_xml for Qwen3-Coder, llama3_json for Llama 3.1 and 3.2, llama4_pythonic for Llama 4 and mistral for Mistral, Qwen’s guide uses hermes for Qwen3, and Ollama tags tool-capable models in its library. A matching parser reads the call, while the model decides whether it is the right one: vLLM lists known issues such as Llama 3 serialising arrays as strings, and Qwen’s guide warns that a model may not always follow the tool-call protocol. Test candidates in a pilot with your own tools and count wrong tool choices and malformed arguments.
Can an n8n AI agent ask for human approval before it acts?
Yes, through the Human review section in the AI Agent’s Tools Panel, which pauses the workflow before the tools connected to it run, until a person approves or denies in n8n’s chat, Slack, Telegram, Gmail or another of nine channels. Approval runs the tool with the arguments the model chose, and a denial cancels it and is reported back to the agent. Outside the agent, send-and-wait operations and the Wait node add an approval to any workflow step.
Can a company self-host n8n for internal use?
Yes, under n8n’s Sustainable Use License, which allows use and modification “only for your own internal business purposes or for non-commercial or personal use”; n8n’s licence FAQ adds that the Community licence applies when workflows are only created or modified by people in your organisation. It rules out hosting n8n for clients to build workflows, letting external users build workflows through your product and white-labelling n8n. SSO, log streaming, external secrets and the sharing of workflows and credentials are not in the Community edition and come with paid plans.

Send us the routine task you would hand to an agent, the systems it reads and changes, who approves changes today and the model you have in mind. We reply within one business day with next steps, starting with a first call that leaves you with two or three possible solution scenarios. The first call is free of charge.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry – see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna