MCP server security for a private LLM: the Model Context Protocol, OAuth 2.1 and tool risks
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- The Model Context Protocol (MCP) connects an LLM application, the host, through one MCP client per server to servers that offer tools, resources and prompts over JSON-RPC 2.0; revision 2026-07-28, current as of October 2026, removed the handshake and sessions and deprecated sampling, roots and logging
- Anthropic announced on 9 December 2025 that it was donating MCP to the Linux Foundation’s Agentic AI Foundation; the project is a Series of LF Projects, LLC, its maintainers take part as individuals, and changes to the specification go through SEPs
- Authorisation is optional and defined for HTTP transports: the MCP server is an OAuth 2.1 resource server that must publish protected resource metadata (RFC 9728), accept only tokens issued for itself and never pass a client’s token on
- Tool descriptions are text the model reads, so a server can hide instructions in them or change them after approval; OWASP’s LLM01:2026 asks to pin, sign and verify every MCP server and audit tool descriptions for hidden instructions
- On vLLM’s chat API the server parses the model’s tool calls and the caller runs them, so the MCP client is the front end or workflow tool, such as Open WebUI (Streamable HTTP since v0.6.31) or n8n’s MCP Client Tool, where allow-lists, scoped credentials, approvals and logs belong
Eurokommerz × Vixen.UNO: Private AI/ML Talk to an expert →
What the Model Context Protocol is and who maintains it
The Model Context Protocol (MCP) is an open protocol that connects an LLM application to tools and data. The application, the host, runs one MCP client per server, and each server offers tools, resources such as files or database schemas, and prompt templates over JSON-RPC 2.0. As of October 2026 the current revision of the specification is 2026-07-28. For a private LLM, MCP links the model to internal systems, so its security depends on which servers you allow and what each may do with whose credentials.
Anthropic introduced MCP and announced on 9 December 2025 that it was donating the protocol to the Agentic AI Foundation, a directed fund under the Linux Foundation in Anthropic’s description. The project’s governance page names it Model Context Protocol a Series of LF Projects, LLC, with lead maintainers who hold final decision authority and core maintainers who steer the specification. Membership is personal rather than per company, and specification changes go through Specification Enhancement Proposals (SEPs).
The 2026-07-28 revision, published on 28 July 2026, made the protocol stateless: the initialize handshake and protocol-level sessions are gone, each request carries its protocol version and client capabilities, and servers must answer server/discover. Revisions up to 2025-11-25 use the handshake, so clients that also talk to older servers handle both.
Hosts, clients and servers: tools, resources and prompts
Tools are model-controlled, meaning the model discovers them and decides when to call them, so they are the feature to restrict first. Resources are application-driven, with the host deciding what enters the context, and prompts are user-controlled templates, offered for example as slash commands, with content from the server. Sampling, through which a server could request a completion from the client’s model, is deprecated in 2026-07-28, as are roots and logging.
The specification sets principles that MCP cannot enforce itself: users consent to all data access and operations, hosts obtain explicit consent before invoking any tool, and tool annotations are untrusted unless they come from a trusted server. Its tools page asks for “a human in the loop with the ability to deny tool invocations”, and clients should show tool inputs before a call, validate results before passing them to the model and log tool use.
Transports: stdio for local servers, Streamable HTTP for remote ones
With stdio, the client starts the server as a subprocess and exchanges newline-delimited JSON-RPC messages over stdin and stdout. The local server security guide on modelcontextprotocol.io states that the client and a stdio server “share one trust domain”: the server runs with the user’s rights, inherits the environment variables, can read SSH keys and cloud credentials and can open outbound connections anywhere. The specification’s authorisation flow is not meant for stdio servers, which take credentials from the environment.
With Streamable HTTP, the server is an independent process with one endpoint that accepts POST and answers with JSON or a Server-Sent Events stream. Servers must validate the Origin header against DNS rebinding, should bind to 127.0.0.1 when they run locally and should authenticate all connections; the guide adds that binding to 127.0.0.1 is no authentication boundary, since other local processes and, through the browser, websites can reach the port. Since 2026-07-28 there is no Mcp-Session-Id header, and the Mcp-Method and Mcp-Name headers, which servers must check against the body, let a gateway route and inspect calls without parsing it; one that enforces policy on them should reject older protocol versions. The older HTTP+SSE transport is deprecated.
MCP authorisation: OAuth 2.1 and protected resource metadata
The specification’s “Authorization” section makes authorisation optional and defines it for HTTP-based transports. Where it is used, the MCP server is an OAuth 2.1 resource server (the specification cites draft-ietf-oauth-v2-1-13) and the MCP client is the OAuth client. Every such server must publish OAuth 2.0 Protected Resource Metadata (RFC 9728) naming at least one authorisation server, announced in the WWW-Authenticate header of a 401 response or at a well-known URI such as /.well-known/oauth-protected-resource.
The client runs the authorisation code flow with PKCE and must send the resource parameter (RFC 8707) with the MCP server’s canonical URI in the authorisation and token requests. The server must check that each token was issued for it and must not “accept or transit any other tokens”. Without pre-registration, the 2026-07-28 revision recommends Client ID Metadata Documents and deprecates Dynamic Client Registration (RFC 7591), and clients must check a returned iss value against the issuer they recorded (RFC 9207).
For least privilege, the server names the scopes an operation needs in its challenge, rejects a token without them with 403 insufficient_scope and keeps scopes_supported minimal. For companies, the optional Enterprise-Managed Authorization extension lets the corporate identity provider decide, through an ID-JAG, which MCP servers an employee may use, so access is granted and revoked centrally where clients and authorisation servers support it.
MCP security best practices: confused deputy, token passthrough, SSRF
The security best practices for 2026-07-28 describe attacks on MCP implementations, many in the OAuth flows. A confused deputy arises at an MCP proxy server that reaches a third-party API with one static client ID. With dynamic client registration on the proxy and a consent cookie at the third-party authorisation server, a crafted link can send an authorisation code to an attacker without a new consent screen, so the proxy must collect consent per client and match redirect URIs exactly. Token passthrough, forwarding a client’s token unchanged to a downstream API, is forbidden, since it bypasses the server’s controls and hides which client made a call.
Server-side request forgery abuses OAuth discovery, in which the client fetches URLs that a malicious server controls and that can point to internal hosts or the cloud metadata address 169.254.169.254, so clients deployed on a server should block private ranges and consider an egress proxy. State handle hijacking replaces session hijacking (covered in the page’s 2025-11-25 version); servers must not treat possession of a state handle as authentication and should bind each handle to the authenticated user. Further sections cover local server compromise, authorisation URL validation, stdio proxies, mix-up attacks, localhost redirect URI impersonation, trust policies for Client ID Metadata Documents and scope minimisation.
Tool poisoning, rug pulls and injection through tool results
Tool names and descriptions are part of the text the model reads, so a server’s tool list is an injection surface before any tool runs. The local server guide calls hidden directives in them tool poisoning, or tool shadowing when they target another server’s tools, and a definition changed after approval a rug pull, which neither a one-time consent dialog nor a pinned package catches. All configured servers share the model’s context, so each new server widens what the others are exposed to.
The OWASP GenAI Security Project’s cheat sheet A Practical Guide for Securely Using Third-Party MCP Servers, version 1.0 of 23 October 2025, defines tool poisoning as “a form of indirect prompt injection” and recommends a central registry of approved servers, a hash or checksum on tool descriptions, containers, narrowly scoped permissions per task and human approval for actions not performed before.
Tool results are the second channel. A ticket or a web page returned by a tool enters the context like any other text, and OWASP’s entry LLM01:2026 Prompt Injection lists “an MCP server’s output” among the sources of indirect prompt injection. The entry asks to pin, sign and verify every MCP server, audit tool descriptions for hidden instructions and monitor tool composition, and notes that pinning does not stop a payload shipped in the pinned version. OWASP’s separate MCP Top 10, in beta, lists MCP03:2025 Tool Poisoning and MCP09:2025 Shadow MCP Servers. Why no filter stops injection reliably is explained in our article on prompt injection and LLM security.
MCP in the enterprise: risks, examples and first controls
| RISK | EXAMPLE | FIRST CONTROL |
|---|---|---|
| Tool poisoning | a description tells the model to attach a configuration file to each call | read descriptions before approval; drop servers with unrelated instructions |
| Rug pull | the server changes an approved tool’s description | a hash of approved definitions; new approval when they change |
| Injection in tool results | a ticket asks the agent to mail customer data outside | approval before calls that change systems or send data out |
| Over-broad token | one admin-scoped token serves every tool | a token per server with the smallest scope |
| Token passthrough | the server forwards the user’s token to the ticket system | audience check; a separate downstream token |
| Local stdio server | a developer’s server reads ~/.ssh | a container with one project directory; no network unless needed |
| SSRF in OAuth discovery | metadata points the front end at 169.254.169.254 | an egress proxy that admits only approved MCP and authorisation servers |
| Shadow MCP servers | staff add servers to their client configuration | inventory of configurations; an allow-list |
MCP specification 2026-07-28, its security best practices and local server security guide; OWASP’s MCP cheat sheet 1.0, LLM01:2026 and MCP Top 10, read on 6 October 2026; examples and controls are our summary.
For a fleet, the local server guide adds an allow-list of vetted servers with pinned versions, a default isolation posture, central deployment of servers that only wrap a SaaS API, an inventory of client configuration files, an audit log of which server ran which tool and when, and a plan to block a vulnerable server’s egress and rotate its credentials. Finding unapproved AI tools is covered in our article on shadow AI policy and controls.
Our Private AI/ML service includes data and permissions management and a log of queries and answers. Write to us with the internal systems an assistant should reach through MCP and who approves changes to them today.
MCP with a private LLM: vLLM, Open WebUI and n8n
vLLM’s OpenAI-compatible server turns the model’s output into tool calls when started with --enable-auto-tool-choice and the --tool-call-parser of the model family, and its documentation leaves defining the tools and handling the calls to the caller’s application. The MCP client is therefore usually the front end or workflow tool in front of vLLM, which lists the servers’ tools, runs the calls the model chooses and returns the results.
| COMPONENT | MCP ROLE | TRANSPORT | DOCUMENTED DETAIL |
|---|---|---|---|
| vLLM chat completions | parses tool calls; the caller runs them | none | a --tool-call-parser per model family |
| vLLM with gpt-oss | MCP client via --tool-server for /v1/responses | MCP SSE servers per the recipe; HTTP+SSE at http://host:port/sse in 0.31.0 | the demo Python tool runs model code in Docker without network isolation by default |
| Open WebUI | MCP client since v0.6.31 | Streamable HTTP only; mcpo exposes stdio or SSE servers as OpenAPI endpoints | servers added under Admin Settings, External Tools |
| n8n MCP Client Tool | gives an AI Agent the tools of one server | HTTP Streamable; SSE marked deprecated | Tools to Include: All, Selected or All Except |
| n8n MCP Server Trigger | offers n8n tools and workflows to MCP clients | SSE or streamable HTTP; no stdio | None, Bearer auth or Header auth |
vLLM 0.31.0 documentation and source, the gpt-oss recipe (22 September 2026), Open WebUI MCP documentation, n8n documentation and node source of n8n 2.42.3 (latest release), read on 6 October 2026.
When self-hosted, Open WebUI and n8n run their MCP clients on a server inside your network, the case the SSRF guidance addresses; since internal MCP servers have private addresses, put that host behind an egress proxy that admits only the approved MCP and authorisation servers instead of blocking all private ranges. Open WebUI’s documentation also asks for a review of authentication, proxy and rate-limiting policies before MCP is exposed externally, and n8n’s page for the MCP Client Tool still names an SSE endpoint, although the node in n8n 2.42.3 defaults to HTTP Streamable. Approvals and credentials per tool in n8n are covered in our article on AI agents with n8n on a private LLM, sign-in and query logs in Open WebUI in our guide to a private ChatGPT alternative, and vLLM’s own API in our guide to securing a GPU server.
Our Private AI/ML service deploys open and commercial models on-premise with vLLM, Ollama or NVIDIA AI Enterprise. Describe in the form below the tools you want to connect first and the systems behind them.
What we do
Our Private AI/ML service includes agent workflows and ITOps automation with n8n and Ansible under rules you control, protection of models against prompt injection, data and permissions management and logging of queries and answers. Models run on your servers or on dedicated hardware in a Tier-3 data centre in Lithuania, and nothing goes to public services unless you enable it. We start with a pilot on one process with clear metrics and scale only what has proved its value. Eurokommerz holds the contract and supplies the hardware, with engineering by our partner Vixen.UNO; the first call is free of charge, and the price of the technical assessment is fixed before work begins. Data handling during a project is set out on our security and compliance page.
FAQ
What is the Model Context Protocol (MCP)?
Who maintains the Model Context Protocol?
How does MCP authorisation work?
What is tool poisoning in MCP?
How do you secure MCP servers in an enterprise?
Can MCP be used with a local or private LLM?
--enable-auto-tool-choice and a --tool-call-parser for the model family, while Open WebUI, natively over Streamable HTTP since v0.6.31, or n8n’s MCP Client Tool connect to the MCP servers and run the calls. With the model, the front end and the servers on your network, prompts and tool results stay there unless you connect a server outside it or a tool sends data out.Send us the internal systems an assistant or agent should reach through MCP, the model server and front end you run or plan, and who approves changes to those systems today. We reply within one business day with next steps, starting with a first call that leaves you with two or three possible solution scenarios. The first call is free of charge.
Talk to an expertWe reply within one business day