BLOG · GUIDE ·

Prompt injection and LLM security for internal assistants: OWASP Top 10, RAG and agents

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Prompt injection is text that changes what a language model does, typed by a user or hidden in a document, email, web page or tool result the model reads; OWASP lists it first, as LLM01, in its Top 10 for LLM Applications 2026
  • Indirect injection needs no account on the assistant: whoever can write to a wiki, file share, mailbox or web page it reads can plant instructions there, including white text on a white background or text inside an image
  • Permission checks at retrieval keep other people’s documents out of the context, but an injected instruction can still try to send the user’s own data out through a rendered image link, an email or another tool
  • For agents, OWASP’s entry on excessive agency, LLM03 since the 2026 edition, asks for minimal tools and permissions, actions run in the requesting user’s context, authorisation outside the model and human approval for high-impact actions
  • The UK NCSC wrote in December 2025 that prompt injection may never be totally mitigated the way SQL injection can be, and Meta calls it an unsolved weakness in all LLMs, so the design goal is limited damage

Eurokommerz × Vixen.UNO: Private AI/ML  Talk to an expert →

Prompt injection: what it is and what it can do

Prompt injection is an attack in which text that reaches a language model changes what the model does. A user can type it, or it can sit in a document, email, web page or tool result that the assistant reads. OWASP lists it first, as LLM01, in its Top 10 for LLM Applications 2026 and states that “no reliable prevention mechanism exists today”. For an internal assistant, LLM security therefore depends mainly on limiting what an injected instruction can reach: the documents retrieved for the user, the tools the assistant may call and the places its output goes.

The UK National Cyber Security Centre (NCSC) explained the cause on 8 December 2025, writing that inside an LLM “there’s no distinction made between ‘data’ or ‘instructions’”. The system prompt, the question and every retrieved passage reach the model as one sequence of tokens. With access to mailboxes or APIs, an injected instruction can act in the user’s name. The server and its inference API are covered in our guide to securing a GPU server.

Direct and indirect prompt injection

In OWASP’s 2026 text, an injection is direct when a user, or an attacker with the user’s access path, “supplies input that changes model behavior in undesired ways”, and jailbreaking is the subset of prompt injection that aims to make the model violate its safety protocols. Indirect injection arrives when the model “ingests content from an external source” that “contains data which acts as prompt injection”. Whoever can write to a source the assistant reads can plant instructions there: a supplier’s PDF, an editable wiki page, an email to a monitored mailbox, a ticket or a search result. Such inputs “need not be visible in the rendered interface to influence the model”, and one of OWASP’s scenarios hides an instruction in an image below the human visual threshold.

NIST AI 100-2 E2025, NIST’s taxonomy of attacks on machine learning systems of March 2025, treats direct prompting attacks (section 3.3), indirect prompt injection (3.4) and the security of agents (3.5) separately. Its index files both prompt injection (NISTAML.018) and indirect prompt injection (NISTAML.015) under availability violations, integrity violations and privacy compromises in generative AI, and prompt injection also under misuse violations.

OWASP Top 10 for LLM applications 2026

OWASP’s resource page dates the 2026 edition 3 August 2026. Hidden context exposure, which OWASP calls “a broader framework”, replaces system prompt leakage, and excessive agency moves up to third. The table adds each entry’s 2025 number, an example from an internal assistant and the control we would apply first.

2026 IDRISK2025 IDEXAMPLEFIRST CONTROL
LLM01Prompt InjectionLLM01hidden text in a supplier PDF instructs the assistant to email the chat to an outside addressleast privilege for tools; approval for actions
LLM02Sensitive Information DisclosureLLM02the answer quotes a salary file the user may not openauthorisation inside the retrieval query
LLM03Excessive AgencyLLM06an agent with full mailbox rights forwards mail after reading a crafted emailread-only scopes; approval before sending
LLM04Supply ChainLLM03a backdoored adapter, or a pickle file that runs code when loadedverified sources, pinned revisions, file hashes
LLM05Data and Model PoisoningLLM04a planted wiki page steers the answers on one topiclimit who writes to indexed sources; validate before indexing
LLM06Unbounded ConsumptionLLM10a looping agent fills the GPU queue for everyoneper-user token limits; step and time limits for agents
LLM07MisinformationLLM09a confident answer that cites an invented clausecitations with document dates; human review
LLM08Hidden Context ExposureLLM07a user extracts a system prompt that contains an API keyno credentials or access rules in the system prompt
LLM09Vector and Embedding WeaknessesLLM08one collection without access filters serves every departmentaccess filter inside the index query
LLM10Improper Output HandlingLLM05a Markdown image in the answer sends chat content to an external serverencode output; images only from your own domains

OWASP Top 10 for LLM Applications 2026 (PDF; resource page dated 3 August 2026) and the 2025 risk pages on genai.owasp.org, read on 6 October 2026; OWASP writes the IDs as LLM01:2026 and LLM01:2025; examples and first controls are our summary.

OWASP’s separate Top 10 for Agentic Applications for 2026, published on 9 December 2025, starts with ASI01 Agent Goal Hijack, ASI02 Tool Misuse and ASI03 Identity & Privilege Abuse.

RAG security: permission checks and planted documents

Retrieval decides what an assistant can disclose. If the search returns a passage the user may not open, the model can repeat it whatever the system prompt says. OWASP’s LLM02 therefore asks to “enforce document- and chunk-level authorization inside the index query, not at the application layer after retrieval”. Our article on RAG on company data explains why the permission filter belongs in the search, and our comparison of vector databases for RAG shows how pgvector, Qdrant, Milvus and OpenSearch apply it; the identity it filters on comes from the sign-in, never from the prompt.

The indexed documents are the second risk. Anyone with write access to an indexed source can plant instructions, so the list of writers belongs in the threat model. OWASP’s entry on vector and embedding weaknesses, LLM09 in 2026, asks to “strip zero-width characters, white-on-white text, and Unicode homoglyphs at extraction”. Store the source and the last editor with every chunk, so that a suspicious passage can be traced and removed.

Our Private AI/ML service covers protection of models against prompt injection, data and permissions management and logging of queries and answers. Tell us which sources your assistant reads and who can write to them.

Agents with tools: excessive agency and human approval

OWASP’s LLM03 defines excessive agency as the vulnerability “that enables damaging actions to be performed in response to unexpected, ambiguous or manipulated outputs”, caused by excessive functionality, permissions or autonomy. Against these, OWASP recommends giving an agent only the tools it needs, avoiding open-ended ones such as a shell or a fetch of any URL, limiting each tool’s permissions and running actions with the rights of the user who asked. Authorisation is enforced in application logic, “rather than relying on an LLM to decide”, and a human approves high-impact actions. The joint guidelines for secure AI system development, published on 27 November 2023 by the UK NCSC, the US CISA and 21 international partners, ask for “appropriate restrictions to the possible actions” of AI components.

The NCSC takes up a rule from a public discussion that an LLM processing information from a party should drop to that party’s privileges, and advises “don’t let an LLM processing emails from random external people have access to privileged tools”. Meta published an “Agents Rule of Two” on 31 October 2025. Within one session, an agent should combine no more than two of three properties: processing untrustworthy input, access to sensitive systems or private data, and changing state or communicating externally. An agent that needs all three “should not be permitted to operate autonomously”.

An approval protects only if the reviewer sees “the exact rendered action rather than a summary”, in the words of OWASP’s LLM01. In n8n, a Human review step on an agent’s tools pauses the workflow until a person approves or denies the call, and its message can show the tool’s name and the parameters the model chose; our article on AI agents on a private LLM with n8n covers this step and credentials per tool.

Our Private AI/ML service includes agent workflows and ITOps automation with n8n and Ansible, so that routine operations run under rules you control. Describe the actions an agent should take and who approves them today.

Output handling and system prompt leakage

Model output is untrusted input for whatever consumes it next. OWASP’s 2025 page on improper output handling, LLM10 since 2026, says to treat the model “as any other user, adopting a zero-trust approach”, which means encoding output for the browser, using parameterised queries where output feeds a database and validating tool arguments before execution.

In OWASP’s second LLM01 scenario, hidden instructions in a web page make the model insert an image whose URL leaks the private conversation. If the chat front end renders Markdown images, the browser requests that URL as soon as the answer appears. A Content Security Policy whose img-src directive lists only your own domains blocks the request, and links to unknown domains are safer shown as plain text. Open WebUI sets no security headers by default, and its hardening guide, which adds the policy through CONTENT_SECURITY_POLICY, warns that “An overly strict CSP will break the frontend”.

For hidden context exposure, LLM08, OWASP asks to design “under the assumption that hidden context is discoverable”. Credentials, secrets and access rules stay out of the system prompt, and checks such as privilege separation run outside the model.

What cannot be fully solved

The sources agree that no filter or prompt design stops injection reliably. The NCSC wrote that “it’s very possible that prompt injection attacks may never be totally mitigated in the way that SQL injection attacks can be”, and recommends reducing “the risk and the impact” instead. Meta calls prompt injection “a fundamental, unsolved weakness in all LLMs”. OWASP expects some of its own controls to “degrade against adaptive attackers”. NIST’s taxonomy notes that many mitigations against adversarial machine learning attacks “tend to be empirical and limited in nature”.

Detection models lower the risk without removing it. Meta’s Llama Prompt Guard 2 labels a text malicious if it “explicitly attempts to override prior instructions” and, unlike Prompt Guard 1, has no separate label for text that may cause unintentional instruction-following. It reads 512 tokens at a time, so longer texts are split, and was evaluated in eight languages, German among them but not Polish, Czech or Hungarian. Its model card says it focuses on “explicit, known attack patterns” and warns that “adversaries may develop sophisticated attacks specifically to bypass detection”. OWASP’s LLM01 measure to pass external content through “a structurally separate, provenance-labeled channel”, such as a separate message for retrieved text, has the same limit.

An injection will sometimes succeed, so the design has to keep the harm small. Permission filters keep other people’s data out of reach, but they do not stop the exfiltration of what the user may read; that takes closed output channels and approved actions. Whoever owns the use case should accept the remaining risk in writing.

Logging, adversarial testing and a review before go-live

The joint guidelines ask providers to “monitor and log inputs to your system (such as inference requests, queries or prompts)”, and the NCSC suggests logging enough to identify suspicious activity, “potentially including full input and output of the LLM”, as well as tool use and API calls. For an assistant, that is one record per request with the user, the prompt, the IDs of the retrieved passages, each tool call with its arguments and result, each approval and the answer. The log holds personal data and needs access rules and a retention period. Open WebUI’s audit log does not contain the answers, and our guide to a private ChatGPT alternative shows where a complete record comes from.

Someone should read the log at fixed intervals, looking for retrieved passages that address an AI with instructions, tool calls to unusual destinations, rejected approvals and sudden peaks per user. OWASP’s LLM01 also asks to test “against adaptive attackers who have read the deployed defense”. NVIDIA’s open-source garak, which its README calls an “LLM vulnerability scanner”, probes for prompt injection, jailbreaks and data leakage, among others, and can target almost anything reachable over REST. A review before go-live takes seven steps.

  1. List every source the assistant reads and who can write to it.
  2. List every tool, its permissions and whether it can change state or send data out.
  3. Split sessions that combine untrusted input, sensitive data and external actions, or add an approval step.
  4. Keep secrets and access rules out of the system prompt and enforce permissions at retrieval.
  5. Allow images and links in the chat front end only from your own domains.
  6. Turn on the log described above, with a retention period.
  7. Test with planted documents and a scanner such as garak before go-live and after every change.

What we do

Our Private AI/ML service includes protection of models against prompt injection, data and permissions management and logging of queries and answers, so that security and legal see who accesses what. Nothing goes to public services unless you enable it, and what goes there is visible in the query log. The first call is free of charge; the price of the technical assessment is fixed before work begins. Eurokommerz holds the contract and supplies the hardware, with engineering by our partner Vixen.UNO, and how we handle data during a project is set out on our security and compliance page.

FAQ

What is prompt injection?
Prompt injection is an attack in which text that reaches a language model, typed by a user or hidden in a document, email, web page or tool result, changes what the model does. OWASP lists it first, as LLM01, in its Top 10 for LLM Applications 2026. In an internal assistant it can produce misleading answers, disclose data or, where the assistant has tools, trigger actions in the user’s name.
What is indirect prompt injection?
Indirect prompt injection places instructions in content the model processes rather than in the user’s own prompt: a document in the RAG index, an email, a web page or the result of a tool call. The attacker needs no access to the assistant, only write access to something it reads, and the instructions can be invisible to people, for example white text on a white background. NIST AI 100-2 E2025 treats it separately from direct prompting attacks, in its section 3.4.
What is the OWASP Top 10 for LLM applications?
It is the OWASP GenAI Security Project’s list of the main security risks of applications built on large language models; the 2026 edition, dated 3 August 2026, names prompt injection, sensitive information disclosure, excessive agency, supply chain, data and model poisoning, unbounded consumption, misinformation, hidden context exposure, vector and embedding weaknesses and improper output handling, numbered LLM01 to LLM10. Hidden context exposure replaces system prompt leakage, LLM07 in the 2025 edition, and excessive agency moved up from sixth to third place. For agents, OWASP published a separate Top 10 for Agentic Applications for 2026 in December 2025.
Can prompt injection be prevented?
OWASP states that no reliable prevention mechanism exists today, and the UK NCSC wrote in December 2025 that prompt injection attacks may never be totally mitigated in the way SQL injection attacks can be. Filters, classifiers and hardened system prompts lower the success rate, while least privilege for tools, permission checks at retrieval, closed exfiltration channels and human approval for actions limit the damage when an injection succeeds. Logging and regular adversarial tests show whether those limits hold.
How do you secure a RAG system?
Enforce each user’s access rights in the search itself, so that the retriever never returns a passage the user cannot open in the source system, and take the identity from the sign-in rather than from the prompt. Control who can write to indexed sources, validate documents before indexing with text extraction that detects hidden content, and keep the source of every chunk. Treat retrieved text as untrusted, because a planted document can carry instructions, and log which passages were retrieved for whom.
What is excessive agency in LLM applications?
OWASP’s LLM03, numbered LLM06 in the 2025 edition, describes excessive agency as the vulnerability that lets an LLM-based system perform damaging actions in response to unexpected, ambiguous or manipulated outputs, caused by excessive functionality, permissions or autonomy. The controls are fewer and narrower tools, permissions limited to what each task needs, actions run with the requesting user’s rights, authorisation checks outside the model and human approval for high-impact actions. Meta’s Agents Rule of Two adds that an agent should not combine untrusted input, access to sensitive data and external actions in one session without supervision.

Send us a short description of the assistant or agent you run or plan: the sources it reads, the tools it may call and how users sign in. We reply within one business day with next steps, starting with a first call that leaves you with two or three possible solution scenarios. The first call is free of charge.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry – see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna