01
What is prompt injection?
Prompt injection is an attack on an application built on a language model in which an attacker’s text, typed directly or hidden in content the model reads, is taken as instructions. Because the model reads its developer’s instructions and untrusted text in one stream, it can follow the attacker instead of its developer.
NIST’s definition names the mechanism: “An attack which exploits the concatenation of untrusted input with a prompt constructed by a higher-trust party such as the application designer” 1. OWASP puts prompt injection first in its 2025 list of the most serious risks for applications built on large language models, as LLM01:2025, and describes a vulnerability that “occurs when user prompts alter the LLM’s behavior or output in unintended ways” 2.
The name was proposed in September 2022, by analogy with SQL injection, a classic flaw in which a program builds a database query by pasting user input into its own code 3. The analogy explains the attack well. As section 03 shows, prompt injection lacks the clean fix that SQL injection has.
02
What is the difference between direct and indirect prompt injection?
In a direct prompt injection, the attacker types the instruction into the application themselves. In an indirect prompt injection, the attacker plants it in content the model will read later, such as a web page, a document or an email, and the person using the application may never see it.
OWASP draws the same line. Direct injections “occur when a user’s prompt input directly alters the behavior of the model in unintended or unexpected ways”, while indirect injections “occur when an LLM accepts input from external sources, such as websites or files” 2. A typical direct attack tells the model to ignore its instructions, either to make it do something else, which Perez and Ribeiro call goal hijacking, or to reveal the instructions themselves, which they call prompt leaking 4.
Indirect injection is the larger problem for assistants that read the web, because retrieval-augmented generation places text from outside directly in the model’s context. Greshake and colleagues showed in 2023 that attackers can “remotely (without a direct interface) exploit LLM-integrated applications by strategically injecting prompts into data likely to be retrieved”, and demonstrated the attacks against real-world systems as well as applications built for testing 5. The attacker needs no access to the assistant or to its user, only to something the assistant will read. Figure 1 traces both paths.
- A model reads one context. The app’s instructions, the user’s message and any page it retrieves are joined into a single stream of text before the model reads any of it.
- Direct injection: the person typing is the attacker. Their message tells the model to drop the app’s instruction, and nothing in the text marks it as less trustworthy than the app’s own words.
- Indirect injection: the attacker never talks to the model. They plant an instruction where a model is likely to read it later, such as a web page, a document or an email.
- An ordinary request brings the planted line in. The user asks for a summary, the assistant fetches the page, and the planted line enters the context beside the real reviews.
- The model may follow the planted line as an instruction. The summary misreports the page, and because it cites the page, it looks sourced. The user never saw the line that changed it.
Planted text doesn’t have to be visible to people. It can be white text on a white background, a comment in a page’s code, or any other text that a page’s extraction picks up but a reader overlooks. What matters is whether the text reaches the model, not whether a person would notice it.
03
Why can’t a model simply ignore injected instructions?
Because a language model has no reliable way to tell instructions from data. Everything in its context, the developer’s prompt, the user’s message and any retrieved text, reaches it as one sequence of tokens, and following instructions written in text is what it was trained to do.
Hines and colleagues describe the root of the problem. Applications combine several inputs “by concatenating them together into a single stream of text”, and the model “is unable to distinguish which sections of prompt belong to various input sources” 6. Greshake and colleagues put it more briefly: applications built on language models “blur the line between data and instructions” 5.
SQL injection was solved structurally. Parameterised queries send the command and the user’s data through separate channels, so the database never runs data as code 3. A language model has no equivalent separation. Delimiters, labels and warnings such as ‘the following text is untrusted’ are themselves more text, which a well-written injection can imitate or argue against. Research systems are building a separate channel, as section 05 describes, but none is yet standard.
Behaviour is also statistical. A defence that blocks an attack most of the time can still fail on a rephrased version, and attackers can try as many phrasings as they like. OWASP’s guidance is frank about this: “Given the stochastic influence at the heart of the way models work, it is unclear if there are fool-proof methods of prevention for prompt injection” 2.
04
What can a prompt injection attack achieve?
What an injection can achieve depends on what the application lets the model do. Against a model that can only write text, it can distort answers or leak the model’s instructions; against one that can read private data, send messages or take actions, it can steal data or act in the user’s name.
Greshake and colleagues group the consequences into categories that include data theft, “worming”, in which an injection spreads itself to other users or systems, and contamination of the information people receive. Their demonstrations showed that processing retrieved prompts can manipulate an application’s functionality and “control how and if other APIs are called” 5. Table 1 sets out how the risk grows with what the model is allowed to do.
| If the model can | An injection can cause | Example |
|---|---|---|
| Only write answers | Misleading or distorted answers | A summary that calls every review positive |
| Read private data | Leaks of that data | The user’s details placed in a link to the attacker’s site |
| Browse or call tools | Requests the user didn’t ask for | A visit to a page that carries further instructions |
| Act for the user | Actions the user didn’t approve | A message sent, a purchase made or a file deleted |
Agents raise the stakes, because one system then reads untrusted text, holds private data and can act. AgentDojo, an evaluation environment with 97 realistic tasks, such as managing an email client or making travel bookings, and 629 security test cases, found that existing attacks “break some security properties but not all”, and that capable models failed many tasks even with no attack present 7.
A model that can only write answers limits the damage but doesn’t remove it. A distorted answer that cites its source looks trustworthy, and a person may act on it.
05
How do you defend against prompt injection?
No single defence is reliable, so applications combine several: mark untrusted text, train models to rank instructions by their source, keep data from changing what the system does, limit what the model can do, and require a person’s approval for consequential actions. OWASP’s guidance recommends measures of this kind, together with input and output filtering and adversarial testing 2.
- Mark untrusted text. Spotlighting transforms retrieved text, for example by marking or encoding it, to give the model “a reliable and continuous signal of its provenance”. In its authors’ tests it cut the success rate of indirect attacks from over 50% to under 2%, with little effect on the task 6. Like every defence inside the prompt, it changes the odds rather than the architecture.
- Train the model to rank instructions. The instruction hierarchy trains a model to give its developer’s instructions priority over lower-priority text and to ignore conflicting lower-priority instructions; its authors report that this “drastically increases robustness”, even against attack types not seen in training 8. StruQ goes further, separating the prompt and the data into two channels and training the model to follow instructions only in the prompt channel 9.
- Keep data from changing the plan. CaMeL takes the system’s control flow from the user’s trusted request alone, so “the untrusted data retrieved by the LLM can never impact the program flow”, and it enforces security policies when tools are called 10. This kind of defence lives in the code around the model, not in the model.
- Limit what the model can do. Least privilege, a long-standing security principle, applies directly: an assistant that can’t send email can’t be tricked into sending it. OWASP recommends enforcing privilege control and least-privilege access 2.
- Ask a person before consequential actions. OWASP recommends requiring human approval for high-risk actions 2. Approval protects people only if the request is shown in plain terms and appears rarely enough that they don’t approve by habit.
- Filter, then test. Classifiers that flag likely injections, and checks on what the model produces, catch known patterns. Adversarial testing shows which attacks still work; AgentDojo is one public environment for it 7.
Each layer reduces the risk and none removes it. Measures in the code around the model, such as least privilege, confirmation and separating the plan from the data, limit what an injection can cause even when the model is fooled.
06
Why should text on a page be treated as data, not instructions?
Because the author of a web page is not the user. Treating retrieved text as data means a system may quote it, summarise it and reason about it, but never takes orders from it: instructions come only from the person using the assistant and from the assistant’s developer.
The rule is simple to state and hard to enforce, because the model itself can’t be relied on to keep it, for the reasons in section 03. In practice it becomes several mechanisms working together. Retrieved text is labelled and kept apart from instructions, the model is trained to ignore instructions found in data, consequential actions are planned from the user’s request alone, and anything irreversible waits for the person to confirm it.
The rule also gives evaluation a clear test. A system that keeps it should behave the same whether or not a page it reads contains planted instructions, so any change in behaviour caused by planted text counts as a failure. Our explainer on AI safety evaluations describes how tests of this kind fit into a wider evaluation.
07
How does Brello approach prompt injection?
Brello 1.0 limits what an injection could do rather than claiming to stop one: its model can only write answers, and an answer that uses the web shows its sources as cards, with inline citations the model is asked to add. For Brello Super Intelligence, which is in development, we are designing four layers around the rule that text on a page is data, not instructions.
The four layers, and how we intend to test them before release by seeding pages with instructions and counting any change in behaviour as a failure, are described in ‘Evaluate first, then ship’. Our safety page lists every guard in Brello 1.0, and ‘Answering from the open web, without a server’ describes how Brello 1.0 reads and ranks web pages.
References
Reviewed . Web pages were checked on that date.
- National Institute of Standards and Technology. “Prompt injection.” Computer Security Resource Center glossary, citing NIST AI 100-2e2025. Accessed 5 October 2026. csrc.nist.gov/
glossary/ term/ prompt_injection - OWASP GenAI Security Project. “LLM01:2025 Prompt Injection.” OWASP Top 10 for LLM Applications 2025. Accessed 5 October 2026. genai.owasp.org/
llmrisk/ llm01-prompt-injection - Willison, S. (2022). “Prompt injection attacks against GPT-3.” simonwillison.net, 12 September 2022. Accessed 5 October 2026. simonwillison.net/
2022/ Sep/ 12/ prompt-injection - Perez, F. and Ribeiro, I. (2022). “Ignore Previous Prompt: Attack Techniques For Language Models.” arXiv preprint. arxiv.org/
abs/ 2211.09527 - Greshake, K. et al. (2023). “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.” Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec 2023). arxiv.org/
abs/ 2302.12173 - Hines, K. et al. (2024). “Defending Against Indirect Prompt Injection Attacks With Spotlighting.” arXiv preprint. arxiv.org/
abs/ 2403.14720 - Debenedetti, E. et al. (2024). “AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents.” NeurIPS 2024 Datasets and Benchmarks Track. arxiv.org/
abs/ 2406.13352 - Wallace, E. et al. (2024). “The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.” arXiv preprint. arxiv.org/
abs/ 2404.13208 - Chen, S. et al. (2025). “StruQ: Defending Against Prompt Injection with Structured Queries.” Proceedings of the 34th USENIX Security Symposium. arxiv.org/
abs/ 2402.06363 - Debenedetti, E. et al. (2025). “Defeating Prompt Injections by Design.” arXiv preprint. arxiv.org/
abs/ 2503.18813
Version history
- 1.0First published.



