
What Is Prompt Injection? A Plain-English Defense Guide
What is prompt injection? Learn direct vs indirect attacks, 9 real scenarios, and the 7 OWASP-backed defenses that keep AI assistants and agents safe.
Prompt injection is the security problem every AI assistant has to solve, and the good news is that the industry now has a clear, layered playbook for it. If you use a chatbot that reads your email, an agent that browses the web, or a coding assistant that opens files, understanding prompt injection will help you pick safer tools and set them up well. This guide explains what prompt injection is, how the two main types work, and the specific defenses that OWASP, Microsoft and Google use today.
- Definition: OWASP ranks prompt injection as LLM01 in its 2025 Top 10 for LLM applications, describing it as input that changes a model's behavior in unintended ways.
- Two types: direct injection comes from the user's own prompt; indirect injection hides instructions in content the AI reads, such as a web page, file or email.
- Seven core defenses (OWASP): constrain behavior, validate output formats, filter inputs and outputs, enforce least privilege, require human approval for risky actions, label untrusted content, and test adversarially.
- Big-platform practice: Microsoft and Google both describe layered, defense-in-depth approaches that combine classifiers, content isolation and user confirmation.
Quick Picks: The Five Habits That Matter Most
- Give AI tools the least access they need. An assistant that cannot send email cannot be tricked into sending email.
- Keep a human in the loop for anything irreversible. Payments, deletions, outgoing messages and code merges deserve a confirmation step.
- Treat everything the AI reads as untrusted data. Web pages, PDFs, tickets and emails are information, never commands.
- Prefer tools that block silent data exfiltration. Look for products that refuse to render unknown external images and links.
- Test your own setup. A short red-team session against your own agent teaches more than any checklist.
What Is Prompt Injection, Exactly?
Large language models take in one stream of text and decide what to do next. The developer's instructions, the user's request and any documents the model reads all arrive as words in that same stream. Prompt injection happens when some of those words, often written by a third party, persuade the model to do something its owner never intended.
OWASP's GenAI Security Project defines prompt injection as user input that alters a model's behavior or output in unintended ways, and notes that the input does not even need to be visible to a human, as long as the model parses it. That is what makes the problem interesting: a white-on-white sentence in a web page or an instruction buried in image pixels can still reach the model.
Jailbreaking is a close cousin. OWASP treats it as a form of prompt injection aimed at getting the model to ignore its safety rules entirely.
Direct vs Indirect Prompt Injection: What Is the Difference?
Direct prompt injection comes from the person typing. Someone asks a customer-service bot to ignore its rules and reveal internal data. Sometimes it is even accidental, when an ordinary request happens to collide with the bot's instructions.
Indirect prompt injection is the one that matters most for AI agents. Here the attacker never talks to the AI at all. They plant instructions in content the AI will later read: a product review, a shared document, a calendar invite, a code comment, an email signature. When a user asks their assistant to summarize that content, the hidden text rides along and tries to take control.
The more an AI system can do, the higher the stakes. A chatbot that only answers questions can, at worst, give a bad answer. An agent with access to your inbox, files and payment tools can be steered into taking real actions, which is why modern defenses focus so heavily on limiting and confirming what agents do.
Real-World Prompt Injection Scenarios
OWASP's LLM01 entry walks through nine example scenarios. A few illustrate the range well:
- Summary-based data leaks: a web page tells the AI summarizing it to embed the user's conversation in an image link, sending data to an outside server when the image loads.
- Poisoned knowledge bases: documents inserted into a retrieval-augmented generation (RAG) store quietly skew the answers the assistant gives.
- Split payloads: instructions broken across different sections of a résumé combine to push an AI screener toward a favorable rating.
- Multimodal tricks: instructions hidden inside an image are read by a vision-capable model.
- Encoded instructions: text in another language, Base64 or emoji is used to slip past simple keyword filters.
None of these require breaking encryption or exploiting memory bugs. They exploit the model's helpfulness, which is why the fixes are as much about system design as about the model itself.
How Do You Prevent Prompt Injection Attacks?
There is no single switch that turns prompt injection off. Every serious source agrees on defense in depth: several independent layers, so that if one misses, the next one catches it. OWASP lists seven strategies.
1. Constrain the Model's Role
Write a clear system prompt that defines what the assistant is for, what it must never do, and that it should ignore instructions that try to change those rules. On its own this is not enough, but it sets the baseline.
2. Define and Validate Output Formats
If your app expects a JSON object with three fields, check that it gets exactly that. Strict output validation stops many injected instructions from turning into unexpected behavior downstream.
3. Filter Inputs and Outputs
Scan incoming content for known injection patterns and scan outgoing responses for sensitive data. OWASP also points to the RAG Triad, which checks whether retrieved context is relevant, whether the answer is grounded in it, and whether it actually answers the question.
4. Enforce Least Privilege
This is the single most effective habit. Give the model only the tools and data it truly needs, and enforce those limits in ordinary application code rather than asking the model to police itself. An injected instruction cannot misuse a permission the agent never had.
5. Require Human Approval for High-Risk Actions
Sending money, deleting records, emailing people outside the organization and merging code should all pause for a person to confirm. This turns a potentially serious incident into a harmless pop-up the user can decline.
6. Separate and Label Untrusted Content
Clearly mark external content so the model can tell it apart from real instructions. Microsoft's Spotlighting technique, described below, is a practical example.
7. Test Adversarially, and Keep Testing
Run regular attack simulations against your own system. Treat the model as an untrusted user and see what it can be talked into. New tricks appear constantly, so this is ongoing work, not a one-time audit.
How Microsoft and Google Defend Their AI Assistants
The two largest productivity-suite vendors have both published how they handle indirect prompt injection, and their approaches line up closely with OWASP's list.
Microsoft's Security Response Center describes a mix of probabilistic and deterministic controls:
- Spotlighting marks untrusted text using randomized delimiters, interleaved marker tokens, or an encoding such as Base64, while the system prompt tells the model never to follow instructions inside that marked text.
- Prompt Shields, a classifier in Azure AI Content Safety, detects injection attempts and can surface alerts in Microsoft Defender.
- TaskTracker, a research technique, looks at the model's internal activations to spot when its task has drifted.
- Data governance uses sensitivity labels and access controls so an assistant cannot reach data the user should not see anyway.
- Human-in-the-loop design means, for example, that Copilot in Outlook drafts a message but the user must send it.
- Deterministic blocking shuts down known exfiltration channels, such as the markdown image technique, as a whole class rather than one case at a time.
Google describes five layers for Gemini in Workspace: prompt injection content classifiers that filter malicious instructions out of emails and files, security thought reinforcement that reminds the model to stay on the user's task, markdown sanitization that refuses to render external images and redacts suspicious links using Safe Browsing, a user confirmation framework for risky actions such as deleting calendar events, and end-user notifications when a threat has been blocked.
The common thread is reassuring: neither company relies on the model alone. They wrap it in ordinary, testable software controls.
What Can Everyday Users Do?
You do not need to be a security engineer to benefit from all this. A few practical steps:
- Review connected permissions. If an assistant has access to your email, drive and calendar, ask whether it really needs all three.
- Read confirmations instead of clicking through them. Those approval prompts exist precisely to catch injected actions.
- Be thoughtful about what you ask an agent to read. Summarizing an unknown web page is usually fine; letting an agent act on it automatically deserves more care.
- Keep tools updated. Vendors ship new classifiers and blocks regularly, and updates are how you receive them.
What Should Builders and Developers Do?
If you are shipping an AI feature or running agents in your pipelines, start with least privilege and human approval, then add layers. Our guide to securing AI coding agents in CI pipelines covers the permission side in depth, and the CodeQL 2.26 prompt-injection detection release shows how static analysis can now flag risky data flows into model prompts before they ship.
The Bottom Line
Prompt injection is real, but it is a well-understood engineering problem with a well-understood answer: limit what AI systems can do, label what they read, confirm what matters, and keep testing. The industry's biggest platforms already work this way, and the same habits scale down to a single developer's side project. For more practical defense coverage, browse our AI security section.
Sources: OWASP GenAI Security Project — LLM01:2025 Prompt Injection — accessed October 9, 2026; Microsoft Security Response Center — How Microsoft defends against indirect prompt injection attacks — July 29, 2025; Google — Mitigating prompt injection attacks with a layered defense strategy — June 13, 2025.
More Ai Security Stories

Anthropic OSS Scanner: Free AI Bug Scans for Open Source
Anthropic's OSS Scanner gives open-source maintainers free AI vulnerability scans with fixes; 85 of 97 tested severe findings met its disclosure bar.

GitHub AI Secret Detection: How Push Protection Gets Smarter
GitHub is adding a ModernBERT classifier to push protection that spots unstructured passwords in under 2 ms and could more than double blocked secrets.

Anthropic Cyber Verification Program: Which Tier Fits You?
Anthropic split its Cyber Verification Program into 3 tiers, Defense, Red Team and Specialized, giving vetted defenders broader Claude access for security.
