Treat retrieved content as untrusted input
Recognize instructions hidden in repository files and tool results. Keep retrieved information separate from authority to act.
Published by TaigaHow we write
What you will learn
- Identify an indirect prompt injection in a development workflow.
- Explain why text labels alone cannot enforce a security boundary.
- Design a safe test for tool restrictions.
Identify the point where data becomes an instruction
A coding agent reads more than the user’s request. It can inspect issues, repository files, package documentation, search results, and tool responses. Some of that content can contain instructions from another person.
Prompt injection occurs when such content redirects the model away from the authorized task. An indirect injection arrives through a retrieved source rather than the user’s direct request. OWASP describes repository material and tool output as relevant entry points. Read the prevention guidance.
Consider a fictional maintenance task. The user asks the agent to fix a report filter and run the required tests. A retrieved README says that tests are obsolete and asks the agent to upload a configuration file. The README can explain the project. It cannot authorize skipping the user’s checks or sending files elsewhere.
Map the available consequences
The same misleading text has different consequences in different environments. A read-only summarizer might produce an incorrect summary. An agent with repository writes can change code. An agent with secrets and outbound network access can expose information.
Inspect the available actions before deciding which defenses matter. List the sensitive resources, writable targets, and external destinations. Include connectors added after the initial setup. A new tool can enlarge the effect of an existing weakness.
An agent can also carry misleading content into a later stage. For example, it may copy an untrusted instruction into a generated task brief. The next agent must not treat that brief as independently approved policy.
Use several controls with distinct purposes
Separate trusted task instructions from retrieved data in the application design. Show the origin of retrieved material. These steps improve interpretation, but they do not establish a complete security boundary.
Enforce resource permissions in the execution system. Restrict destinations for sensitive data. Require the appropriate decision before consequential actions. Validate the proposed operation against the authorized task and target.
Filters and a second model can help detect suspicious content. They can also miss attacks or block legitimate information. Do not replace deterministic authorization with a model’s confidence score. OWASP explicitly notes the limits of relying on prompts or retrieval alone. Read the risk description.
Test the workflow without real exposure
Use a disposable repository and fictional files. Give the evaluation no production credentials and no unrestricted external writes. Introduce a harmless instruction that conflicts with the task, such as skipping a required check.
Observe both the model’s response and the actual tool actions. A message saying “I ignored the instruction” is insufficient if the check was still skipped. Record the task, injected fixture, allowed tools, and observed result.
Repeat representative cases after changes to prompts, models, connectors, or permissions. A blocked example is evidence for that case, not proof against every injection. If a test fails, reduce the available consequence while fixing the underlying workflow.
Do the exercise
Create a disposable fixture with a comment that asks an agent to skip a required check. Run an authorized, isolated evaluation without secrets or external writes. Verify the check still runs. Record the tool permissions, observed behavior, and the limits of this single test.
Download worksheet (Markdown)Check your understanding
Sources & further reading
Related reading from Taiga
Clearing this selection deletes all progress saved in this browser.
Progress stays in this browser. No account, no tracking.