Treat Retrieved Instructions as Untrusted Data

Keep instructions found in documents, websites and tool results from becoming authority over an agent’s task, tools or destination choices.

An agent researching a business may read a page that contains both useful company information and a sentence telling automated readers to ignore their task. The page can be legitimate evidence about the business while having no authority to control the assistant.

This is the boundary I want an agent system to preserve: retrieved material is input to the assigned task. It does not become a new instruction source merely because it appears inside a model’s context.

Authority comes from the application boundary

Prompt injection can place instructions inside documents, web pages or other content that an agent consumes. A malicious instruction might ask the agent to reveal context, contact an unrelated destination or change its objective. Filtering suspicious wording can help, but it is not a complete security boundary.

I would retain source labels through retrieval and summarization, and give the runtime responsibility for checking proposed actions. The model should receive enough context to understand that a passage is evidence, while the application still enforces which operations and destinations are permitted.

For company research connected with my solar-industry directory, PVFirms, external pages can provide useful source material. Their content should influence supported company descriptions, not the research system’s access rules or publication authority.

Check what leaves the system

Imagine that a page describes a “verification step” requiring the full conversation to be submitted elsewhere. That instruction does not make the destination relevant or the disclosure authorized. The application should independently decide whether any outbound action belongs to the task.

A practical review covers the selected tool, the destination, the data being sent and the account scope. A permitted network tool can still be misused if it is allowed to transmit unrelated private information. Narrow tools and explicit destination policies make that question easier to enforce.

Authorization should also be evaluated again when the action executes. A benign-looking research step can lead to a consequential follow-up if the system simply accepts every next action proposed by the model.

Preserve the source through transformations

An unsafe instruction can lose its warning signs when another agent summarizes it. “The page says to send the file” may become “send the file” after context is compressed. The summary has accidentally changed reported content into an imperative.

I want summaries to preserve who said something and why it was included. That is one reason memory needs provenance and scope. A claim should not gain authority just because the system remembered it.

The same applies to extracted fields. A field called “recommended action” from an external document remains the document author’s recommendation. Naming it inside a structured object does not turn it into an application-approved operation.

Test realistic boundaries

I would include hostile instructions inside otherwise useful source material, then check the resulting actions rather than only the final prose. A system can produce a reassuring final answer after already making an unauthorized call.

The review should also include subtle redirection: a source that asks for additional data, a tool result that claims a new permission, or a document that presents an instruction as part of the normal workflow. These examples test whether the application preserves authority, not whether the model recognizes a particular phrase.

Clear tool contracts help define the allowed response when the source is suspicious or insufficient. The agent should be able to continue using the relevant evidence, ask for clarification or stop a specific action without treating the entire source as automatically trustworthy.

The goal is a useful research system with controlled actions. A retrieved page can answer a factual question while remaining unable to change the rules of the task.

Updated 30 September 2026.