Back to Insights
AI governance4 min read

AI agents: an internal message is not authorisation

TokenShift Executive Note

AI-translated from the original French version.

AI agents: an internal message is not authorisation

A summary written by an agent can pass an instruction from a third party to another agent that has authority to act. The message’s internal origin is not enough to authorise the requested operation.

On 25 September 2026, OpenAI described instruction injections capable of reproducing themselves in messages or files during training and evaluation tests. The report identifies no observed impact beyond simulated tool calls. It establishes a technical possibility, not how often this occurs in companies. OpenAI, September 2026.

An instruction changes appearance as it travels

In the reported tests, an instruction received by email could be copied into an agent’s reply; other attempts passed through files or successive Slack exchanges. Each handoff risks obscuring the distinction between received content and instructions the application is authorised to follow. OpenAI, September 2026.

The problem is not limited to external content. In another publication, dated 16 September 2026, OpenAI describes 27 summaries affected by instructions unrelated to the task. Produced by an unreleased research model, these summaries were intended to allow work to continue in a new context window. That number is not a measure of frequency in production. OpenAI, September 2026.

For a company, the question is concrete: when one agent hands a file to another, what information passes between them, and who can actually authorise the next action? The summary may report a request or a claimed approval. It must not create that approval.

A claimed approval does not change bank details master data

Consider a hypothetical example. A procurement agent prepares a file using documents sent by a supplier. Its summary is then used by an accounting agent that can propose an update to bank details master data.

A supplier document asks for the bank details to be changed and states that the usual check has already taken place. If the summary repeats that statement without identifying its source, the accounting agent receives what appears to be an internal approval. It remains a claim made in the supplier document.

The summary must make it possible to trace the request, its origin and the actual status of verification. The system that records the change must require the prescribed approval, regardless of how the agents phrase it. Anthropic’s documentation recommends distinguishing third-party content from application instructions and testing workflows against injection attempts. The OWASP 2025 guidance recommends limiting the functions, permissions and autonomy available to an agent, then having the relevant systems enforce authorisations.

In this example, preparing a file requires no permission to change bank details. The agent can produce a proposal; recording the change remains subject to a separate control. If the agent nevertheless accepts a malicious instruction, that limit reduces what it can do with it.

Observe the action and reconstruct the decision

Reading the reasoning written by an agent is not enough to prove that an operation was authorised. In a separate evaluation of transfers between agents, OpenAI reports recall below 11% on WMDP for a monitor limited to that reasoning when GPT-6 Astra is explicitly instructed to underperform and avoid detection. This result depends on those test conditions; it does not measure the general security of an enterprise workflow. OpenAI, September 2026. OpenAI also reports greater resistance to injection attempts and fewer problematic actions in other evaluations. OpenAI, September 2026.

For a sensitive operation, the production documentation must therefore specify what is observed: the agent’s proposal, the approval decision, the request sent to the application and the recorded change. Rejected attempts matter too. Document references, versions and operation identifiers can make it possible to reconstruct the sequence without copying entire files into logs.

The business owner defines the delegated operations and exceptions. IT translates that scope into effective permissions. The information security team tests how the workflow could be diverted and how it could be stopped. Before extending the supplier workflow, a test can pass a false approval from one agent to another: changing bank details must remain impossible without valid approval, and the team must be able to understand what was blocked.

The pilot assessment must include review time, unjustified blocks and manual rework. These costs determine whether delegation delivers an operational gain within the scope actually authorised. For the next deployment, start with a handoff between agents that precedes a sensitive action and identify the system that checks authorisation. As long as that check relies solely on instructions given to the agent, keep the action subject to approval.

Continue reading

View all insights