AI Agent Governance: Six Controls Before You Open the Catalog
TokenShift Executive Note

The risk in a plugin catalog isn't how many tools it offers. It's how easily the right to read, write or send can pass for a routine software option. AI agent distribution is scaling up; governance now has to bite at the level of each real workflow.
The catalog industrializes access, not accountability
Since 9 July 2026, OpenAI has replaced its app directory with a plugin directory. A plugin can bundle instructions, applications and application templates; it can therefore wire an agent into a CRM, an inbox, a document workspace or a code repository. The platform states that existing permissions still apply and that administrators can distinguish between read, act and confirm-before-execution (OpenAI, July 2026).
On 2 June, OpenAI had announced six business plugins bundling 62 applications and 110 skills (OpenAI, 2026). Discovery and installation become a distribution layer.
In its 2025 survey, McKinsey finds that 62% of respondents are experimenting with or deploying agents, but that no more than 10% are deploying them at scale within any given function (McKinsey, 2025).
The catalog lowers the effort of getting access. It decides nothing about purpose, risk level or expected value.
The Agent Passport: six controls before installation
We propose a simple unit of governance: the Agent Passport. One page, six fields, one business owner. Without a complete passport, a plugin can be trialled in a sandbox — it does not get production rights.
1. Purpose and decision boundary
Spell out the trigger, the steps executed, the expected output and the forbidden actions. "Help customer service" doesn't cut it. "Draft a reply from three approved sources, without sending it or altering the case file" is a boundary you can check.
2. Owner and escalation
Name the person who signs off on the business outcome, approves any change in rights, and can shut the agent down. The owner is not "the AI committee". Provide a deputy and an escalation deadline.
3. Data scope
List every source, its sensitivity class, the basis for access, the retention period and any transfers. Blanket access to the Drive or the CRM is not a scope.
4. Action rights
Separate read, create, edit, send, delete and commit spend. Start read-only; then open one specific right at a time. The technical identity used, the confirmations required and the revocation mechanism all belong in the passport.
5. Evidence and human oversight
For every run, keep the relevant input, the sources consulted, the tools called, the output, the human sign-off and the final action. Since 2 August 2026, the transparency obligations of Article 50 of the AI Act apply to systems in scope (European Commission, 2026).
A plugin is not automatically "high risk"; the intended use drives the classification. Where Article 26 applies, it requires among other things oversight by people who are competent and duly authorised, monitoring of the system, and — for logs under the deployer's control — retention of at least six months (EUR-Lex, 2024).
6. Outcome and full cost
Set a baseline before the trial: workflow cycle time, cost per completed case, rework rate, backlog and incidents. Add the cost of the model, the connectors, the human sign-offs and the oversight itself. A faster agent that shifts the work into correction has not improved the workflow.
Going live in four steps
- Map it. Walk through the current workflow, its exceptions and its systems. Name the owner and measure the baseline.
- Watch without acting. Run the agent in shadow mode on a closed sample. It proposes; the team still executes. The gaps become your test set.
- Probe the limits. Throw in missing data, an ambiguous instruction, a duplicate, an out-of-scope access attempt and a connector outage. Each case must produce a defined response: refusal, request for sign-off, or escalation.
- Open in stages. Move from read to reversible write, then to a confirmed external action. Any new source, action or major version reopens the passport review.
Mini-case: the claims agent that doesn't decide the payout
Take a motor insurer. The agent reads the dedicated inbox, matches documents against the policy, flags a missing item and drafts an acknowledgement of receipt. It doesn't rule on coverage, doesn't adjust the reserve and doesn't trigger any payment.
The first stage covers 200 closed cases, read-only. The second allows a note in a test environment. The third permits sending the acknowledgement after human sign-off.
Moving from one stage to the next depends on observed results: rate of corrected fields, unsourced references, human rejections, time per case, cost per completed case and share of traceable actions.
Mistakes to avoid
- The catalog seal of approval. Mistaking a listing in a directory for internal clearance of the vendor, the contract or the use case.
- Maximalist OAuth. Granting every right requested to save a week, then never scaling them back.
- The pilot with no baseline. Counting tasks produced without comparing cycle time, quality, backlog and full cost.
- Ownership by committee. Briefing five functions without giving one person the authority to stop the agent.
- Compliance by PDF. Describing a control that exists neither in the permissions, nor in the logs, nor in the escalation procedure.
The markers the board needs to see
The monthly dashboard should fit on one page: value per completed outcome; quality, with correction rate and error severity; control, with unauthorised actions, log coverage and time to revoke; change, with versions of the plugin, the model, the connectors and the workflow.
An agent goes into production when its rights, its evidence and its outcome have an owner — not when its plugin installs.
This week, ask for an inventory of the agents already connected to internal systems. Then pick one workflow and require its Agent Passport before granting any new write access. The metric that matters isn't how many plugins are installed; it's how many useful outcomes were produced through an action that is attributable, reversible and measured.
Sources
- OpenAI, Plugins in ChatGPT and Codex, July 2026
- OpenAI, Codex for every role, tool, and workflow, 2 June 2026
- McKinsey, The state of AI in 2025, 5 November 2025
- European Commission, guidelines on Article 50, 20 July 2026
- European Union, Regulation 2024/1689, Article 26
Scope of this briefing: the most recent material reviewed dates from 9 August 2026. It documents distribution and permissions, not incident rates or the ROI of plugins in regulated production environments; those results have to be measured company by company.