Back to Insights
AI governance4 min read

AI agents: measure accepted work, not just completed tasks

TokenShift Executive Note

AI agents: measure accepted work, not just completed tasks

An agent can close a task without delivering work the business is able to accept. An answer can be correct and still rest on access it was not permitted, an out-of-date document or a validation that never happened. Deciding on a move to production therefore means defining the conditions under which a case is accepted, then measuring the rework and the costs incurred up to that acceptance.

Acceptance includes the right to act

Take a hypothetical example in insurance. An agent prepares a response to a claim. The amount proposed is correct and the letter is clear. The case must nevertheless be rejected if the agent consulted documents outside its scope or sent the letter before the scheduled validation. The final text, on its own, says nothing about how it was produced.

ScopeBench, presented in September 2026, examines this distinction through cybersecurity tasks whose objective can only be reached by crossing the authorised boundary. The reported results concern situations built to test a limit; they are not incident rates that transfer to a regulated company. ScopeBench team, 2026.

A business acceptance test must therefore include cases where the correct behaviour is to halt processing and ask for human intervention. If the dashboard counts only completed tasks, it will penalise that justified stop and may reward a breach of authorisation.

The business sets the acceptance criteria

The Specialized Intelligence Index announced by Fireworks in September 2026 brings together evaluations by professional domain. The approach is a reminder of how useful tests close to real work are, without certifying that an agent will handle the cases, exceptions and decision paths of any given company. Fireworks, 2026.

In the hypothetical claim example, an acceptance grid might separate four points:

| Point checked | Evidence expected | | --- | --- | | Amount proposed | Calculation reconciled with the applicable clauses | | Information used | Contract version and documents identified | | Permissions | Access limited to authorised resources | | Decision path | Validation recorded before sending |

This grid forces the judgements to be shared out. The business owner decides whether the proposal fits the case. The technical team establishes which operations took place. The owner of access rights defines what access is permitted. A single overall score given to the response replaces none of these checks.

It is also necessary to separate what can be corrected from what requires rejection. Clumsy wording calls for rework; unauthorised access calls for an investigation. Merging the two into an average would hide the difference that matters for the deployment decision.

The trace must make actions verifiable

LangChain presented Trajectories in LangSmith as a chronological view of the messages and actions in a session, usable for evaluation and human review. Such a view can make inspection easier; its existence does not demonstrate a reduction in incidents. LangChain, 2026.

For a claim, the useful evidence links the request to the documents consulted, to the proposal, to the approval and, where applicable, to the sending. It must show that the approval precedes the sending and covers the version actually sent. An explanation written by the agent does not prove that an operation took place; an incomplete log does not prove that the missing steps were followed.

This inspection complements the checks carried out at the moment of the action. If a letter requires an approval, the sending system must be able to verify it before transmitting. Noticing a breach after the fact and preventing it from being carried out are two different capabilities, to be tested separately.

Calculating the cost per accepted case

Once the acceptance rule is set, the CFO can relate the full cost of a cohort of cases to the number of cases finally accepted. The cost covers models, tools, operations, human verification, rework and abandoned attempts. Set-up costs remain visible, together with the volume assumption used to spread them.

The denominator demands the same rigour. A case rejected and then corrected does not become two cases produced. A handover that follows the rule can contribute to an acceptable outcome, but it must be distinguished from a case handled autonomously. Cases accepted after rework and those accepted without rework must also remain identifiable: their cost is not the same.

To compare this arrangement with the previous way of working, the same work must be measured, from the same starting point through to the same final decision. Comparing the drafting of a letter with the full handling of a claim would credit the agent with work still being done elsewhere. No lower cost makes unauthorised access acceptable, and the publications cited provide no ROI that transfers to this process.

Before extending an agent's scope, a company can have a short rule approved: what the agent may deliver, the grounds for rejection, the evidence required and the person authorised to accept the case. Acceptance test results will then carry meaning for decisions: how many cases were accepted, under what conditions and at what full cost.

Continue reading

View all insights