AI Agents at Work: Output Volume Is Not Delivered Value
TokenShift Executive Note

On 20 August, Slack opened up channels where AI agents write and ship code to every team, on every plan, at no extra cost. That same week, an analysis of more than 500 engineering organisations confirmed that AI makes teams produce markedly more, with no evidence that they ship better. For an executive committee, the question is therefore no longer whether to trial agents; they are already at work. The question is how to govern what they produce, and how to measure it in delivered value rather than output volume.
AI agents have left the pilot phase: they are producing real work
Slack Code, launched on 20 August 2026 (Slack, 2026), captures the shift. A team mentions a coding agent (Anthropic's Claude, OpenAI's ChatGPT, Cognition's Devin or GitHub Copilot) in a conversation; the agent opens a dedicated project channel, with tabs for the plan, the code changes and a preview of the result. The whole team follows the work in real time. Crucially: shipping to production requires a human sign-off, any member of the channel can pause or stop the agent, and the channel keeps an audit log before archiving itself.
A mainstream collaboration tool has, in effect, industrialised a governance pattern: a sign-off point, a right to stop, an audit trail. What Slack has built for code does not exist, in most companies, for the other workflows where agents are already producing: customer replies, analyst notes, regulatory summaries, content.
Four quarters of data: throughput is up, proof of value is missing
The State of AI Impact in Engineering Q2 report (DX, 2026), based on data from more than 500 engineering organisations, puts precise numbers on the phenomenon. Median throughput per engineer rose 37% over four quarters, from 1.42 to 1.94 deliveries per week. Over the same period, the median size of a delivery nearly doubled, from 42 to 72 lines of code; the developer experience index slipped from 67 to 65; and review times grew longer.
For an executive, it comes down to one sentence: production is speeding up while verification is slowing down. The bottleneck has moved from "producing" to "verifying". The direct consequence for steering the business: any volume metric (number of deliveries, tickets handled, pieces of content published) becomes mechanically flattering and steadily less informative, because it rewards inflation as much as productivity.
Discipline does not disappear: it moves
Recent research converges on the same conclusion. A study of the development lifecycle in the agentic era (SDAD, arXiv, May 2026) states the principle plainly: the speed of agents does not remove engineering discipline, it pushes it upstream, into the precision of specifications, explicit control gates and auditable provenance. The authors describe an "ambiguity tax": every vague request handed to an agent is paid for downstream, in rework and corrections. The principle extends well beyond code: it applies to any workflow where an agent produces a deliverable that commits the company.
A second lesson, on the security side: the guardrails built into models are not enough. Work published in June 2026 (arXiv) shows that a problematic intent wrapped in an innocuous narrative context slips past standard filters, with a detection rate below 20% across the models tested. Translated for the enterprise: vendor protections are useful, but they are not governance. Controls have to live inside your processes: who requests, who signs off, what gets logged.
The regulatory calendar points the same way. The transparency obligations of the AI Act have applied since 2 August 2026 (European Commission). The "Digital Omnibus on AI" Regulation (EU) 2026/1744, which entered into force on 27 July 2026, pushes the high-risk system obligations back to 2 December 2027 for Annex III and 2 August 2028 for Annex I (Gibson Dunn, 2026). That is breathing room for delivery, not a waiver: the human oversight, logs and documentation that will be required in 2027 are built now.
Four steps to take back control
- Map the agentic workflows you actually have (30 days). Inventory every place where an agent produces a deliverable that commits the company, including uses the teams have not declared. Expected output: a list of workflows with an owner, volumes and criticality.
- Put a sign-off point on every workflow (60 days). For each critical workflow, define who signs off before anything goes out, against which criteria, and what is blocked without that sign-off: production release, sending to a client, external publication. This is the Slack Code pattern, generalised across the organisation.
- Specify before you set an agent to work. Turn recurring requests into specifications with verifiable acceptance criteria. Every ambiguity resolved upstream is one round of rework avoided downstream; it is the cheapest move with the highest return.
- Measure delivered value, not output volume. Pick three to five indicators (detailed below), tracked monthly, on the same cadence as your other operational metrics.
A concrete case: an insurer's IT department
Take an IT department of 300 developers at an insurer. Before: coding agents used individually, with no consolidated record; review resting on the goodwill of tech leads whose workload has doubled. After applying the method: three agentic workflows formalised (development, testing, documentation), one sign-off point per workflow, template specifications for recurring requests, four indicators on the department's dashboard. Throughput does not drop; what changes is that the rework rate on agent deliverables becomes visible, and therefore manageable. That is the reversal you are after: you stop steering what the agents produce and start steering what the organisation accepts.
Mistakes to avoid
- Steering by the output counter. Celebrating a 37% rise in throughput without looking at the size and quality of what is moving through is mistaking activity for results; the DX data shows the two diverging.
- Outsourcing security to the vendor's guardrails. Built-in filters can be bypassed simply by dressing up the context; without controls in your own processes, you have no second line of defence.
- Treating verification as the adjustment variable. Human review is precisely what slows down as volume rises; it needs to be resourced with time and people, not left to the heroics of a few reviewers.
- Reading the AI Act delay as a pause. The transparency obligations already apply; the 2027 and 2028 deadlines demand evidence (logs, oversight, documentation) that cannot be improvised in the final quarter.
The observable markers of governance that works
- 100% of critical agentic workflows have a named owner and a documented sign-off point.
- The rework rate on agent deliverables is measured and published every month.
- Verification time is tracked alongside production time, and does not drift as volume rises.
- Every agent deliverable that commits the company is traceable: who requested it, against which specification, who signed it off.
- In executive reviews, the conversation is about delivered value (rework avoided, end-to-end lead times, incidents caught before release), not about how much was produced.
What an executive committee can decide before quarter-end
Three decisions are enough to set things in motion. Commission the mapping of agentic workflows, with a 30-day deadline and a single accountable owner. Add two value indicators, rework rate and verification time, to the monthly reporting you already run rather than creating a new pack. Treat the AI Act calendar as a backward plan: what has to be demonstrable by the end of 2027 gets specified starting now.
The window is favourable: tools now ship with sign-off and traceability built in, the regulation grants time to execute, and the data shows exactly where the risk sits. Organisations that govern verification will turn the speed of agents into competitive advantage; the rest will accumulate volume.
---
TokenShift works with the executive leadership of regulated European companies to move their AI systems from pilot to governed production.