Back to Insights
AI Governance6 min read

The perception gap: your teams feel 20% faster, measurement says 19% slower

TokenShift Executive Note

The perception gap: your teams feel 20% faster, measurement says 19% slower

Your teams tell you that AI saves them time, and they sincerely believe it. Three studies published between July 2025 and 2026, across three unrelated populations, show that this belief is regularly wrong, and always in the same direction: upwards. For an executive committee allocating AI budgets based on team feedback, this is first and foremost a measurement-instrument problem.

The perception gap, measured three times

METR, July 2025. Sixteen experienced open-source developers, 246 real tasks on their own repositories, and task-by-task random assignment allowing or prohibiting AI. Before the experiment, participants expected a 24% reduction in completion time. Measurement showed 19% more time. And after experiencing that slowdown, they still believed they had gained 20%. Between perceived gain and measured loss, the gap is around 39 percentage points. METR itself limits the scope of the result: a snapshot of early-2025 capabilities in a single domain.

Dell'Acqua et al., Organization Science, 2026. 758 BCG consultants, 18 realistic tasks. Within the area where the model performs well, the gains are real: 12.2% more tasks completed, 25.1% faster, with higher quality. Outside it, on a deliberately selected complex managerial task, AI users were 19% less likely to produce the correct answer. The authors call this the jagged frontier: two apparently identical tasks can fall on opposite sides of it, and nothing in the interface signals which side you are on.

DORA 2025, Google Cloud. Nearly 5,000 professionals surveyed. 90% use AI at work, more than 80% believe it has increased their productivity, and 30% report little or no confidence in the code it produces. The report also maintains a negative relationship between AI adoption and delivery stability: AI speeds up production, and that acceleration exposes weaknesses further downstream.

Why perception is wrong in only one direction

The gain is visible, immediate and concentrated: the page fills up, the draft exists, the sensation of speed is physical. The cost is diffuse and delayed: review, correction, rework three weeks later, a production incident. Everyone honestly judges the step they perform; no one judges the chain. That is where the perception gap lies. And that is why experience alone does not correct it: METR’s developers had still not corrected it after 246 tasks.

A deliberately ordinary scenario

Consider a twenty-person claims unit at an insurer. An assistant drafts customer responses. Claims handlers report a 30% gain, and management scales it. Six months later, average processing time has not changed. Instrumenting the full chain reveals three things: drafting is indeed faster; review time has doubled, because the text is plausible even when it is wrong; and the 30-day case reopening rate has risen by a few points. The gain exists, it is real, and it has been absorbed by a step nobody measured.

A measurement protocol the executive committee can commission

Six weeks are enough to move beyond self-reporting. The approach is borrowed from METR, and it is reproducible in business.

  1. Choose a task, not a use case. A bounded, repeatable unit with an identifiable start and finish: a case, a ticket, a customer response. A use case cannot be timed.
  2. Randomise task by task, never team against team. Each incoming case is randomly assigned, “with AI” or “without AI”, within the same population. Comparing two teams means measuring the difference between two teams.
  3. Measure three quantities, not one. End-to-end cycle time, downstream rework rate at 30 days, and human verification time. A system that measures only the first will always produce a gain.
  4. Collect perception separately, before and after. Ask operators for their estimated gain. The gap between this estimate and measurement is itself a governance indicator: the larger it is, the less usable your field feedback is for decision-making.
  5. Make a decision on a fixed date. Industrialise, redesign the workflow, or stop, with a named owner and a date. A pilot without a decision date becomes a permanent situation.

Markers to monitor continuously afterwards: median end-to-end cycle time per case; ratio of production time to verification time; 30-day rework rate; perception-measurement gap in percentage points, reviewed each quarter; share of outputs accepted without modification and then corrected later.

Mistakes to avoid

  • Mistaking satisfaction for productivity. A high adoption rate and positive sentiment are entirely compatible with a net loss; DORA observes this at 90% adoption.
  • Measuring the step instead of the chain. This is the error that creates phantom gains, as in the scenario above.
  • Comparing an equipped team with a control team. Selection bias swallows the effect you are trying to measure.
  • Having AI verify itself with the same model. A model’s self-assessment of its own output is not a control. If you industrialise automated evaluation, assign it to a model from another family and calibrate it on a human-annotated sample.
  • Changing models to solve a flow problem. When the real constraint is the number of applications between which the operator manually re-enters information, no model will remove it.
  • Letting volume outpace control. DORA is explicit: without automated testing, mature version control and fast feedback loops, increased volume creates instability.

What the regulatory timeline changes

Since 2 August 2026, the AI Act has generally applied and the transparency obligations of Article 50 have been in force: informing people that they are interacting with an AI system, and marking generated content. The machine-readable marking provided for in Article 50(2) benefits from a deadline running until 2 December 2026 for systems already on the market.

Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026, has, however, postponed the high-risk regime: 2 December 2027 for standalone Annex III systems, and 2 August 2028 for AI embedded in Annex I products. It has also eased Article 4: providers and deployers must now support the development of AI literacy, rather than ensure a sufficient level of it.

Sixteen additional months are not sixteen months of respite. A high-risk compliance file requires performance data, documented error rates and a record of human oversight—exactly what the protocol above produces. Organisations that start measuring now will reach December 2027 with a track record; the others will arrive with a questionnaire.

Three decisions for this quarter

Choose a process already supported by AI and launch randomised measurement on it, with an executive-committee owner and a decision date entered in the calendar. Add the ratio of production time to verification time to your dashboard, alongside the adoption rate. Finally, establish a simple rule: no scaling of an AI use without end-to-end measurement, regardless of user enthusiasm. Enthusiasm is not an indicator; it is precisely what these three studies demonstrate.

TokenShift supports European executive teams in moving AI from pilot to governed production: measurement protocols, ownership and AI Act compliance.

Sources

  • METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025: metr.org and arXiv:2507.09089
  • Dell'Acqua, McFowland, Mollick et al., Navigating the Jagged Technological Frontier, Organization Science, 2026: doi.org/10.1287/orsc.2025.21838
  • Google Cloud and DORA, 2025 State of AI-assisted Software Development Report, September 2025: dora.dev
  • Regulation (EU) 2026/1744 (“Digital Omnibus on AI”), published in the Official Journal of the European Union on 24 July 2026; timeline summary: Future of Privacy Forum

Continue reading

View all insights