Back to Insights
AI Governance7 min read

Token prices are falling, your AI bill is rising: manage cost per outcome

TokenShift Executive Note

Token prices are falling, your AI bill is rising: manage cost per outcome

On 13 August 2026, Google halved the price of Gemini 3.7 Flash and OpenAI unveiled Ultrafast, a mode that runs GPT-5.6 Sol up to fourteen times faster. Tokens have never been cheaper or quicker. And yet, in companies moving their AI agents into production, the invoices keep overshooting the budget. This paradox now has a name, documented by research: token inflation. It changes how an executive committee must budget for, contract for and govern its AI systems.

Token inflation: why the list price doesn't predict the bill

A research paper published on arXiv in July 2026 (Fu et al., "Not All Tokens Are Equal") formalises what teams see in practice: when an agent fails on its first attempt, the system retries, reloads the context, escalates to another model. Every cycle burns extra tokens. The authors define token inflation as the ratio between the full cost of the workflow and the cost of a single call. Their measurements: up to 4.25 times for a small model on multi-step question-answering tasks, and standard routing tools that underestimate the real cost by more than a factor of two on hard tasks.

Put plainly, a business case built on the price per million tokens is optimistic by design. That bias is anything but marginal. MIT (project NANDA, 2025) showed that 95% of generative AI pilots delivered no measurable impact on the bottom line, and as early as June 2025 Gartner predicted that more than 40% of agentic AI projects would be scrapped by the end of 2027, citing runaway costs and unclear business value as the leading reasons. Mispriced cost is not an engineering footnote; it is one of the top killers of these projects.

Price lists that no longer hold still: promotions, speed tiers, doublings

13 August 2026 also illustrates a second reality: vendor prices are no longer constants. Google launched Gemini 3.7 Flash at 0.75 dollars per million input tokens and 3.75 dollars for output, a promotional rate running until the end of 2026; on 1 January 2027, those prices double (VentureBeat, August 2026). OpenAI, for its part, now sells speed as a product: Ultrafast hits 750 tokens per second on Cerebras chips, in preview for selected customers, with no published price; its Fast mode already bills at twice the Standard rate (The Decoder, August 2026).

For a CFO, the implication is immediate: a use case that turns a profit at the 2026 promotional rate can slip into the red at the 2027 rate, and a poorly qualified latency requirement can push a workflow into a higher price tier with no business gain whatsoever. Contracts and business cases need price-review clauses and sensitivity scenarios, exactly as they do for energy or commodities.

The five-step method for steering on cost per outcome

  1. Define the unit of outcome. One unit per use case, in language the business understands: the claim handled end to end, the contract reviewed, the customer query resolved without human rework. That is what you budget — not the token.
  2. Instrument every workflow. Track per run: tokens consumed, number of attempts, model escalations, context reloaded. This telemetry doubles as the traceability documentation the AI Act requires for high-risk systems.
  3. Measure the inflation factor. It is the ratio between the full observed cost and the nominal cost of one successful call. Below 1.5, the workflow is healthy; above 2, you have a design problem, not a vendor problem.
  4. Put guardrails on escalation. The arXiv paper delivers a counter-intuitive result: handing a more powerful model the reasoning chain of a model that has just failed degrades its accuracy by as much as 34.8 points. Effective escalation starts from a clean context. According to the authors, this disciplined routing reaches 94.7% accuracy on a mathematics benchmark, against 91% for FrugalGPT, the reference approach, while using 31% fewer tokens.
  5. Govern it at board level. A monthly one-page review: cost per outcome, first-pass completion rate, sensitivity to vendor pricing. The business sponsor owns the number; IT owns the instrumentation.

Mini case: the 0.90 euro claim that actually cost 2.70

A representative case, figures rounded for illustration. The operations function of an insurer budgets its claims-handling agent at 0.90 euros per file, based on the pilot. Six weeks after go-live, the observed cost hits 2.70 euros: 30% of files trigger retries, every escalation ships the entire file to the higher-tier model, and the orchestrator re-reads the attachments at each step. Three fixes are enough: route simple files to a lightweight model, restart escalations from a clean context, cap the number of attempts with a handover to a human handler. The cost settles around 1.10 euros, and first-pass completion rate becomes the weekly steering indicator. None of these fixes required changing vendor.

Mistakes to avoid

  • The list-price business case. Budgeting on today's price per million tokens, with no inflation factor and no price-increase scenario. This is the mistake that manufactures MIT's 95%.
  • Dirty escalation. Passing failed reasoning chains up to the higher-tier model: you pay more for worse accuracy.
  • The speed premium by default. Paying for an Ultrafast or Fast tier on back-office workflows where latency carries no business value.
  • The pilot on promotional life support. Validating an ROI on a rate that doubles on 1 January 2027 without ever testing the full-price scenario.
  • No per-workflow meter. Without per-run telemetry, cost per outcome is unknowable and the budget conversation runs on guesswork.

The indicators your board should see every month

  • Cost per outcome, in euros, by use case, benchmarked against the cost of the equivalent human process.
  • Token inflation factor per workflow (actual cost divided by nominal cost).
  • First-pass completion rate, with no retry and no human rework.
  • Share of spend driven by retries and escalations.
  • Price sensitivity: the impact on the use case's margin if vendor prices double.

These five numbers fit on a single page. If they don't exist in your reporting today, that is the most useful takeaway from this article.

Where to start this week

Ask your teams for the cost per outcome of your three largest AI workflows in production. If the answer lands within 48 hours, together with an inflation factor, your setup is genuinely governed. If the answer is a price per million tokens, you are budgeting for a raw material when what you are buying is an outcome. Within 30 days: mandate a unit of outcome per use case, instrument the existing workflows, and write a price-review clause into every new vendor contract.

TokenShift works with the executive teams of regulated European companies to take AI from pilot to governed production.

Sources

  • Fu et al., "Not All Tokens Are Equal: Inflation-Aware Routing for Agentic LLM Systems", arXiv, July 2026. https://arxiv.org/abs/2608.13571
  • OpenAI, "Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed", August 2026. https://openai.com/index/previewing-ultrafast/
  • The Decoder, "GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras", August 2026. https://the-decoder.com/gpt-5-6-sol-goes-14x-faster-as-openai-launches-ultrafast-mode-powered-by-cerebras/
  • VentureBeat, "Google's Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut", August 2026. https://venturebeat.com/technology/googles-gemini-3-7-flash-targets-coding-and-agents-with-a-50-introductory-price-cut
  • Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027", June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
  • MIT NANDA, "The GenAI Divide: State of AI in Business 2025", Fortune coverage, August 2025. https://finance.yahoo.com/news/mit-report-95-generative-ai-105412686.html

Continue reading

View all insights