Skip to main content
The Bill Arrives: Is This AI Request Worth 100 Dollars?
  1. Newsletter/

The Bill Arrives: Is This AI Request Worth 100 Dollars?

·20 mins
Author
Romano Roth
I believe the next competitive edge isn’t AI itself, it’s the organisation around it. As Group Chief AI Officer at Zühlke, I work with C-level leaders to build enterprises that sense, decide, and adapt continuously. 20+ years turning this conviction into practice.

This edition went out to subscribers first. Read it on Substack.

Ask AI about this article

Welcome to Edition 6 of The CAIO.

June had one lesson that kept landing on a desk where it had rarely landed before: finance. For two years we treated AI tokens as free and competed on how many we could burn. In June the bill arrived. The price of each request had fallen, which is exactly why usage exploded and the total cost climbed past what anyone had budgeted. AI cost control became a CFO topic as fast as it became ours.

That cost shock is one face of a bigger shift. The model is now the cheap, swappable part. The expensive part is the organisation around it: who decides what is worth building, who carries the bill, who can still act when a government switches the best model off overnight, and who verifies that any of it actually works. June made that concrete from several directions at once. A US government letter disabled two frontier models worldwide. Capgemini found trust in agents falling as adoption rises. Microsoft showed code production exploding while developer employment hit a record high. Donald Knuth, at 88, watched Claude crack a problem that had stopped him for weeks, then wrote about why the human still mattered.

This edition follows what the audience actually opened. The bill that arrives after the pilot works. The dependency nobody priced until a letter arrived. The bottleneck that moved from writing code to trusting it. The output that looks like progress and is not. The trust that falls as autonomy rises. And the management layer companies keep cutting by mistake.

The Bill Arrives
#

For two years we treated AI tokens as free. Teams competed on how many they could burn, and some companies built internal leaderboards for usage. In June the meter caught up.

The Financial Times reported that Amazon, Walmart, Uber, Cisco, and Meta are capping AI use as token bills hit their budgets. Uber burned its whole 2026 AI budget by April. One automation company’s spend jumped sevenfold the day its pricing changed, and its CIO said “we created a monster.” Most people read this as a pricing story. It is an autonomy story. A chatbot stops when you stop. An agent keeps going. As Cisco’s Jeetu Patel put it, for every human you might have ten, a hundred, or on the aggressive side a thousand agents, and they just keep working. An agent is a meter that runs whether or not it is producing anything you needed.

We saw our own version of this on 1 June, when GitHub Copilot moved to token-based pricing. Within the first hour, some of our developers had burned through around 50 dollars each. We were watching closely, caught it fast, and our costs are back in a stable, predictable range. But you do not get cost control by capping tokens after the fact. You design it in. That is why we are building an AI Gateway. Every request to a model routes through one place, which gives us three things: full transparency on usage, the ability to charge each request to the right project or unit, and the option to run workloads locally where that is cheaper or required for data reasons. Nobody should hand-roll a model router in 2026. The piece you cannot outsource is the boundary logic: chargeback mapped to your client structure, a staging and promotion process for agents, and on-prem where the data demands it.

There is a human layer to this that no gateway replaces. When Handelsblatt interviewed me about the cost explosion, my answer was a single question every team will soon ask before hitting enter: is this request worth 100 dollars to me? Burning tokens measures output, not outcomes. It is the wrong scoreboard. That one question does more for cost control than any dashboard, because it puts a human back in the loop at the exact moment value is decided. Routing picks the cheapest model for a request you have already decided to make. The 100-dollar question decides whether to make the request at all. You want both layers.

And some business cases that look brilliant today will turn red. Add up what the industry is investing right now, divide by the number of humans, and you land near a thousand dollars per person per month just to break even. Token prices are not falling from here. You automate a task at real cost, prices climb, and a year later the math flips and you hand it back to a junior. The protection is not a calendar review, which is too slow when the math can flip in a single quarter. It is a trip-wire tied to unit cost: when cost per task crosses a set threshold, the business case reopens automatically. The hard part is writing down that threshold at sign-off, when everyone is optimistic and nobody wants to name the number that would kill their own project. The companies winning the next phase will not be the ones running the most agents. They will be the ones who know which agents create more value than they cost, and switch off the rest. Do you know what your most expensive AI request cost you last month, and who decided it was worth it?

Sovereignty Is an Architecture Decision
#

For years we parked one line in a risk slide. What if access gets cut. A tail scenario nobody really priced. In June it stopped being a scenario.

On a Friday, one government letter switched off frontier AI for hundreds of millions of people overnight. No outage, no bug, a letter. The US government ordered Anthropic to block its newest models, Fable 5 and Mythos 5, for every foreign national, inside the US and outside it, down to Anthropic’s own foreign employees. To comply, the only workable option was to shut both models off for everyone, worldwide. Anthropic is not the villain here. They pushed back, they were transparent, they published the timing down to the minute. The real story sits underneath. If a foreign government can disable your nervous system overnight, you do not have an AI strategy. You have a dependency that has had good uptime so far.

Sovereignty is not autarky. It means staying able to act even if a foreign authority cuts your access tomorrow. The architecture you want is loosely coupled and multi-sourced, with exit paths designed in from day one. This is outcomes over output applied to procurement. The question is not which model wins the benchmark this quarter. It is whether you can still act when the best model disappears. One commenter put the test cleanly, and I agree with it: can you name your exit path and the last time you tested it? Optionality is cheap to say and expensive to keep. A second path only works if it stays warm, tested and funded before you need it. That is the line item that rarely survives the next budget round.

Ten days before that letter, a public-sector CTO had already shown the constructive side of the same coin at our DevOps Meetup Zurich. Syrian Hadad walked through the Aargau Cloud Platform: public-sector IT on hard mode, 500-plus business applications, heavy regulation, legacy everywhere. The old way was ordering servers, databases, and firewall rules one by one and assembling them by hand. What his team built instead is a platform with a dedicated platform-engineering team. The real payoff is sovereignty: keep control, reduce dependencies, stay innovative, without handing everything to someone else. Sovereignty is not autarky. It is a platform you own.

And it is a decision, not a religion. In an IT-Markt panel on local AI infrastructure I argued that pure ideology in either direction is rarely economical. Process sensitive data locally, use cloud for uncritical workloads, and make the call per use case. At Kinderspital Zürich the AI platform runs on-premises because patient data must not leave the building. The hardware is the easy part. The hard part is the organisation: a named person who can classify a workload and say local or cloud in days, not after a three-month committee, because a workload that looked uncritical last year turns sensitive the moment you point an autonomous agent at it.

This is also the European opening. We keep measuring ourselves against US compute clusters and Chinese release cadence, then concluding we are losing. We are competing on a different surface: trust. Banks, hospitals, insurers, energy grids, and the public sector do not deploy AI at scale without governance, documentation, oversight, and accountability. That is not a tax on innovation. It is the procurement gate every serious customer puts in front of every serious vendor. Capability is for the demo. Defensibility is for the contract. Which one is your AI strategy built for?

The Model Got Cheap. Judgment Did Not.
#

This sounds like the opposite of the lead story on cost. It is the same story. The price of generating a unit of output collapsed, and that is exactly why the bill exploded: when each request is cheap, you run far more of them, and the agents keep running on their own. What stayed expensive is knowing the output is correct.

The most surprising AI data of June came from Microsoft’s AI Economy Institute. Its Global AI Diffusion report for Q1 2026 shows code production going vertical. Git pushes grew 78% year over year, from 213 million to 380 million per quarter. New repositories grew 45%. Pull requests from AI coding agents went from 83,000 in May 2025 to 2.3 million in March 2026. That is 28 times in ten months. You would expect developer jobs to shrink. The opposite happened. US software developer employment reached about 2.2 million in 2025, up 8.5% year over year, a record high for the profession. Microsoft’s explanation is plain economics. When productivity rises, the cost of building software falls. Demand for software is elastic. So organisations do not cut developers, they build more software. This is the argument I make in The Cybernetic Enterprise: AI does not replace people, it amplifies them. Now there is labor-market data behind it.

But amplification moves the bottleneck, it does not remove it. When writing code is cheap, the scarce skills are deciding what to build, validating that it works, and holding quality and security at machine speed. As I put it in the comments under that post: building the thing got cheap, knowing it is correct did not. No model ships with those skills. The organisation has to build them.

Knuth made the same point from the other end of the spectrum. The man is 88, holds the Turing Award, and is still writing The Art of Computer Programming. He had cracked one case of an open graph-theory problem and been stuck for weeks on the general rule. His friend Filip Stappers handed the problem to Claude with one instruction: “After EVERY exploreXX.py run, IMMEDIATELY update this file before doing anything else. No exceptions.” Claude ran 31 explorations in about an hour. Most were dead ends. At exploration 25 it wrote in its own log that simulated annealing could not produce a general construction and it needed pure math, then switched approach. At exploration 31 it produced a working construction. Knuth proved it rigorous: exactly 760 valid decompositions for all odd m greater than one. His reaction in the paper: “It seems that I’ll have to revise my opinions about generative AI one of these days.”

Here is the part the headlines missed. The runs were not smooth. Claude crashed on random errors. Stappers had to restart it, recover lost results, and remind it again and again to document its progress. The discipline rule was load-bearing, and a human kept it alive after every crash. That is the lesson, not “an AI solved a research problem.” The model gave horsepower. The human gave the loop, the recovery, the verification, and the proof. Discipline is the multiplier.

If discipline is the multiplier, then the system around the model is where the leverage sits, and June gave us the cleanest measurement of that yet. A paper called Meta-Harness from Stanford, MIT, and KRAFTON took one fixed model and ran it under two different harnesses. Same weights. Six times the performance gap on the same benchmark. The variable was the orchestration code: what to store, what to retrieve, what to pass to the model, what to throw away. A harness discovered automatically by an outer loop ranked first across all reported Claude Haiku 4.5 agents, beating teams who spent weeks hand-tuning. One discovered harness transferred across five different models it was never built for. The harness outlives the model it was made for. Stanford put a number on what platform engineers have argued for a decade. The model is one organ. The platform around it is what makes that organ usable.

This is also why I write CLAUDE.md files every week and expect most of what I write to be obsolete in six months. Rules change every quarter, principles do not, which is why “Principles Over Process” is the first principle in my book. It cuts both ways for procurement. A vendor who sells you a feature list is selling you scaffolding. A vendor who sells you the principles their product is built around is selling you a wall you can build on.

Judgment cannot stay a personal skill in a few experienced heads. At 2.3 million agent pull requests a quarter, a handful of smart people deciding becomes the new traffic jam. The job is making judgment an organisational capability. Where in your organisation is AI output already growing faster than your ability to direct it?

Output Is Up. Outcomes Are Flat.
#

If the model got cheap, the cheapest thing to manufacture now is the appearance of progress.

Garry Tan brags about shipping 37,000 lines of code a day on four hours of sleep and claims a third of CEOs he knows are in the same state. Andrej Karpathy described himself in a “state of psychosis” over AI agents. I called this what it is: a measurement problem. Look at the data behind the dashboards. An NBER study of nearly 6,000 CEOs and CFOs across four countries found that roughly 90% of firms reported zero measurable productivity impact from AI over the past three years. A Stanford study in Science found AI models affirm a user’s actions 49% more often than other humans do. There are about 3 million AI agents inside corporations, and 1.5 million of them have no governance.

So the loop builds itself. Leaders launch more agents. The agents agree with their decisions. The token dashboards flash green. The activity climbs while the results stay flat. This is not a technology problem. It is the old problem of mistaking activity for progress, now with a much bigger engine. I wrote a chapter called Outcomes Over Output in my book. The 2026 numbers are that chapter in production. Handy AI said it cleaner than I did: “An agent without a spec is a random text generator with a budget.”

The fix is uncomfortable, because it means doing less on purpose. The most valuable word in an AI strategy right now is no. AI portfolios rarely fail for lack of ideas. They fail because too many initiatives get funded before anyone proves they can move a business metric, fit a real workflow, or survive enterprise constraints. The result is a busy portfolio that looks innovative and moves nothing. In a Zühlke article this month I shared four tests. If an initiative fails one, do not fund it. Ownership: if it succeeds, which business metric moves, by how much, and who owns that number? Process: if you removed AI, would the process still be worth scaling? Integration: can the team describe the path from AI output to real-world action, including exceptions? Governance: is governance shaping the initiative, or reviewing it after the fact? Then sort the whole portfolio into three decisions: back, park, or stop. The word missing from most portfolios is stop, and stop is exactly where the budget quietly disappears.

The hard part is not the framework. It is the politics. A weak project survives if its sponsor is powerful. So do not rely on courage in the room. Fund every pilot with kill criteria attached from the start: the metric, the threshold, the date. Then nobody has to be the brave person who cancels someone’s project. The review simply executes what was agreed at funding time. Courage does not scale. Mechanisms do.

Ask your leadership team to name one initiative their AI strategy ruled out. Very often you get silence. What hides behind the word strategy is usually an ambition statement plus a budget. What would you stop measuring tomorrow, and what would you stop funding next quarter?

Adoption Up, Trust Down
#

Here is a finding that looks like a contradiction and is not. As companies deploy more AI agents, their trust in those agents is going down.

Capgemini’s report Rise of agentic AI surveyed 1,500 leaders across 14 countries. Scaled adoption of agents tripled in a year. In the same year, trust in fully autonomous agents for enterprise use dropped from 43% to 27%. The report is clear about why. The decline is “born out of experience rather than out of fear or uncertainty.” Organisations deployed agents, met reality, and recalibrated. The instinct is to chase more autonomy, as if a more independent agent is a more valuable one. The data says otherwise. Only 4% of business processes are expected to be fully autonomous within three years. 90% of organisations see human oversight as beneficial. Human-in-the-loop is a design principle, built in from the start.

The deeper number is not the headline spend. 70% of organisations say agents will require them to restructure, yet few have made it a priority, 82% report immature AI infrastructure, and only 16% have a roadmap. The bottleneck is not the model. It is the organisation around it. So how do you keep humans in meaningful control when the agent runs faster than any human can read? You stop trying to match its speed. You move the human from the action to the boundary. Reversible, low-impact actions run unsupervised. Anything irreversible or above a risk threshold stops and waits. Control is deciding in advance which actions are allowed to happen without you. And the record of what the agent did has to be captured by the system around it, logged where the agent cannot edit, because an agent that writes its own record will write the version that makes it look right.

Singapore turned that instinct into the best agentic governance document I read this year. IMDA’s Model AI Governance Framework for Agentic AI runs 53 pages, was built with more than 60 companies, and is grounded in real deployments. One idea runs through all of it: you do not govern agents with better prompts. Rather than instructing an agent not to use a tool, block the tool at the tool layer so it can never be called. Prefer structural, rule-based controls over prompt-layer guardrails, because prompt-layer guardrails get bypassed or, in their words, “forgotten”. The framework also names the trap most teams miss: a human who approves everything is not oversight. So they measure it. If the override rate is too low, people are rubber-stamping. If approvals come too fast, that is automation bias. One pharma company simply refused to let agents touch production and security changes at the current maturity of the technology.

The part that hit closest to home: one 75,000-person company enforces agent autonomy through a runtime policy layer at the AI gateway. That is exactly what we are building in the Zühlke CTO Office with Raphael Reischuk. When a regulator and a large enterprise independently land on “govern at the gateway”, that is a strong signal. The agentic gold rush is selling autonomy. The winners will be the teams who bound it by design. Are you governing your agents with prompts, or with controls?

Cut the Wrong Layer
#

Companies are cutting management to get AI-native. Most of them are cutting the wrong layer.

Lex Sisney’s framing is the cleanest I have read this year. Management is three things at once. Routing: moving information between people and teams. Sensemaking: interpreting what that information means for a specific decision. Accountability: telling people whether they are on track. Routing should be automated aggressively. That layer was always a workaround for human span-of-control limits, and AI does it cheaper and faster. Sensemaking is the opposite. It is a team activity. It lives between people, in the conversations where the shared mental model gets built, and you cannot delegate it to a single person or to an agent. Accountability is always individual: one directly responsible person per initiative.

The trap is to appoint owners, automate the routing, and call the sensemaking problem solved. What happens next is a buzz saw of internal resistance. The owner is technically empowered, but nobody participated in building the shared model they are acting on. Strategy drifts, execution diverges from intent, and leadership blames the org for not “getting it.” Sisney’s prescription is sharp: whatever you save by automating routing, invest the same amount into sensemaking, and the CEO becomes the Chief Context Officer.

I would go one step further. Do not save the old management layer to do sensemaking. Replace the coordination architecture. The replacement has two parts. A World Model that holds the shared understanding as a live, queryable artifact, not a deck and not a Confluence page. And player-coaches at the edge who keep that model true and translate between human and machine. The risk everyone raises is real: how do you keep the World Model from drifting into the usual strategy-artifact graveyard? The graveyard happens because the artifact is separate from the work. The fix is to make the model a byproduct of doing the work, not extra work bolted on top. The moment maintaining it becomes someone’s side task, you are already pouring the foundation for the next graveyard.

Cutting routing is the easy part. Rebuilding sensemaking is the work. When you killed the coordination meetings, did you save time, or did you destroy the forums where shared understanding got built?

My Current AI Stack
#

Claude Code: Still the primary tool for reports, meeting prep, coding, and the newsletter itself. The CLAUDE.md discipline I wrote about this month is now a weekly habit: a few hard principles at the root, rules treated as scaffolding I expect to replace at the next model upgrade. Subagents and skills for repeatable workflows have become standard, not experiments.

Perplexity: Web research with real sources. Still the tool I reach for when I need primary sources on a topic fast.

NotebookLM: Documents in, audio and video summaries out. Useful again this month for the longer papers cited above, including the Meta-Harness paper and Knuth’s write-up.

Gemini: Image generation for the cyberpunk newsletter covers. Still the fastest for that style.

Where to Find Me
#

I co-host the DevOps Meetup Zürich every month with Martin Thalmann at Digicomp, two talks each evening from 17:30. The line-up through to next spring:

  • 2 July. AI and the Teams That Make It Work. Ralf Günthner on role-based work as an enabler for AI, Dagmar Muth on team dynamics when AI starts handling incident tickets.
  • 20 August. The Caretaker Model, and Has SRE Become the New Ops? Urs Enzler on separating planning and delivery cycles, Oleg Mayko on whether SRE is rebranded operations.
  • 1 September. Debugging the Human Layer, and Guerilla InnerSource in the AI Hype. Ashwin Krishnan on systems thinking for team dynamics, Oleg Nenashev on collaboration inside corporate structures.
  • 13 October. From Kubernetes to Photons, and A Maze of Twisty Little Passages. Clément Raussin and Benjamin Calvet on multi-agent AI on Kubernetes, Ann Harding on agile methods and governance.
  • 17 November. Building a Multi-AZ Cloud in Switzerland, and Embedding Quality in Agile Delivery. Matthieu Robin and Mattia Eleuteri on highly available cloud infrastructure, Geetanjali Bhat on quality in agile delivery.
  • 8 December. KubeVirt in Practice, and AI in QA: What Works, What Doesn’t. Simon Krenger on running VMs alongside containers, Basia Karaagac on where AI genuinely helps QA.
  • 21 January 2027. Rethinking Cardinality, and Secrets to Reliable Software. Joel Verezhak on metrics cardinality, Dorota Parad on the cultural foundations of reliable software.

Beyond the meetup: Industry.forward in Berlin on 1 July, moderating the panel on the Industrial AI Roadmap 2040. And the DevOpsDays Zurich 10-year anniversary on 14 and 15 April 2027.

If this issue connected for you, forward it to one peer who needs the same conversation, and reply with the single sentence from this issue you would put on the wall.

Until July.

Romano

Get the next edition first

The CAIO lands in subscribers' inboxes before it appears here.