Skip to main content
The Bottleneck Was Never the Model
  1. Blogs/

The Bottleneck Was Never the Model

Author
Romano Roth
I believe the next competitive edge isn’t AI itself, it’s the organisation around it. As Group Chief AI Officer at Zühlke, I work with C-level leaders to build enterprises that sense, decide, and adapt continuously. 20+ years turning this conviction into practice.
Ask AI about this article

Christoph Gulden invited me onto his Thought Leaders Talk. Christoph reviews investment decisions for a living, I guide companies through technology waves. What came out of it is an argument about the 90 percent of AI pilots that never reach production, and what that means for the next investment decision. At one point Christoph contradicts me openly, and he is right to. The dialogue below is the one we worked out together. The original appeared in German on LinkedIn; this is my English version.

Thought Leaders Talk with Romano Roth: the bottleneck was never the model

Our Conversation
#

Christoph Gulden: Romano, an AI pilot is producing positive signals. Usage is up, people work faster, the calls for a broad rollout are getting louder. What is missing for you?

Romano Roth: The proof that it was a pilot at all.

I have seen this film twice before AI arrived. Once with the internet, once with the cloud. Classically, around 70 percent of proofs of concept never made it past the PoC stage. With AI the rate is above 90.

The technology changes, the trap stays the same.

Christoph: And the trap is?

Romano: That nobody can say which strategic goal the thing contributes to.

A proof of concept validates a business case, and that business case has to contribute to a strategic goal. If it does not, I can almost guarantee it will never go to production. “We want to have AI” is not a strategy.

Technology is the smallest worry here. There is a tool for nearly every use case, and the models get better every few months. The bottleneck was never the model.

Christoph: And yet almost every discussion starts with the tool.

Romano: Because the tool is visible and discipline is not.

When a team is unhappy with the quality of AI-generated code, someone goes looking for a better prompt. But the models are good enough. What is missing is everything around them: no spec, no context discipline, no automated check that catches a wrong answer before it lands in review.

Discipline beats intelligence. The constraint is process and governance.

Christoph: Back to the pilot that looks good. What is it measuring wrong?

Romano: It measures the thing that goes up in the first few weeks with every new tool.

Most AI pilots are built so that they cannot fail. What gets measured is usage and satisfaction. Both rise whether or not value is created at the end. That is not insight, that is a rollout with a measuring device attached.

Usage is not adoption. A tool in use is not yet a changed way of working, and a changed way of working is not yet impact.

Christoph: In my investment checks I always ask the same question at this point: which number would have looked different if the pilot had failed? If nobody can answer that, nothing was measured, it was only observed. What do you measure instead?

Romano: The end of the chain.

Local speed is not an outcome. There is a study of more than 100,000 GitHub developers that shows the curve: coding agents lift commits by 180 percent, at project level 50 percent remains, and for actual releases 30. The gain decays on the way to the customer.

A chain runs at the pace of its slowest step. AI accelerates the writing, while review, integration, test and release still run at human pace today.

Christoph: You say “today”. If those steps get faster too, your bottleneck dissolves.

Romano: It does not dissolve. That is exactly what we are experimenting with right now.

We build deterministic CI/CD pipelines with clear gates and clear instructions. Checking happens where the code is created, not three days later in the pull request. In my book I call this the gated commit, the bouncer of the codebase: static analysis, tests, coverage thresholds, security scans. Whoever fails does not get in.

What was good engineering hygiene for years becomes a load-bearing wall the moment a machine writes the code. So the gate has to get smarter too, with AI-assisted review and generated tests for the gaps. The machine helps against the machine.

We have already learned one lesson: a gate that blocks from day one gets switched off after the first false alarm. So we let new checks run alongside first and log what they would have blocked. Only once they catch a real defect correctly are they allowed to block.

Christoph: And if that works out, what stays human?

Romano: The decision about what the gate demands. And the case that no rule fits.

The pace of execution becomes machine driven, the pace of decisions does not. Outcomes over output. If a metric measures output, it measures the wrong thing.

Christoph: You had a client who wanted to roll out an AI coding assistant across the board. What did you do differently?

Romano: We stopped the rollout before it started. Handing a tool to everyone is a procurement decision, not a strategy.

Instead we asked the question that almost never gets asked: where is the bottleneck in the first place? Not in a room with twenty people, but one on one, and before anybody touches a tool.

In my experience the answer almost always points at requirements. Not at coding, not at testing, not at deployment. Typing was never the problem.

Had we only accelerated the writing, we would have increased the pressure on review and release without more reaching the customer. That is whole-system thinking. You optimise the flow of value, not a single station.

Christoph: So the trial did not confirm productivity, it refuted an assumption. Is that just as valuable?

Romano: More valuable. An experiment is good when it improves the decision.

And this is where it gets uncomfortable. A trial nobody is allowed to stop is not an experiment. It is a rollout with an interim report. Most organisations have already communicated the rollout internally before the pilot even started. From that moment on, the measurement can only confirm.

Christoph: You know my countermeasure: write down the threshold before the start. Which observation would overturn our preferred decision? It costs ten minutes.

Romano: Your ten minutes are right, and they are not enough.

They only work in an organisation that can tolerate an uncomfortable result. Goal, assumption, metric, threshold, consequence, every company writes that down in half a workshop.

The failure happens elsewhere: who is allowed to pull the plug after the result without it costing them their career? As long as the answer is “nobody”, all that preparation is decoration.

Christoph: Here I disagree with you, and specifically on the sequence. The authority to stop is created in exactly those ten minutes. Whoever writes down the threshold has to write down who pulls it, and that is the only moment when this question is cheap to answer, because nobody has invested yet. Six weeks later the same question costs somebody their face.

Romano: I accept that. You see the moment before, I usually see the state afterwards.

I make that role a condition before budget flows. No sponsor with a business KPI, no trial.

And I check the foundations. Many companies are currently trying to use AI to leapfrog digitalisation, agility and a working environment. That is building on sand.

A colleague at a bank had the best image for it: “If you only want to set up a hotdog stand at the station, but you have to build the entire station around it first, then the business case for that hotdog stand simply is not there.”

Christoph: Let us assume the foundations are in place and the result still comes out against the rollout. What is the reason then?

Romano: That the tool works at a point that was never the problem. AI is an amplifier, not a direction.

A team with clean requirements and a working review process gets amplified towards impact. A team without those foundations gets amplified towards chaos. Same tool, opposite outcome.

And then there is the number that occupies me most. At the best software companies in the world, roughly two thirds of all ideas produce zero or negative value. A tool that implements every idea faster is then not progress. It is an accelerator for waste.

Christoph: That inverts the investment logic. Almost everyone calculates with saved development time, because for decades that was the scarce and therefore expensive quantity. Once that quantity becomes cheap, half the business case is built on the wrong factor.

Romano: That is the real shift of our time.

For decades the scarce resource was development capacity. Almost all of our management is built on it: estimates, contracts, project economics, workforce planning. When agents produce code cheaply and in volume, that resource stops being the binding one.

Expensive execution used to be the natural brake against the wrong problem. That brake is gone. The cheaper the building, the more expensive the wrong problem.

Christoph: Which turns a technical question into a leadership question.

Romano: Yes, and that is the core of my book.

AI is the nervous system of an organisation. The soul remains the people. A machine can check, compare and recommend. It cannot decide what matters, and it cannot carry a consequence.

That is why human-in-the-loop is not a precaution for me, it is a design principle. Who owns the business KPI, who decides, who is allowed to stop? If that role is missing, we have an operating model problem translated into software.

Christoph: Your core thesis in one sentence?

Romano: The next competitive advantage is not the AI, it is the organisation around it.

Everyone gets the models. The difference comes from whether an organisation closes its feedback loops, knows its actual bottleneck, and has people who carry a decision.

Speed without feedback loops is not progress, it is risk.

Christoph: What do you advise somebody who has to decide on a broad rollout next week?

Romano: Three things, in this order. The first two are value stream mapping at their core. A method that has existed for decades and almost never comes up in AI discussions.

Make the flow visible. From customer request to delivery, and with the wait times between the steps, not just the working times. Ask five of your people individually where they lose the most time and confidence. Individually, because in a group everyone tells the official version. If the answers do not point at the step your tool accelerates, you already have your decision.

Then measure at the end of that flow. Usage and satisfaction always rise. Lead time to the customer and rework do not.

And name the person who is allowed to stop after the result. If that person does not exist, save yourself the trial and roll out honestly. At least then everyone knows it was a procurement and not an experiment.

“A signal only becomes valuable when it triggers an action that makes the next decision more reliable.”

The Common Core
#

An AI initiative is not ready to scale just because it makes something locally faster.

It becomes ready for a decision when the smallest meaningful trial tests an expensive assumption across the whole value stream, and when it is settled beforehand which result leads to scaling, changing or stopping.

As a short decision rule: clarify the goal. Identify the expensive assumption. Measure the whole system. Bound the smallest trial. Fix the consequence in advance. Assign the accountability.

Takeaway
#

Before the next rollout or budget gate, the accountable person should be able to complete one sentence:

If we observe [observation] in the limited trial, we will [scale / change / stop], because that makes assumption [assumption] for strategic goal [goal] either more reliable or untenable. Accountable for this decision is [role].

If the organisation cannot formulate that sentence, what it lacks is not automatically a bigger pilot. What it lacks first is a learning loop that can produce a decision.

A Note on How This Came About
#

This dialogue is not a verbatim transcript of a live interview. It was editorially developed from a real professional exchange and from my documented positions. I reviewed the statements attributed to me, adjusted them where needed, and approved publication. Publisher and editorial work: Christoph Gulden, 2026.

Thanks to Christoph Gulden for the conversation and for the disagreement in the right place. The original is in his newsletter Thought Leaders Talk on LinkedIn.