A proof of concept is quick to build these days, and that is exactly where the problem starts. Classically around 70 percent of all PoCs got stuck in the PoC stage. With AI, going by the studies Steffen follows, the rate sits beyond 90 percent.
In episode 5 of Prompt & Proper, Steffen Ochsenreither from Swiss Post and I talk about what actually kills these pilots. Technology has the least to do with it. The episode is in German, and this post captures the core of the conversation.
Why There Are Suddenly So Many More Pilots#
Steffen brought the 70 percent along from his digitalization years, where it was the magic number. With AI adoption the rate has gone up. The studies land at different heights, the direction is unambiguous.
Two reasons for it. The barrier to entry has dropped massively. You used to need knowledge to digitalize something, today you take a tool that already does a lot out of the box and get going. On top of that comes the hype: people are primed for it and want to do something. So far more gets started.
The core problem has stayed the same, and that is the uncomfortable truth. Every PoC is cool at first, and when something works you are genuinely pleased. Then the same questions arrive. Is there a concrete business case? Are the people on board? I watch PoCs fail constantly too, and it has nothing to do with AI. It was the same with cloud computing, and with the internet before that.
A PoC Without a Strategic Goal Does Not Reach Production#
A proof of concept validates a business case. A company has strategic goals, it wants revenue, it wants margin. If your PoC pays into none of those goals, I can nearly guarantee it will not go into production.
Steffen supplied the line for it: “we want to have AI” is not a strategy. The usable formulation is, I want to do this thing better, and I will use AI for that. If someone had written “we want to digitalize” into a strategy paper back then, it would have worked just as poorly.
Technology, Strategy, Culture#
Steffen breaks the topic into three elements. You need a technology that can do it. You need a strategy that supports it. And you need a culture that carries the changed way of working.
Nobody has to worry about the technology any more. There is a tool for practically every use case, and where there is not, you assemble one fairly easily. The models also keep getting better. Steffen mentioned an Anthropic study in which Opus solved close to 90 percent of the more complex tasks. Every few months a new model can do something we used to consider very complex.
Strategy and culture are the hard parts. And that is where the things sit that were killing PoCs long before AI. How do I operate this thing? How do I introduce the change? How do I train the people? Who runs first, second and third level support? All of that still shows up. Too often we let ourselves be dazzled by presentations where a ChatGPT cowboy shows how he builds himself a CRM in no time. It is impressive. It is also the same thing as sitting in Microsoft conferences years ago while somebody demonstrated how to spin up an Access database with a few clicks. The effort behind the strategic goal gets trivialized in those demos.
On the culture side came the sentence I wrote down:
“Usage does not equal adoption.”
Trust in the decisions the AI makes, guardrails and guidelines, digital ethics. And in the end the question of how your people actually work with it. That is exactly where many companies get stuck.
The LLM Will Always Lie to You#
“An LLM is only a token prediction engine. Which means the thing will always lie to you. It will always hallucinate.”
There is no way out of that. You can prompt as well as you like and write “don’t hallucinate” in capital letters, it happens anyway. What helps is a harness around it, with checks and loops that reduce the effect. It burns extra tokens, and it will not get you to 100 percent either.
Steffen sees it the same way: with LLMs, 100 percent is effectively not possible today. For business-critical processes that is a real problem, so you need guardrails that keep you from opening a sales office in a country you never wanted to enter just because the model suggested it. His line was that you are still allowed to think. I pushed back: you have to. Switching your brain on is business-critical with these systems.
One clarification: when Steffen talks about AI in a PoC, he means more than LLMs. The more classical machine learning and highly refined algorithms count too. We come back to that later.
Data Quality, the Least Sexy PoC Killer#
Digitalization left all of us with large volumes of data. Not much of it is clean. So for a PoC you first have to clean the data or provide a usable data set.
The most popular entry-level PoC is the company chatbot on your own data. It is a no-brainer, so plenty of people build it. And plenty make the same mistake: they point it at the SharePoint where all the data garbage lives. Then interesting things happen.
Steffen calls data quality and building data products the least sexy topic there is, and the one you cannot do without. If data is not democratized and not everybody has the same access, you get different results depending on what you build on top. The same query gives two people two answers. Putting an LLM on poor data management may help you find things faster, it does not fix the problem underneath.
Without a Foundation You Are Building on Sand#
With data we are already close to the foundation. I was recently at a company that wanted to run a PoC which has to go into a production environment. That company is only now building its cloud environment, nothing runs in production there yet. For the PoC it would be necessary.
I see that more often. In the AI era, companies try to leap over everything that should have come first. Agility, digitalization, clean operations. You can still build the PoC, it just gets harder, and often that is precisely what prevents the go-live.
Steffen sees this as the main difference between start-ups and grown companies, particularly in Europe. A start-up founded three years ago has everything in the cloud, far less data and therefore better data quality by default, and no analog processes. At a company that is 50 or 100 years old it looks different. The focus sits on the core process of the business. Now AI is supposed to go in because everybody is doing AI somehow, and out come isolated PoCs serving one small purpose. The cases where AI would genuinely be strong, back-end processes, entire supply chains, decision support, are left untouched.
This fear of missing out is running pretty wild right now. With AI we can build everything faster and better, the story goes, and along the way people forget what has to be built underneath first. You go to the cloud when it pays into a strategic goal. If you are fast enough and good enough on premise, that is a legitimate answer. Otherwise the build-out around it kills the business case. A colleague of mine put it this way at a bank:
“If all he wants is to put a hotdog stand at the station, but he has to build the entire station around it as well, then the business case for that hotdog stand simply is not there.”
Steffen’s comment: fairly expensive hotdogs.
The way out is the small business case you can copy often. When you put something in front of many people and they work with it, you end up with a high impact you cannot see the first time round.
AI Is Much More Than Generative AI#
Artificial intelligence is an enormous term. It starts in the 1950s with rule-based systems and simple if-then-else queries, and that already counts as AI. Machine learning builds on top of it, your spam filter works that way. Then come deep learning and computer vision, which let us tell cats from dogs.
There are early studies suggesting that around 90 percent of value generation in companies worldwide comes from exactly those areas. About 10 percent originates with the large language models. The entire hype sits on those 10 percent. That is why I tell a lot of companies to look at the other things. In industrial companies the cameras have been standing there for years, and for the processes in that setting machine learning or deep learning is often the better fit.
Steffen even doubts the 10 percent. He had a survey in mind and still owes me the source: fewer than 10 percent of AI users have ever let their model think for longer than a minute. We have models that handle open-ended questions with deep thinking, and we ask them about tomorrow’s weather. For tasks like that you need zero generative AI.
For me the consequence is simple. I look at the business process and ask whether it has to run correctly 100 percent of the time. If the answer is yes, generative AI is probably the wrong approach. You still use AI, only then it is rule-based, machine learning or deep learning. After that it is classical engineering: show what we want to achieve and how we get there.
The Take-Away#
Steffen’s recommendation for anyone who wants fewer PoCs to die:
“Experiment with purpose.”
Experimenting is good, you should get to grips with the technology. The difference sits between an experiment with a purpose and messing around. If you simply let people run free, maybe one in 100 cases turns up something that helps the company.
On top of that come three things Steffen would pass on to everyone. Do not start with the giant use case, run one focused PoC instead. Pick the right technology for your use case rather than hitting everything with the GenAI hammer. And build a change culture: handing people a tool is not enablement, you have to teach them how to work with it. Otherwise the effect stays short-lived.
Do that, and you end up killing considerably fewer than 90 percent of your PoCs.
If you enjoy the episode, subscribe to Prompt & Proper on Spotify, Apple Podcasts, Amazon Music or YouTube and tell us which AI topic you are wrestling with right now. All episodes: promptandproper.ai.
