The demo? Flawless. The team, ecstatic. The CEO declared it the company’s future. And, as always, someone wrapped up the meeting with that familiar phrase: “We’re finally moving forward.”
The pilot ran for four months and cost less than a single trade show booth. Then autumn arrived, and with it budget season. Suddenly, the usual questions surfaced: What’s the daily run cost? Whose cost centre pays? And, most importantly, how much real money did we save - not just impressions, but hard numbers? The room went silent.
The pilot wasn’t rejected. No one vetoed it. No dramatic emails. It just quietly vanished from the spreadsheet. That’s how most AI projects die: not in IT, but in the budget. If you want a pilot to survive, treat it like a project from the start.
Nine in ten use it; three in ten can see it
At the end of August, McKinsey published its global survey on the state of artificial intelligence, conducted in May and June of this year across 1,719 respondents in 97 countries.
Almost nine out of ten organisations now use AI in at least one part of their business. Eighty percent say they’re more productive, half say their decisions have improved, and the number of companies scaling AI, not just piloting it, has nudged up from 38 to 44 percent.
Yet only 37 percent can attribute any effect to their company’s EBIT. That figure is practically unchanged from a year earlier, and the gap remains.
That’s the crux of the problem, in two numbers. The benefits are real, but only at the individual level. Your controller builds reports faster, your sales rep drafts proposals quicker, everyone’s happy. But none of this hits the income statement, because those saved hours aren’t tracked, reallocated, or turned into anything you can measure. If you don't measure the gains, they stay invisible.
Percentages that don’t measure the same thing
You have probably read somewhere that 95 percent of AI pilots fail. That figure comes from a report by MIT’s NANDA initiative, and it says the pilot produced no measurable effect on profit during the period observed. The study was not peer-reviewed, and parts of the academic community have publicly challenged its methodology.
RAND, the American research institute, puts failure above 80 percent, but that finding is qualitative, drawn from interviews with around sixty engineers, and should be read as “the large majority” rather than a decimal. S&P Global measures something else entirely: how many companies abandon most of their initiatives. That number jumped from 17 to 42 percent in a single year. IBM asked two thousand chief executives how many initiatives had delivered the expected return and got 25 percent.
Ninety-five, eighty, forty-two, twenty-five. These numbers don’t contradict each other because they don’t measure the same thing. If you’re about to use one as your argument, check the definition first. In finance, that’s second nature. But in tech, we forget.
What they all say, in the same voice, is this: the tools work; the organisations don’t.
A pilot is an experiment; production is an investment
Pilots and production are approved under two completely different regimes, and companies notice this far too late.
A pilot passes as an experiment: a small amount, usually from an innovation budget or somebody’s reserve, with no formal business case, justified by the phrase “we have to try it.” IBM’s research shows that nearly two-thirds of chief executives openly admit they invest in technologies before they understand the value those technologies bring, simply because they fear falling behind. That is not a character flaw. It accurately describes how pilots get funded.
Production is treated as an investment. Here the rules change: payback period, an owner, a cost centre, a multi-year projection. The CFO rarely says no. The CFO says: bring me the numbers. It sounds like an open door, and it is, in fact, a verdict, because a team that never measured the baseline before the pilot has nothing to build that case from. You cannot claim you cut a process by thirty percent if you never measured how long it took before.
So pilots aren’t usually killed; they’re put off. But defer something twice, and it’s as good as dead. The takeaway is simple: treat delay as a budget decision, because delay often ends the project.
Five questions before you sign
Which leaves the question of why nobody has those answers. Not because people are careless, but because the company, as a rule, does not know its own starting position.
The typical planning error looks like this. Of six months allocated, five go to building the solution and one to “going live.” In reality, everything that constitutes production - integrations, data cleanup, access rights, monitoring, quality control of outputs - takes roughly as long as the solution itself. This becomes clear in month five, when the money is spent, and management’s patience is gone.
RAND’s root causes of failure run in this order: the problem was framed wrongly, the data was inaccessible or poor, the technology was chased instead of solving the problem, and the infrastructure was inadequate. Each of those would have been visible in advance if anyone had looked.
So, before greenlighting any pilot, here are five questions I’d ask. None are technical. None require an outside consultant.
Where exactly does the data this solution needs live, and what condition are they in?
Who owns them - by name, not by department?
Which specific process are we changing, and is that process documented at all?
What is the measured value today, before we do anything?
Who pays for this next year, if it works?
A company that can answer all five isn’t necessarily digital-savy; it’s just honest with itself. If most of these questions get blank stares, forget the pilot. First, figure out what you really have. Answer the basics before you approve the project. You’ll do that inventory eventually; it’ll just cost more once the project stalls halfway.
Why finance pays off first
If those five questions sounded familiar, you probably work in finance. That’s no coincidence, and it leads to a conclusion I rarely hear in conversations about artificial intelligence: the finance function is the best place for a first serious project, precisely because it already has everything missing everywhere else.
The baseline exists, because we measure how long the close takes and how far the forecast lands from the actual. The indicators are defined. The processes are documented, because the auditor would not accept them otherwise. The owner of every number is known, because somebody signs off on it. Forecasting, variance commentary, and preparing materials for management are tasks with clear inputs, clear outputs, and a clear test of success.
Put simply: in finance, it’s tough to fudge the numbers, which is exactly why it’s the best place to prove results first.
Then you look at Eurostat’s figures for Serbia. A tenth of domestic companies use some AI technology, compared with a European average of 20 percent. Of those that do, only 4.6 percent apply it in accounting and finance. The bulk goes to marketing and sales — where the effect is hardest to isolate and easiest to overstate.
The cost nobody put in the spreadsheet
One more item is almost always underestimated, and it started to bite this year.
A fifth of respondents in the McKinsey survey report that the operating costs of running AI, including token costs, forced them to limit its use. The explanation is simple: the price per token is dropping, but the number of tokens consumed is rising faster than the price falls. A pilot with ten users and production with three hundred are not the same cost multiplied by thirty. It is a different calculation.
That calculation also includes a crucial detail: where does the solution run? For companies serving banks, hospitals, or the public sector, data location isn’t an abstract debate; it’s a regulatory and contractual must, and it comes with a price tag. If your usage is high and steady, owning the infrastructure makes sense. If not, rent it. But make no mistake: this is a planning question, not an IT one.
Not against experiments - against experiments with no name on them
To be clear, I am not against trying things. Harvard Business Review warned last year about the trap companies already fell into once, during the digital transformation era: the damage isn’t done by experimentation; it’s done by unfocused experimentation, ten thousand flowers planted and not one gardener. The point is not that there should be no pilot without a proven return up front. The point is that there should be no pilot without a measured baseline, without a name beside it, and without an estimate of what comes after it costs. Make every experiment accountable before you start it. If you can’t name it, measure it, and own it, don’t start it.
The six percent of companies in the McKinsey survey that genuinely see AI in their EBIT aren't distinguished by their models; they use the same ones as everyone else. They are distinguished by being twice as likely to have a defined process for measuring the effect of their initiatives, and by the fact that nearly three-quarters of them redesign the work process itself, against just a quarter of everyone else. They do not add a tool to the existing way of working. They change the way of working. That is the point: measure the effect, redesign the work, and make the result visible. If you want AI to show up in the numbers, stop treating it like a tool and start treating it like a change in work.
And one more figure, perhaps the most important one for us. Among companies with over a billion dollars in revenue, the share of companies scaling AI agents jumped from 27 to 40 percent in a single year. Among smaller companies, it stayed at 22 percent, unchanged. The gap is not closing. It is widening. In a country where small and medium enterprises make up 99 percent of the economy, the challenge is not abstract; it is local. It is ours. And if we want AI to move from pilot to production, the work has to begin there.
A large company that gets it wrong on five million euros gets another chance next year. A small one that gets it wrong on fifty thousand concludes that “this doesn’t work here” and goes back to the old way for another three years. That is how the gap compounds: the bigger company can absorb the mistake, while the smaller one often cannot afford a second try.
So the real question isn’t if you’ll try; everyone will. It’s whether your experiment has a name, a number, and a spot in the spreadsheet. If not, it’s not a project. It’s a hobby, one your company is paying for.