A board buys the licences, a team runs three pilots, the demos go well, and a year later nothing has changed in the numbers the company reports. This is now the most common situation I get called into, and the diagnosis is almost always the same.
The pilots are not failing. They are succeeding at the wrong step.
Measure one thing first
Before any of the debate about models, platforms or governance, measure the share of lead time consumed before anyone writes code. Take twenty initiatives that reached production in the last year. For each, mark the date the idea was first proposed, the date a specification existed that someone could build from, and the date it went live.
In the technology organisations I have taken through this, the first gap is consistently larger than the second, and it is not close. Deciding, arguing about ownership and defining precisely enough to start routinely consumes well over half the elapsed time, sometimes considerably more.
That number is the whole story. A tool that makes the build faster is operating on the smaller half. And because it makes building cheap, it invites more half-defined work into a system that was already carrying more than it could finish.
What the measurement usually exposes
Nothing is ever killed. The business case functions as an entry ticket. Returns are never revisited, no initiative is stopped for underperformance, and the portfolio grows until the number in flight is several times real capacity. Everything then moves slowly, which is read as a delivery problem rather than a portfolio problem.
Nobody owns anything end to end. Each function optimises its own stretch and hands over a date rather than a working outcome. Ask who is accountable for whether a given capability actually produced the revenue it promised, and the answer is a committee.
Definition happens during development. Requirements arrive late and imprecise, and get resolved by developers guessing. This is survivable with humans, who ask questions and apply judgement. It is not survivable with an agent, which will confidently build the thing you described rather than the thing you meant.
Review capacity was never planned. Generation goes up by an order of magnitude. The number of people who can competently judge whether generated work is correct stays exactly the same. The queue does not disappear; it moves from writing to waiting, and it is now staffed by your most senior people.
Controls only ever accumulate. Every incident in the last decade added a gate. None was ever removed. Each was reasonable on the day. Together they mean that shipping something small costs nearly as much as shipping something large, which is why nothing small ever ships.
The uncomfortable implication
None of the five is an AI problem, and none is solved by a better model, a different vendor or a centre of excellence. They are all decisions about how the organisation works, which means they are slow, political and owned by people who did not ask for an AI programme.
This is why so many AI initiatives quietly relocate into tooling. Tooling is purchasable, demonstrable and nobody has to change their behaviour. It also does not work, and eighteen months later the board asks the question again with less patience.
What to change first
Start with the portfolio, because it is the fastest to move and it unblocks everything else. Cut work in progress hard, to something close to genuine capacity. This is unpopular and it is the single highest-leverage action available, because every other improvement is invisible while everything is queued behind everything else.
Then make definition a named job with a named owner. Not a business analyst function restored out of nostalgia, but specification treated as the discipline that determines output quality, done by people close to the business and assisted by the same models everyone is excited about. If an organisation cannot write down what it wants precisely, it has no mechanism for turning an agent into value, and no amount of tooling will supply one.
Then plan review capacity explicitly, as a number, in the same way you would plan build capacity. If generation is going up tenfold, the constraint is now judgement, and judgement has to be developed deliberately because it was previously acquired as a by-product of writing code that nobody writes any more.
Then, and only then, worry about which model you are using.
A reasonable test
If someone proposes an AI initiative, ask two questions. What specification will the agent act on, and who wrote it? Who will review the output, and do they have the time allocated?
An initiative that cannot answer both is a pilot. It will demo well, it will not reach production, and the reason will have nothing to do with the technology.