AI pilot design: bu ne anlama geliyor
Most AI pilots succeed. Most AI programmes stall after the pilot. Those two facts are related, and the relationship is the pilot’s design.
A pilot on extracted historical data, in an environment where nothing is at stake, tests whether a model can fit a pattern. That was never seriously in doubt.
What the pilot did not test
Whether the data arrives on time in production. Whether anyone acts on the output. Whether the action is possible in the system where it has to happen. Whether the organisation accepts the decision when it contradicts an experienced person.
Every one of these is where real programmes fail, and none of them are exercised by a backtest.
A pilot worth running
One real decision, made repeatedly, on live data, with the action taken and the outcome measured. Small scope, real stakes, agreed success criteria before it starts.
That pilot can fail. That is the point: a pilot that cannot fail has not tested anything.
Why the failing pilot is the useful one
A pilot exists to reduce uncertainty. If it cannot come out badly, it removes none — you learn that a model fits data, which was the least uncertain thing in the programme.
The uncertainties worth resolving are organisational. Will the data arrive by 06:00 every day, including the day the source system is patched? Will a planner accept a recommendation that contradicts them at the moment they are busiest? Does the ERP allow the write that the decision implies, or does it require a permission nobody has?
Each of these can be discovered in six weeks or in the second year. The pilot is the choice between those.
Success criteria have to be written before the start
Not “demonstrate feasibility”. A number, a threshold, and a date.
Fill rate on the pilot SKUs against the control group, measured after eight weeks, with a stated margin that counts as success. Override rate below a level agreed in advance. Data availability on time on at least nineteen days in twenty.
The discipline here is not bureaucratic. Criteria written afterwards are always met, because the result becomes the criterion — and a pilot that always succeeds is a pilot that never informed a decision.
The scope that works
One decision, made at least weekly, on live data, with the action executed and the outcome recorded.
Weekly matters: it produces eight observations in eight weeks, which is enough to see a pattern. A monthly decision gives two, and two of anything is an anecdote.
Live data matters more than volume. A pilot on fifty SKUs with real feeds teaches more than one on five thousand SKUs of extract, because the failures it surfaces are the failures production will have.
What to do with a pilot that fails
Read it rather than repeat it. A pilot that failed on data timeliness has told you the integration is the project, and no amount of model work will change that. One that failed on adoption has told you the output is arriving in the wrong form or at the wrong moment.
Both findings are worth the eight weeks. The outcome to avoid is the pilot that succeeded, told you nothing, and left the organisation confident about a programme that has not yet met its real constraints.
The pattern in stalled programmes
They almost always have a successful pilot behind them, and the stall arrives at the first contact with production conditions.
That is not bad luck. It is the design of the pilot arriving on schedule.
The question to ask before approving one
What would have to happen for us to stop this programme? If nobody can answer, the pilot is not a test — it is a demonstration with a budget.