Decision Intelligence

The pilot that proves nothing

It ran on clean historical data, in a sandbox, against a metric agreed afterwards. It succeeded. It told you nothing about whether this works.

AI pilot design: bu ne anlama geliyor

AI pilot design: from data through prediction to a recorded decision

Most AI pilots succeed. Most AI programmes stall after the pilot. Those two facts are related, and the relationship is the pilot’s design.

A pilot on extracted historical data, in an environment where nothing is at stake, tests whether a model can fit a pattern. That was never seriously in doubt.

What the pilot did not test

Whether the data arrives on time in production. Whether anyone acts on the output. Whether the action is possible in the system where it has to happen. Whether the organisation accepts the decision when it contradicts an experienced person.

Every one of these is where real programmes fail, and none of them are exercised by a backtest.

A pilot worth running

One real decision, made repeatedly, on live data, with the action taken and the outcome measured. Small scope, real stakes, agreed success criteria before it starts.

That pilot can fail. That is the point: a pilot that cannot fail has not tested anything.

Why the failing pilot is the useful one

A pilot exists to reduce uncertainty. If it cannot come out badly, it removes none — you learn that a model fits data, which was the least uncertain thing in the programme.

The uncertainties worth resolving are organisational. Will the data arrive by 06:00 every day, including the day the source system is patched? Will a planner accept a recommendation that contradicts them at the moment they are busiest? Does the ERP allow the write that the decision implies, or does it require a permission nobody has?

Each of these can be discovered in six weeks or in the second year. The pilot is the choice between those.

Success criteria have to be written before the start

Not “demonstrate feasibility”. A number, a threshold, and a date.

Fill rate on the pilot SKUs against the control group, measured after eight weeks, with a stated margin that counts as success. Override rate below a level agreed in advance. Data availability on time on at least nineteen days in twenty.

The discipline here is not bureaucratic. Criteria written afterwards are always met, because the result becomes the criterion — and a pilot that always succeeds is a pilot that never informed a decision.

The scope that works

One decision, made at least weekly, on live data, with the action executed and the outcome recorded.

Weekly matters: it produces eight observations in eight weeks, which is enough to see a pattern. A monthly decision gives two, and two of anything is an anecdote.

Live data matters more than volume. A pilot on fifty SKUs with real feeds teaches more than one on five thousand SKUs of extract, because the failures it surfaces are the failures production will have.

What to do with a pilot that fails

Read it rather than repeat it. A pilot that failed on data timeliness has told you the integration is the project, and no amount of model work will change that. One that failed on adoption has told you the output is arriving in the wrong form or at the wrong moment.

Both findings are worth the eight weeks. The outcome to avoid is the pilot that succeeded, told you nothing, and left the organisation confident about a programme that has not yet met its real constraints.

The pattern in stalled programmes

They almost always have a successful pilot behind them, and the stall arrives at the first contact with production conditions.

That is not bad luck. It is the design of the pilot arriving on schedule.

The question to ask before approving one

What would have to happen for us to stop this programme? If nobody can answer, the pilot is not a test — it is a demonstration with a budget.

Request a Demo