Decision Intelligence

The pilot that proved nothing

It ran on clean historical data, in an isolated environment, against a metric agreed afterwards. It succeeded. It told you nothing about whether this works.

Pilot that proved nothing: bu ne anlama geliyor

Pilot that proved nothing: from data through prediction to a recorded decision

Most AI pilots succeed. Most AI programmes stall after the pilot. The two facts are related, and what relates them is how the pilot was designed.

A pilot that runs on exported historical data, in an environment where nothing is at risk, tests whether a model fits a pattern. That was never seriously in doubt.

What the pilot did not test

Whether the data arrives on time in production. Whether anyone acts on the output. Whether the action is even possible in the system where it has to happen. Whether the organisation accepts a decision that contradicts an experienced person.

Every one of these is where real programmes break, and none of them is tested by a backtest.

A pilot worth running

A real decision, made repeatedly, on live data, acted on, with the outcome measured. Narrow scope, real stakes, and success criteria agreed before it starts.

Such a pilot can fail. That is the point: a pilot that cannot fail has tested nothing.

Why the pilot that fails is the pilot that works

A pilot exists to reduce uncertainty. If it cannot come out badly it reduces none — you learn that a model fits data, which was the least uncertain thing in the programme.

The uncertainties worth resolving are organisational. Will the data land at 06:00 every day, including the day the source system is patched? Will a planner accept a recommendation that contradicts them at their busiest moment? Does the ERP permit the write the decision requires, or does it demand an authority nobody holds?

Each of these can be discovered in six weeks or in year two. The pilot is the choice between those two.

Success criteria belong in writing, before it starts

Not “demonstrate feasibility”. A number, a threshold and a date.

Fill rate on pilot SKUs against a control group, measured after eight weeks, with the share that counts as success stated openly. Override rate below a level agreed in advance. Data ready on time on at least nineteen of twenty days.

The discipline here is not bureaucratic. Criteria written afterwards are always met, because the result is what shaped them.

Scope is not the same as stakes

Narrow scope is right. Low stakes are not. A pilot covering forty SKUs in one warehouse is narrow; if those forty are ordered on the recommendation and someone is accountable for the outcome, the stakes are real.

The common failure is the reverse: broad scope, no stakes. Six hundred SKUs, a dashboard nobody orders from, and a report at the end saying the model was accurate.

What to do with a pilot that failed

Read what it says. Failure at the data layer is an integration problem and is solvable. Failure at the override layer is a trust problem and needs explanation, not a better model. Failure at the action layer — the ERP will not accept the write — is an architecture problem and is the most expensive to find in year two.

A pilot that surfaces one of these in six weeks has done its job. A pilot that surfaces none of them has usually avoided all three.

Request a Demo