Demand Intelligence

What data do you need for demand forecasting?

Forecast accuracy depends on data before it depends on the model. What cannot be forecast without, what measurably helps, and what only adds noise.

Demand forecasting data requirements: bu ne anlama geliyor

Demand forecasting data requirements: from data through prediction to a recorded decision

Most demand forecasting projects stall on data preparation, not on model selection. “Which algorithm is better” is an unanswerable question when the data in hand cannot feed either one.

The three you cannot forecast without

Sales history at transaction level. Monthly totals are not enough to build a forecast on; the data has to be weekly, preferably daily, and split by product and location. Aggregated history has already erased the variation you are trying to predict.

Stock-out records. This is the most commonly skipped input, and skipping it produces systematic error. If a product was not sold because it was not on the shelf, the sales data reads “low demand” — when demand existed and went unmet. Without flagging out-of-stock days, the model learns your inventory mistakes and treats them as customer behaviour.

Calendar and promotion history. Public holidays, back-to-school, year end, discount periods. Without them the model sees seasonality as random noise.

What measurably helps

Price history and competitor pricing, promotion type (discount, multibuy, visibility), new product launches and lifecycle stage, and weather — particularly in food, beverage and seasonal goods.

What usually adds noise

Broad macroeconomic indicators and social media sentiment. Both sound powerful; when their contribution to an operational SKU forecast is actually measured it is usually close to zero, and they add complexity nobody can explain later.

How much history is required

At least two full years to capture seasonality. A model built on one year mistakes seasonality for trend.

How data quality is measured

Collecting data is not the same as having data a forecast can use. Three practical checks:

Coverage. What share of product-location combinations has at least two years of unbroken history? Below 60%, the remainder needs a different method — deriving from a similar product profile, or forecasting at product group level.

Continuity. Are there gaps? Missing days caused by a system change, an ERP migration or a store closure are read by the model as zero demand, and pull the forecast down.

Consistency. Did the product change code over time? A code change splits two years of history into two unusable six-month fragments.

Starting without complete data

New products and new locations have no history. Similarity-based forecasting is used here: the launch curve of a product in the same category, price band and location profile is taken as a reference. It gives a workable answer for the first 8-12 weeks, after which the product’s own data takes over.

The point is not to exclude a product from forecasting because it lacks history — the stock decision will be made regardless, either by a model or by a guess.

The common mistake

Waiting for the data quality project to finish. In practice data quality never finishes; it improves through use. A first forecast built on what exists shows which data is genuinely missing, and makes prioritising possible.

The question to answer before collecting anything

Which decision is the forecast for? The answer determines the granularity required.

A purchasing decision needs product-supplier at monthly horizon. An allocation decision needs product-store at weekly. The second is roughly ten times the data, and collecting it for the first is wasted effort.

Most projects begin with “let us gather all the data” and are still gathering six months later. Working backwards from the decision reduces that to weeks.

Data that adds noise

Not every extra field improves a model. A macroeconomic indicator updated monthly behaves as a constant in a model making weekly decisions, and carries no information at all.

The same applies to product descriptions, supplier notes and free-text fields: they contain information, but it cannot enter a model unstructured, and the cost of structuring it usually exceeds the gain.

A simple test: would the decision change if we did not know this field? If not, leave it out.

Demo Talep Edin