Demand Intelligence

Why store-level demand forecasting differs from chain-level

A model that forecasts total demand accurately can be systematically wrong per store. The problem is not the model — it is the level of aggregation.

Store-level demand forecasting: bu ne anlama geliyor

Store-level demand forecasting: from data through prediction to a recorded decision

You may be forecasting a product’s weekly demand across the chain to within 5%. The same model may be 40% wrong at store level. Both are true, and the difference comes from aggregation.

Aggregation hides error

Over-forecasting in one store and under-forecasting in another cancel out. The total looks right, while a wrong decision was made in both places: dead stock in one, lost sales in the other.

Because the stock decision is made at store level, accuracy has to be measured there.

What is different at store level

Local effects. A school, a factory shift pattern, a hospital or a newly opened competitor changes that store’s demand independently of the rest of the chain.

Small numbers. A product selling 5,000 units a week chain-wide may sell 3 units a week in one store. At small numbers, chance is larger than pattern, and classic time series models are weak there.

Stock-out effects are sharper. A stock-out in one store disappears in the chain total; in that store’s own data it erases demand entirely.

Cannibalisation and substitution. A customer who cannot find one item buys a similar one. That corrupts the data of two products at once, at store level.

The practical approach: hierarchical forecasting

Rather than forecasting at a single level, a common and effective method is to forecast across a chain → region → store hierarchy and then reconcile. The stability of the upper level is combined with the locality of the lower.

For very slow-moving items, a probability distribution beats a point forecast: “80% likely between 1 and 6 units” rather than “3 units a week”.

Store clustering

Modelling hundreds of stores individually is impractical; treating them identically is wrong. Clustering is the middle path: stores with similar demand patterns are grouped, forecast at cluster level, and allocated down.

Useful clustering dimensions: store size, location type (mall, high street, motorway), customer profile, seasonal pattern and product mix.

Cluster counts usually sit between 5 and 15. Fewer erases the differences; more reintroduces the small-numbers problem.

New stores and new products

A newly opened store has no history. What works best in practice is to take a starting curve from the closest cluster profile and transition gradually to the store’s own data over the first 8-12 weeks.

Critical point: opening-period data is not normal demand. Launch promotion, curiosity traffic and first-week intensity, once fed into the model, cause systematic over-forecasting for months afterwards.

Which error metric

MAPE is misleading at store-SKU level, because percentage error explodes on small numbers: selling 2 units instead of 1 is a 100% error.

WMAPE (weighted) or MAE (mean absolute error) are more appropriate at this level. Choosing the metric to match the decision level is a precondition for tracking accuracy at all.

Aggregation hides error rather than removing it

A forecast showing 90% accuracy at chain level may be 60% at store level. Errors cancel as they aggregate; a surplus in one store covers a shortfall in another and the total looks sound.

Where it does not cancel is the allocation decision. Stock goes to a specific store and is either there or not. The chain average means nothing to the customer standing in that store.

Is the store difference real

Every store believes it is different, and is partly right. The distinguishing test: does the difference persist across seasons, or did it appear last quarter.

A store deviating in the same direction for three consecutive quarters is a pattern. One that deviated last quarter is a coin flip that landed.

Clusters of comparable stores

For slow-moving products, store-level forecasting collapses for lack of data. The answer is not to revert to a central forecast but to cluster comparable stores: similar size, similar customer profile, similar region.

Enough data accumulates at cluster level, and the store difference is not entirely erased. Which store belongs to which cluster should be recorded, because a category manager may dispute a cluster and is often right.

Demo Talep Edin