Retail & consumer

Cutting stockouts across 240 stores with a forecast the buyers trust

A hierarchical demand forecast, a promo-aware feature store and a replenishment workflow the category team can override — because a model nobody overrides is a model nobody uses.

Stockouts on A-class SKUs
cut by roughly a third
Forecast accuracy (WAPE)
mid-thirties → high teens
Buyer overrides accepted
zero → most plans
Sector
Multi-category retail, India
Estate
240 stores, 3 DCs, ~28,000 SKUs
Engagement
Discovery sprint, then delivery pod
Duration
22 weeks to first full rollout

Stack

  • Airflow
  • dbt
  • Snowflake
  • LightGBM
  • Prophet
  • Feast
  • FastAPI
  • React
  • Power BI

Practices involved

Discuss a similar problem

The situation

The chain ran replenishment from a weekly spreadsheet cycle. Category managers pulled sales from the ERP, adjusted by memory for festivals and promotions, and emailed order quantities to three distribution centres. The process worked when the business had 60 stores. At 240, the same three people were making roughly 28,000 decisions a week, so most SKUs were simply reordered at last week's number.

The symptoms were the ones you would expect: fast movers went out of stock in the last four days of the month, slow movers accumulated in the DC, and the two problems were invisible to each other because nobody reported them together.

The constraint

An earlier forecasting pilot had been rejected. It was more accurate than the spreadsheet on paper, but it produced numbers the category team could not explain to their own management, and it had no way to absorb the thing they knew and it did not — a competitor opening opposite Store 118, a distributor strike, a school holiday shifted by a week.

So the requirement was not really accuracy. It was accuracy the buyers would accept ownership of.

What we built

A feature store that knows what a promotion is

Promotions lived in three systems and a shared drive. We modelled them once — mechanic, depth, participating SKUs, store list, date range — and made that the single input to both training and serving. Festival calendars, local holidays, weather and competitor-opening flags landed in the same place. Training-serving skew disappears when both paths read the same table.

A hierarchical forecast, reconciled

Separate models at SKU-store, SKU-cluster and category-region level, then reconciliation so that the sum of the store forecasts equals the number the category head is held to. Gradient-boosted models on intermittent-demand SKUs, a seasonal statistical model on the long-tail slow movers where there was not enough signal to justify anything heavier.

An override that teaches

The replenishment screen shows the recommendation, the three factors that moved it most, and a field for the buyer to change it with a reason code. Every override is stored. Two things happen with that data: recurring reason codes become candidate features, and the acceptance rate per buyer became the metric we managed the rollout by.

A rollout designed to be reversed

Six weeks of shadow mode with no orders placed, then one region, then three, then all. At every stage the previous process stayed live and switchable within a day.

What changed

The number the client cares about is not the WAPE improvement, it is that most replenishment plans now go through with a buyer adjustment rather than a wholesale rejection. The forecast became a starting position instead of an opinion competing with theirs.

What we would do differently

We spent four weeks building a clean master-data mapping between ERP item codes and the point-of-sale catalogue before we understood how often the mapping changed. It changed weekly. We should have built the reconciliation job first and the model second; instead we rebuilt the mapping twice.

Outcomes

Stockouts on A-class SKUs
cut by roughly a third
Forecast accuracy (WAPE)
mid-thirties → high teens
Buyer overrides accepted
zero → most plans

Client identity withheld under a mutual NDA. Figures are illustrative — rounded and directional, meant to show the shape of the change rather than an audited result. We will walk through the real numbers, and how they were measured, under NDA on a call.

Next step

Tell us what you're trying to ship.

Send the brief, the RFP, or three messy sentences about the problem. You get a written point of view from an architect within two working days — not a sales deck.