Problem Statement: Forecasting Demand for Tens of Millions of Store-SKU Series
Frames retail demand forecasting as a hierarchical, intermittent, promotion-driven ML production system, not a single model call.
Problem statement
Design a production pipeline that ingests daily and weekly retail sales history for thousands of SKUs across many stores, trains and tunes forecasting models per SKU or per store cluster, publishes quantile forecasts for upcoming windows, absorbs mid-cycle data (promotions, weather, partial closures), and feeds a replenishment engine that reorders inventory before stockouts. The brief names Prophet, LSTM, and Transformer-family models; a strong answer treats those as candidates inside a model zoo rather than as the architecture.
The defining property of this problem is concurrency across series. A retailer does not forecast one time series; it forecasts a hierarchy. The M5 forecasting competition, run on Walmart's published sales hierarchy, exposed 42,840 related series over 1,941 days: item-store leaves rolled up through department, category, store, and state. Walmart reports roughly 10,500 stores and clubs worldwide and serving about 270 million customers each week, and a single supercenter carries on the order of 120,000 SKUs, while Costco publicly runs its warehouses at roughly 3,800 SKUs. Those public figures bracket the design space: the number of leaf series is store-count times active-SKU-count, and every architectural choice below is a consequence of that product.
Why this is not a generic batch job
Three forces separate retail forecasting from textbook time series. First, intermittency: most store-SKU pairs sell zero units on most days, so mean-squared-error models collapse toward zero and service levels die; Croston-style or quantile models are required. Second, calendar and promotion effects: holidays shift, promotions spike demand 3x to 20x, and a missing promotion flag is a silent accuracy killer. Third, hierarchy coherence: a store-level forecast that disagrees with the chain-level plan produces purchase orders that finance and supply chain reject, so reconciliation is a first-class compute stage.
The four planes
- Data plane: POS close-out ingestion, cleaning, calendar and promotion alignment, hierarchy registry, feature store.
- Training plane: series tiering, model zoo, hyperparameter search, evaluation gates, model registry with lineage.
- Serving and planning plane: batch quantile inference, hierarchical reconciliation, atomic forecast publication, planner overrides, replenishment consumption.
- Governance and learning plane: accuracy and calibration monitoring, drift detection, drift- or calendar-triggered retrains, override value-add accounting, audit.
Public baseline versus design assumptions
Public evidence proves the category operates at scale: Amazon published DeepAR, a global probabilistic model trained across groups of related series, and shipped GluonTS plus a managed forecasting service whose documented algorithm menu included DeepAR+, NPTS, Prophet, ARIMA, and ETS. Meta published Prophet from its Forecasting at Scale work. Google Research published the Temporal Fusion Transformer with retail-style exogenous inputs. For capacity planning this answer explicitly assumes a mature regional retailer with 1,800 stores, 50,000 active SKUs, 22 million active store-SKU pairs, and a nightly close of 38 million POS line items. Unless tied to a named public source, every number in this answer is a stated assumption, budget, or target.
Key Highlights
- •The unit of scale is the store-SKU series count, not request QPS: M5 published 42,840 related Walmart series over 1,941 days.
- •Intermittency, promotion calendar, and hierarchy coherence are the three forces that break naive model-per-series designs.
- •The architecture splits into data, training, serving/planning, and governance planes with explicit ownership.
- •Amazon DeepAR, Meta Prophet, and Google TFT are cited as public evidence that global models, additive baselines, and attention hybrids all belong in the zoo.
- •All uncited scale values are explicit design assumptions: 1,800 stores, 50,000 SKUs, 22M active pairs, 38M POS lines per day.
Section Rescue Kit
Buzzwords to use:
Safe statements:
- "Let me size the problem as a hierarchy of series before choosing any model family."
- "I will separate what the data pipeline must guarantee from what the model must learn."