The bottom line
In most FMCG businesses the demand plan is a spreadsheet with a statistical column nobody trusts and a judgement column that overrides it — and nobody can prove which is right, because last quarter's accuracy was never measured. A six-week starter sprint on Microsoft Fabric builds a baseline statistical forecast for one bounded slice, at a stated grain, and backtests it against your current plan using WMAPE and bias — producing an evidenced proceed decision, not a system. It excludes deployment, integration, retraining and new-product models; those belong to rollout, funded only once the sprint has answered the question. Measure forecast value add, never a single headline accuracy figure.
In This Article
A statistical column nobody trusts, a judgement column that overrides it
In most FMCG businesses I walk into, the demand plan is a spreadsheet with a statistical column nobody trusts and a judgement column that overrides it. The judgement column is usually right about the big things and wrong about the small ones — and nobody can prove which is which, because last quarter's accuracy was never measured.
The plan is built on primary sales — what was dispatched to distributors — because that number arrives cleanly and on time. Consumer demand is what you actually need to predict, and it is the number you have least of.
The consequences show up on the balance sheet before they show up in a report: working capital sits in the wrong SKUs while the fast movers stock out. This is a scope document, not a brochure — what a six-week demand forecasting sprint on Microsoft Fabric genuinely delivers, what it does not, how the result should be measured, and what a full rollout adds.
What a demand forecasting starter sprint is
A demand forecasting starter sprint is a six-week, fixed-scope engagement that builds a baseline statistical forecast for one bounded slice of the business — one category or one region — at a stated grain, backtests it against your current planning baseline using WMAPE and bias, and produces an evidenced proceed-or-not decision.
The framing matters commercially: you are not buying a forecasting system in six weeks. You are buying a defensible answer to whether a forecasting programme is worth funding, backed by a backtest on your own data.
The tooling is Fabric Data Science — notebooks reading Delta tables in OneLake, MLflow experiment tracking and the model registry, forecasting libraries such as Prophet added through a Fabric environment, AutoML on FLAML for a fast baseline, PREDICT for batch scoring, and Power BI in Direct Lake mode for consumption.
What six weeks genuinely covers
A sprint runs on a bounded slice — one category, region or channel, not the whole portfolio.
- Weeks 1–2 — historical data assessment: assemble primary sales, secondary sales where uploads support it, price, the scheme/promotion calendar, stock positions and the master data that joins them
- Weeks 2–4 — a baseline statistical forecast at a grain agreed in writing before modelling (SKU–region–month, or SKU–distributor–week)
- Weeks 3–5 — a promotional uplift view, where the history genuinely supports it
- Weeks 4–6 — backtesting against the current planning baseline: hold out recent periods, forecast as of the historical decision date, and compare against what planners actually submitted
- Week 6 — the proceed decision: which SKU segments the model beats the current plan on, which it does not, what the data would need to look like for the rest, and a costed rollout scope
What a sprint explicitly excludes: production deployment, integration into planning or replenishment systems, automated retraining, new-product introduction models, hierarchical reconciliation across grains, a consensus planning process, distributor onboarding for missing data, and any guaranteed accuracy figure.
Starter sprint vs full rollout
The sprint answers the question; the rollout acts on the answer. They are different engagements with different risk profiles.
| Dimension | Starter sprint (6 weeks) | Full rollout (4–8 months, staged) |
|---|---|---|
| Scope | One category, region or channel | All categories and channels, phased |
| Grain | One stated grain | Multiple grains, reconciled to be coherent |
| Models | Baseline statistical; uplift if history supports | Segmented portfolio; ML where data volume justifies |
| New products | Excluded | Analogue-based, with a written selection rule |
| Promotions | Retrospective uplift view | Forward-looking scenario planning on the trade calendar |
| Validation | One backtest vs the current plan | Rolling backtests plus live accuracy each cycle |
| Process | None — analysis only | Consensus demand review, named owner, logged overrides |
| Integration | None — outputs read in Power BI | Forecast written back to ERP/planning; feeds replenishment |
| Output | A decision, evidenced | A forecast that changes POs and production plans |
| Risk | Bounded and small | Real — worth taking only after the sprint answers the question |
Whether your data can support a sprint
Some businesses should not run this sprint yet. Five checks, answerable this week from your own systems:
- History depth — at least 24–36 months at the intended grain, so the model sees two or three cycles of the same seasonality
- Demand signal — secondary sales, or only primary?
- Upload completeness — if secondary is the signal, what percentage of distributors uploaded complete, on-time data in each of the last six months?
- Promotional record — a scheme master with dates, mechanics and SKU coverage that someone who was not in the room could reconstruct
- Master-data stability — does a single SKU code survive a pack change and a promotional variant?
Failing the upload and promotion checks does not stop a sprint — it narrows what the sprint can claim. Failing history depth and master-data stability usually means the six weeks are better spent on the lakehouse foundation first.
How accuracy should be measured, and reported
This is where most forecasting proposals become unfalsifiable, so be precise. Use WMAPE, not MAPE — weighted mean absolute percentage error weights errors by sales volume, so a 40% miss on a rounding-error SKU does not flatter the headline. Report bias alongside error: a forecast with acceptable WMAPE and persistent positive bias builds inventory every cycle.
Measure forecast value add against a naive baseline — FVA captures the change in performance attributable to a step or participant, with a naive forecast as the control. And remember the grain determines the number: accuracy at category–month grain and at SKU–week grain are different measurements, and the second is materially worse.
A single headline accuracy percentage is a warning sign. If someone quotes one number without asking your grain, horizon, promotional intensity and SKU mix, they have not measured anything.
Where this breaks — what a forecast does not fix
Promotional intensity sets a ceiling: in a portfolio where a large share of volume moves on scheme, baseline demand is a small part of the signal. Distributor upload gaps become model input — if a distributor stops uploading for three weeks, the model reads it as a demand collapse. Structural breaks defeat history: a new distributor in a state, a price reset, a competitor entry, a pack change, a regulatory shift.
Two organisational truths matter more than any model. A forecast nobody owns changes no decision — the most common failure, and it is organisational, not technical. And a better forecast only pays back if supply can act on it: if your minimum production run is a month of cover, your import lead time is ten weeks, or your co-packer needs six weeks' notice, improving a four-week forecast has limited financial effect. Fabric also runs on a capacity model — training and backtests over multi-year SKU-level history consume Spark compute, so budget for it.
What to do first
Answer four questions this week, with your own data:
- What is the WMAPE and bias of your current plan over the last six months, at the grain you actually plan on? If nobody has calculated it, that is the first finding
- Which single category or region has the cleanest 24–36 months of history and the best-documented scheme calendar? That is your sprint slice
- Would a better four-week forecast change a purchase order, a production run or a promotional commitment — and how quickly can that decision be changed?
- Who owns the number? Name the person, not the department
If the third question has no answer, do not run the sprint. If the fourth has no answer, fix that first — it costs nothing and determines everything else. We build these on Microsoft Fabric, OneLake and Power BI, with Power Automate closing the loop from forecast to action, under a Fractional Data Consultant model.
Start with the honest number: the WMAPE and bias of your current plan, at the grain you actually plan on. If nobody has calculated it, that is finding number one — and a good reason for a 30-minute conversation. Book a diagnostic with Amit — no slides, no pitch deck, no obligation to proceed. You will leave knowing whether your data can support a sprint, and which slice to run it on.
Free Assessment
Where does your operation sit on the data maturity curve?
8 questions. 3 minutes. You get a scored breakdown across data infrastructure, analytics readiness, and automation potential — with a specific next step for your industry.