The bottom line
A Microsoft Fabric discovery sprint exists to produce two decisions: whether to build at all, and the exact scope, sequence and fixed price of the first slice. It is not a roadmap exercise. Over two to three weeks it should contain five named activities: source inventory and access provisioning, profiling of production data (full tables, not a 1,000-row sample), metric definition interviews, capacity sizing with the arithmetic shown, and a working technical spike that moves real data end to end. Four deliverables justify the fee — a profiled data assessment naming actual defects with counts, a fixed-price first-slice scope, a capacity recommendation with reasoning, and an architecture decision record. Maturity models, generic roadmaps and unrequested tooling comparisons are filler. A proportionate discovery costs 8–15% of the build it scopes and narrows the estimate band from ±50% to ±15%. A discovery that concludes "not yet" and says what must be true first has saved you the build budget.
In This Article
A technical assessment and a fortnight of workshops look identical on paper
One tells you that your goods receipt table has 41,000 rows with a null supplier code and that OTIF is defined differently by the plant and the commercial team. The other tells you data is a strategic asset and recommends a phased approach — something your own team could have written in a day, for which you paid four weeks.
This is a scope document written by the person who runs these: what a real discovery contains, the deliverables worth paying for, the filler to refuse, and the client time it actually costs.
What is a Microsoft Fabric discovery sprint actually for?
A discovery sprint exists to produce two decisions: whether to proceed with a build at all, and if so, the exact scope, sequence and fixed price of the first slice. It is not a roadmap exercise. Its worth is measured by the accuracy of the estimate that follows and by the defects it finds early.
A roadmap is a statement of intent, and intent is free. An estimate is a commitment, and a commitment is only as good as the evidence under it. Discovery buys that evidence at a fraction of what it costs to find mid-build, when the fix arrives as a change request. The test: at the end, can the partner quote a fixed price for the first slice and stand behind it?
What happens week by week in a Fabric discovery sprint?
Five named activities across two to three weeks. Week 1 — source inventory (half a day) and, more importantly, access. Read credentials on a production ERP need an owner, a security review and a named approver; on-premises sources need the gateway installed inside your network. Ask the partner to send the access request list before the contract is signed — one who does that has run this before.
Week 1–2 — profiling the actual data, not the data dictionary. The output is specific: row counts and date ranges per table, null rates per key column, duplicate rates on every join column, orphaned foreign keys, and the share of records where a mandatory field holds a placeholder. One caution: Power Query profiles the first 1,000 rows by default, so serious profiling runs in a Fabric notebook against the full extract. If the report names no defect with a number attached, ask what was run.
Week 2 — metric definition interviews (the plant, commercial and finance separately on the same three KPIs; you will get three definitions of OTIF), and capacity sizing with the reasoning shown (the F64 licence boundary usually dominates cost more than compute does). Week 2–3 — the technical spike: take one genuinely awkward source, land it in OneLake, model it to the grain the business argues about, and put a number on a page someone recognises. A spike against sample data proves nothing — sample data is clean by definition, because someone chose it to make the demonstration work.
A spike against sample data proves nothing. Sample data is clean by definition — someone chose it because they wanted the demonstration to work. The spike must move a genuinely awkward source.
Who has to be in the room, and for how long
Discoveries run over because client time was assumed rather than booked. Across three weeks:
| Role | Time | Needed for |
|---|---|---|
| Executive sponsor | 2–3 hours | Scope decision, arbitration when definitions conflict |
| ERP / applications owner | 6–10 hours | Table-level guidance, access approval, extract testing |
| IT security / infrastructure | 3–5 hours | Gateway install, firewall rules, service account approval |
| Operations lead | 6–8 hours | Metric definitions, workaround reports, decision context |
| Finance representative | 2–4 hours | Reconciliation rules, period close treatment |
| Power BI author / analyst | 8–12 hours | Report inventory, existing DAX logic, what is actually used |
That is 27–42 hours across six people. If the ERP owner is unavailable for two of the three weeks because of month-end close, the discovery does not compress, it slips. Check the operational calendar before setting the start date — the most common cause of overrun I see, and it sits entirely with the client.
Which discovery deliverables are worth paying for?
Four deliverables justify the fee; three common ones are filler:
| Deliverable | Why it matters / warning sign |
|---|---|
| Profiled data assessment | Names actual defects with counts. Warning: adjectives instead of numbers ("data quality is variable") |
| First-slice scope with fixed price | Converts discovery into a commitment. Warning: an "indicative" ±50% figure |
| Capacity recommendation with reasoning | The F64 boundary and throttling model drive real cost. Warning: a SKU named with no arithmetic |
| Architecture decision record | Records what was decided and rejected. Warning: diagrams only — a picture is not a decision |
| Metric definition register | One written definition per KPI with an owner's sign-off. Warning: names without formulas |
| Maturity assessment | Filler. A 1–5 score against a generic model says nothing about your goods receipt data |
| Generic roadmap | Filler. Three horizons, twelve initiatives, no costings ("Phase 3: Advanced Analytics") |
| Unrequested tooling comparison | Filler. If Fabric is chosen, a Fabric-versus-Snowflake section is padding |
What should a discovery cost relative to the build?
A proportionate discovery costs roughly 8–15% of the first-phase build it scopes. Below that range the assessment is usually too thin to support a fixed price. Above 20%, discovery has become the project, and the buyer is funding analysis the build will redo. The useful framing is not the number but what the proportion buys: discovery should narrow the uncertainty band on the build estimate from something like ±50% to ±15%. If it does not narrow the estimate, it was not discovery, whatever it cost.
Two commercial notes. A partner willing to credit part of the discovery fee against the build has something at stake. And insist discovery is separately terminable — bundled inside a twelve-month contract, the recommendation to proceed was made before any data was profiled.
Questions to ask a partner before commissioning discovery
- Which of your people will be on this, by name, and are they the people who will build it?
- What access do you need, and can you send that list before we sign?
- Will you profile full tables or a sample? Which tool, run by whom?
- What will the spike move, from which source to which output?
- Will you quote a fixed price for the first slice at the end? If not, why not?
- What is your capacity sizing method, and will you show the arithmetic?
- How many hours of our people's time do you need, by role?
- What have you found in previous discoveries that stopped a project?
- Is discovery separately terminable, and is any fee credited against the build?
- If the answer is "do not proceed yet", will you say so in writing?
That last question changes the conversation. Watch the pause.
What counts as success — including the outcome where you should not proceed
Three legitimate outcomes. Proceed on the scoped slice — the data supports it, definitions are agreed, the sponsor is real, and a fixed-price first slice is signed. Proceed on a different slice — the intended use case depends on a source with a defect that takes months to fix commercially, so changing target because the evidence changed is discovery working, not failing.
Or do not proceed yet — this happens and should be said plainly. The usual triggers: no named business owner for the decision the platform supports; an ERP migration already in flight that will invalidate the source model mid-build; an operating decision that cannot be made faster even with better information; or nobody to own the platform after handover. A discovery that concludes "not yet" and states what has to be true first has saved you the build budget.
Where this breaks: what this does not fix
Discovery cannot fix a source that has no data — if shift-level downtime was never captured, no assessment creates it, and instrumentation becomes a prerequisite project with a longer clock. A fixed price is only as good as its boundary — fixed price on a scope that says "integrate the ERP" is not fixed; it has to name tables, entities, metrics and report pages. Profiling shows today's data, not the governance behind it — a table that profiles clean today degrades the moment a new plant, distributor or pack code arrives.
The spike proves feasibility, not performance at volume — moving one month of one table proves the connection works, not that five years of history refresh inside the window. Discovery cannot manufacture a sponsor — if no operating leader will own the decision, the assessment will be accurate and the programme will still stall. And two to three weeks bounds the depth — nine source systems in three weeks is reconnaissance-depth profiling, not audit depth, so narrow the source list or lengthen the sprint, but be told which.
What to do first
Four questions to answer before commissioning:
- What decision is the platform meant to change, and who makes it today? Name the person and the meeting. If neither exists, that is the first finding
- Can we grant read access to our ERP within two weeks? If the answer involves a committee, start now, not on day one of the sprint
- Do plant, commercial and finance define our top three KPIs identically? Ask each separately, before they compare notes
- Is an ERP change, plant commissioning or acquisition in flight in the next six months? If so, the source model is moving and the sequence changes
At MyData Insights, discovery is the Discover phase of a Discover → Prototype → Deploy → Expand engagement, run under a Fractional Data Consultant model. First value in six weeks means a working slice against your real data — and it starts with a discovery honest enough to say when those six weeks should wait.
A real discovery touches your production data and ends with a fixed-price first slice the partner will stand behind — or an honest "not yet" with the conditions that would change it. Both are worth paying for; a roadmap of intentions is not. Book 30 minutes with Amit — no slides, no pitch deck, no obligation to proceed — a straight read on whether your source systems, access position and metric definitions are ready for a discovery.
Free Assessment
Where does your operation sit on the data maturity curve?
8 questions. 3 minutes. You get a scored breakdown across data infrastructure, analytics readiness, and automation potential — with a specific next step for your industry.