The bottom line
Lakehouse migrations fail for organisational reasons far more often than technical ones. The platform works — the data sits in open Delta on OneLake, so a bad pipeline is a rebuild, not a catastrophe. What kills the programme is the governance decisions available for free at the start: an unnamed decision the platform was meant to improve, a legacy stack never retired, lift-and-shift scoping that met a remodelling job, deferred metric definitions, absent sponsorship, a part-time owner, a capacity cost surprise, and no reconciliation deliverable at go-live. Execution failures are recoverable; failures of authority usually are not. The test: ask the sponsor to name the decision that changes, and the date it first changes.
In This Article
The programme rarely dies on a Friday
It dies over about nine months, in a sequence you can usually predict from week three. I sell lakehouse migrations, and in the estates I have worked in — manufacturing, FMCG, packaging, logistics and EPC — the platform almost always works.
That distinction is worth holding on to. Technical failure in a lakehouse migration is rare and almost always recoverable, because the data sits in an open format: OneLake stores tables in Delta Parquet, readable by any engine, so a bad pipeline is a rebuild, not a catastrophe. This names the failure modes in the order they actually unfold, with the early warning sign for each.
Why do lakehouse migrations fail?
For organisational reasons far more often than technical ones. The common causes are an unnamed decision the platform was meant to improve, a legacy stack that was never retired, lift-and-shift scoping that met a remodelling job, deferred metric definitions, absent sponsorship, a part-time internal owner, a capacity cost surprise, a demo-shaped build, data quality that became visible, and no reconciliation at go-live.
None of those are platform defects. Every one is a governance decision that was available for free at the start and becomes expensive in proportion to how long it is left. And the two failure types demand opposite responses — a technical failure wants more engineering; an organisational failure wants a decision, and more engineering makes it worse.
Technical failure versus organisational failure
Technical failure means a pipeline, model or capacity decision produced the wrong result or unacceptable performance. It is diagnosable, bounded and usually fixed in days to weeks — and it announces itself: a Direct Lake model falls back to DirectQuery because a table sits on a non-materialised SQL view, and the report is slow.
Organisational failure announces nothing. The planner still exports to Excel; the legacy report still circulates; the steering meeting is still "broadly positive". By the time it is undeniable, nine months have passed and the sponsor has spent their credibility. That silence is exactly what makes it dangerous.
The failure modes, in the order they unfold
Scope and decision failures appear in weeks 1–6; delivery and definition failures in weeks 6–16; cost and ownership failures around month two to four; acceptance failures at go-live. The eleven, in sequence:
- No decision was ever named — the platform is judged on whether people like the dashboards ("single source of truth" is not a decision)
- The legacy stack was never retired — two numbers circulate and trust splits; permanent parallel running is a slow-motion failure
- Lift-and-shift met a remodelling job — legacy views encode a decade of undocumented exclusions; copying them puts eleven definitions of net sales side by side
- Master data treated as a one-off cleanup — clean for a fortnight, then sales creates three new spellings because nothing changed at the point of entry
- Metric definitions deferred to "phase two" — the first release lands with an OTIF number operations disputes and a margin finance will not sign
- The sponsor left, or was never real — one who approves budget but will not settle a definitional dispute is not a sponsor
- The internal owner got the migration on top of a full-time job — two days a week assumed, never protected
- Capacity cost surprised finance in month two — backfill on top of business-as-usual, and nobody named to watch it
- The team optimised for the demo — a beautiful report on a hand-curated table that survives no schema change
- Data quality became visible and got blamed on the platform — three spellings of one supplier, harmless in silos, wrong once joined
- Go-live had no reconciliation deliverable — a 0.4% unexplained gross-margin variance keeps the legacy report authoritative indefinitely
Early warning signs and the intervention that still works
| Failure mode | Early warning sign | Intervention that still works |
|---|---|---|
| No decision named | Business case says "visibility"; no meeting will change | Stop build. Name one decision, one owner, one date. Re-scope around it |
| Legacy never retired | No decommission date, or it has moved twice | Put the shutdown date and its evidence test in the plan; hold a milestone against it |
| Lift-and-shift met remodelling | Engineers asking "why does this view exclude that?" weekly | Time-box a logic-discovery phase, priced separately; freeze the build until definitions land |
| Master data as one-off | Cleanup sprint, no owner for month three | Assign a data steward per domain with a weekly Purview data-quality scan |
| Definitions deferred | "We'll agree that in phase two" in the minutes | Pull the top ten metrics into scope now. A release without signed definitions is not a release |
| Sponsorship absent | Sponsor missed two steering meetings | Escalate once in writing with a decision list; if no response in two weeks, pause |
| Part-time owner | Workshops rescheduled twice for operations | Buy out their diary formally or replace them; do not proceed on assumed availability |
| Capacity surprise | Nobody can name the capacity owner | Name one; weekly review, surge thresholds, migration workload on separate capacity |
| Demo over estate | Showcase report has an unmentioned manual step | Move to dev/test/prod; rebuild the showcase on the real pipeline before the next demo |
| Quality became visible | A leader says "the platform's numbers are wrong" | Within a week, publish the defect list with owners and dates — silence loses the sponsor |
| No reconciliation at go-live | Cut-over date but no reconciliation pack | Delay go-live by one close; an unaccepted go-live is worse than a late one |
Which failures are recoverable, and which are not
Failures of execution are recoverable — scope, sequencing, capacity sizing, data quality and reconciliation can all be fixed with time and money. Recoverable, in my experience: a scoping error, a capacity surprise, a Silver-layer underestimation, master-data drift, poor reconciliation, a demo-shaped build. Each is expensive and each has a defined path out.
Usually terminal: no named decision after six months of delivery; a sponsor who has stopped attending and has not been replaced; two authoritative numbers circulating past the second close; and a business owner who has told you privately they cannot make the time. Those four are failures of authority, and authority cannot be bought with engineering. The honest test before you spend more is a single question asked of the sponsor, alone: name the decision that will be made differently because of this platform, and the date it will first be made that way. An answer means the programme is recoverable at almost any stage.
The test before spending more: ask the sponsor, alone, to name the decision this platform changes and the date it first changes. No answer means no amount of engineering will save it.
Where this breaks, and what it does not fix
A pause is not free or neutral — pausing a programme that has lost its sponsor is the right call and still burns political capital, sometimes permanently. Naming failure modes does not give an external partner authority to fix them — I can hold the mirror, escalate in writing and refuse to bill into a vacuum, but I cannot make a client decide. Some estates fail for reasons no diagnostic catches — an acquisition mid-programme, an ERP replacement announced in month four, a regulatory change that redirects finance for a quarter.
And none of this addresses history: if downtime reasons were entered as free text for six years, no medallion layer creates clean reason codes retrospectively. Fixed-scope contracting reduces commercial risk, not organisational risk — a fixed price protects your budget from scope drift, not from an absent sponsor.
What to do first
Four questions to answer this week, before the next invoice is approved:
- Name the decision — which recurring operating decision will be made differently, who makes it, and on what date does it first change?
- Name the shutdown date — when is the legacy warehouse switched off, and what evidence would finance accept, one reconciled close or three?
- Name the disputes — which three metric definitions are contested between departments, and who has authority to settle each?
- Name the owner — who owns capacity consumption, who owns master data per domain, and how many protected hours per week does the business owner actually have?
If the first and third have no answer, stop building — those are the cheapest problems on this list today and the most expensive in month six. We run lakehouse migrations on a fixed-scope basis after a short paid discovery, and the discovery exists mainly to answer those two before anyone commits to a build.
The cheapest problems on this list today are the most expensive in month six: a decision nobody named, and metric definitions nobody settled. Answer those before the next invoice. Book a diagnostic with Amit — no slides, no pitch deck, no obligation to proceed. If your programme has stalled, 30 minutes will tell you honestly whether it is a failure of execution or of authority.
Free Assessment
Where does your operation sit on the data maturity curve?
8 questions. 3 minutes. You get a scored breakdown across data infrastructure, analytics readiness, and automation potential — with a specific next step for your industry.