Skip to main content
Data Platform

Common Reasons a Lakehouse Migration Fails (and Which Are Recoverable)

The programme rarely dies on a Friday. It dies over about nine months, in a sequence you can predict from week three. The platform almost always works — the failure is organisational. Eleven failure modes in the order they unfold, the early warning for each, and which are recoverable.

Amit Kumar Singh - Technology Consulting Partner at MyData Insights

Technology Consulting Partner · MyData Insights

14+ years in industrial data · Former Accenture & EY · India, GCC, SEA

19 August 2026 · 13 min read

The bottom line

Lakehouse migrations fail for organisational reasons far more often than technical ones. The platform works — the data sits in open Delta on OneLake, so a bad pipeline is a rebuild, not a catastrophe. What kills the programme is the governance decisions available for free at the start: an unnamed decision the platform was meant to improve, a legacy stack never retired, lift-and-shift scoping that met a remodelling job, deferred metric definitions, absent sponsorship, a part-time owner, a capacity cost surprise, and no reconciliation deliverable at go-live. Execution failures are recoverable; failures of authority usually are not. The test: ask the sponsor to name the decision that changes, and the date it first changes.

The programme rarely dies on a Friday

It dies over about nine months, in a sequence you can usually predict from week three. I sell lakehouse migrations, and in the estates I have worked in — manufacturing, FMCG, packaging, logistics and EPC — the platform almost always works.

That distinction is worth holding on to. Technical failure in a lakehouse migration is rare and almost always recoverable, because the data sits in an open format: OneLake stores tables in Delta Parquet, readable by any engine, so a bad pipeline is a rebuild, not a catastrophe. This names the failure modes in the order they actually unfold, with the early warning sign for each.

Why do lakehouse migrations fail?

For organisational reasons far more often than technical ones. The common causes are an unnamed decision the platform was meant to improve, a legacy stack that was never retired, lift-and-shift scoping that met a remodelling job, deferred metric definitions, absent sponsorship, a part-time internal owner, a capacity cost surprise, a demo-shaped build, data quality that became visible, and no reconciliation at go-live.

None of those are platform defects. Every one is a governance decision that was available for free at the start and becomes expensive in proportion to how long it is left. And the two failure types demand opposite responses — a technical failure wants more engineering; an organisational failure wants a decision, and more engineering makes it worse.

Technical failure versus organisational failure

Technical failure means a pipeline, model or capacity decision produced the wrong result or unacceptable performance. It is diagnosable, bounded and usually fixed in days to weeks — and it announces itself: a Direct Lake model falls back to DirectQuery because a table sits on a non-materialised SQL view, and the report is slow.

Organisational failure announces nothing. The planner still exports to Excel; the legacy report still circulates; the steering meeting is still "broadly positive". By the time it is undeniable, nine months have passed and the sponsor has spent their credibility. That silence is exactly what makes it dangerous.

The failure modes, in the order they unfold

Scope and decision failures appear in weeks 1–6; delivery and definition failures in weeks 6–16; cost and ownership failures around month two to four; acceptance failures at go-live. The eleven, in sequence:

  • No decision was ever named — the platform is judged on whether people like the dashboards ("single source of truth" is not a decision)
  • The legacy stack was never retired — two numbers circulate and trust splits; permanent parallel running is a slow-motion failure
  • Lift-and-shift met a remodelling job — legacy views encode a decade of undocumented exclusions; copying them puts eleven definitions of net sales side by side
  • Master data treated as a one-off cleanup — clean for a fortnight, then sales creates three new spellings because nothing changed at the point of entry
  • Metric definitions deferred to "phase two" — the first release lands with an OTIF number operations disputes and a margin finance will not sign
  • The sponsor left, or was never real — one who approves budget but will not settle a definitional dispute is not a sponsor
  • The internal owner got the migration on top of a full-time job — two days a week assumed, never protected
  • Capacity cost surprised finance in month two — backfill on top of business-as-usual, and nobody named to watch it
  • The team optimised for the demo — a beautiful report on a hand-curated table that survives no schema change
  • Data quality became visible and got blamed on the platform — three spellings of one supplier, harmless in silos, wrong once joined
  • Go-live had no reconciliation deliverable — a 0.4% unexplained gross-margin variance keeps the legacy report authoritative indefinitely

Early warning signs and the intervention that still works

Failure modeEarly warning signIntervention that still works
No decision namedBusiness case says "visibility"; no meeting will changeStop build. Name one decision, one owner, one date. Re-scope around it
Legacy never retiredNo decommission date, or it has moved twicePut the shutdown date and its evidence test in the plan; hold a milestone against it
Lift-and-shift met remodellingEngineers asking "why does this view exclude that?" weeklyTime-box a logic-discovery phase, priced separately; freeze the build until definitions land
Master data as one-offCleanup sprint, no owner for month threeAssign a data steward per domain with a weekly Purview data-quality scan
Definitions deferred"We'll agree that in phase two" in the minutesPull the top ten metrics into scope now. A release without signed definitions is not a release
Sponsorship absentSponsor missed two steering meetingsEscalate once in writing with a decision list; if no response in two weeks, pause
Part-time ownerWorkshops rescheduled twice for operationsBuy out their diary formally or replace them; do not proceed on assumed availability
Capacity surpriseNobody can name the capacity ownerName one; weekly review, surge thresholds, migration workload on separate capacity
Demo over estateShowcase report has an unmentioned manual stepMove to dev/test/prod; rebuild the showcase on the real pipeline before the next demo
Quality became visibleA leader says "the platform's numbers are wrong"Within a week, publish the defect list with owners and dates — silence loses the sponsor
No reconciliation at go-liveCut-over date but no reconciliation packDelay go-live by one close; an unaccepted go-live is worse than a late one

Which failures are recoverable, and which are not

Failures of execution are recoverable — scope, sequencing, capacity sizing, data quality and reconciliation can all be fixed with time and money. Recoverable, in my experience: a scoping error, a capacity surprise, a Silver-layer underestimation, master-data drift, poor reconciliation, a demo-shaped build. Each is expensive and each has a defined path out.

Usually terminal: no named decision after six months of delivery; a sponsor who has stopped attending and has not been replaced; two authoritative numbers circulating past the second close; and a business owner who has told you privately they cannot make the time. Those four are failures of authority, and authority cannot be bought with engineering. The honest test before you spend more is a single question asked of the sponsor, alone: name the decision that will be made differently because of this platform, and the date it will first be made that way. An answer means the programme is recoverable at almost any stage.

The test before spending more: ask the sponsor, alone, to name the decision this platform changes and the date it first changes. No answer means no amount of engineering will save it.

Where this breaks, and what it does not fix

A pause is not free or neutral — pausing a programme that has lost its sponsor is the right call and still burns political capital, sometimes permanently. Naming failure modes does not give an external partner authority to fix them — I can hold the mirror, escalate in writing and refuse to bill into a vacuum, but I cannot make a client decide. Some estates fail for reasons no diagnostic catches — an acquisition mid-programme, an ERP replacement announced in month four, a regulatory change that redirects finance for a quarter.

And none of this addresses history: if downtime reasons were entered as free text for six years, no medallion layer creates clean reason codes retrospectively. Fixed-scope contracting reduces commercial risk, not organisational risk — a fixed price protects your budget from scope drift, not from an absent sponsor.

What to do first

Four questions to answer this week, before the next invoice is approved:

  • Name the decision — which recurring operating decision will be made differently, who makes it, and on what date does it first change?
  • Name the shutdown date — when is the legacy warehouse switched off, and what evidence would finance accept, one reconciled close or three?
  • Name the disputes — which three metric definitions are contested between departments, and who has authority to settle each?
  • Name the owner — who owns capacity consumption, who owns master data per domain, and how many protected hours per week does the business owner actually have?

If the first and third have no answer, stop building — those are the cheapest problems on this list today and the most expensive in month six. We run lakehouse migrations on a fixed-scope basis after a short paid discovery, and the discovery exists mainly to answer those two before anyone commits to a build.

The cheapest problems on this list today are the most expensive in month six: a decision nobody named, and metric definitions nobody settled. Answer those before the next invoice. Book a diagnostic with Amit — no slides, no pitch deck, no obligation to proceed. If your programme has stalled, 30 minutes will tell you honestly whether it is a failure of execution or of authority.

Free Assessment

Where does your operation sit on the data maturity curve?

8 questions. 3 minutes. You get a scored breakdown across data infrastructure, analytics readiness, and automation potential — with a specific next step for your industry.

Data PlatformLakehouseMigrationGovernanceMicrosoft Fabric

Your Data · Our Technology · Our Automation

Get practical insights every fortnight

Amit writes about Microsoft Fabric, Power BI, AI in operations, and digital transformation for manufacturing and supply chain leaders. Practitioner perspective - no fluff, no vendor spin.

No spam. Unsubscribe any time. Also on Substack.

FAQ

Common questions

Why do most data lakehouse migrations fail?

Most fail organisationally rather than technically. The platform works, but no operating decision was named that it was meant to improve, so it gets judged on whether people like the dashboards — and an organisational failure is made worse, not better, by more engineering.

Is a failed lakehouse migration recoverable?

Execution failures are — scoping errors, capacity surprises, Silver-layer underestimation, weak reconciliation. Authority failures usually are not: no named decision after six months, an absent unreplaced sponsor, two authoritative numbers past the second close, or an owner who privately cannot make the time.

How long does it take to notice a lakehouse migration is failing?

The signs appear in weeks 3–6 and are usually acted on in month nine. Early indicators are workshops rescheduled twice, "we'll agree that in phase two" in the minutes, no decommission date in the plan, and a sponsor missing two consecutive steering meetings.

What is the most common technical cause of a lakehouse migration overrun?

Scoping the Silver layer as a lift-and-shift when it is a remodelling job. Legacy views encode undocumented exclusions and overrides that must be rediscovered and agreed before rebuilding.

Should we run the old warehouse and the lakehouse in parallel?

Yes, for one to two monthly closes — long enough for finance to reconcile both. Permanent parallel running is a failure mode, not a safety net: it lets two numbers circulate and splits trust.

Who should own a lakehouse migration internally?

Someone with operating authority over the decisions the platform is meant to improve, with protected time — typically an operations, supply chain or finance leader, not an IT project manager alone. Ownership on top of a full-time job is one of the most reliable predictors of a stalled programme.

Continue Reading

Related Articles

Data Platform

Microsoft Fabric vs a Legacy BI Stack (SSIS + SSAS + Power BI): The Migration Case

The most common estate I walk into is not a mess. It is an on-premises SQL Server, a set of SSIS packages built between 2014 and 2019, one or two SSAS cubes, and Power BI bolted on the front. It runs. Finance closes on it. The reason I get called is a symptom — the person who wrote the packages left, the overnight batch now finishes at 07:20 and the plant meeting is at 07:30. "It is old" is not a business case.

16 min read

Data Platform

Microsoft Fabric vs SAP Datasphere: Which One Do You Actually Need

The SAP account team says the analytics answer is SAP Datasphere, because that is where the business semantics already live. Two weeks later the Microsoft team says Fabric, because that is where Power BI, the MES extracts and the 3PL feeds already live. Both are internally consistent, and neither mentions the other except to dismiss it. The IT Head is asked to pick, and picks badly — because the two products solve different halves of one problem.

16 min read

Data Platform

The Hidden Costs of a Microsoft Fabric Migration Nobody Tells You About

The awkward conversation happens in month five, not month one. The platform works. The first three reports are live. Then the finance business partner circulates the actual run-rate against the approved business case, and the number is 30–50% over — not because the partner overran, but because six or seven cost lines were never in the case at all. I sell Fabric implementations. This names the costs my own proposals have to cover.

15 min read

Want to see how MDI solves this in your industry? Explore industry solutions

Is this the challenge you're facing?

Book a 30-minute call. We'll look at your specific operation and tell you what's achievable - plainly and without slides.