The bottom line
A data lakehouse programme has seven cost lines: platform and compute, storage, implementation, internal business time, parallel running during migration, ongoing run cost, and periodic remodelling and reprocessing. Implementation dominates year one; internal time is the most underestimated; storage is almost always negligible for structured operations data. Cost tracks estate complexity, not company size — two manufacturers with identical revenue can sit a factor of three apart on effort, driven by source count, whether sources are accessible or gated, transaction volume, entity and plant count, and regulatory scope. A single gated source (SAP via BW, an MES with no export, a SCADA historian) can exceed the effort of three accessible ones. Any vendor who gives a single figure before asking about your source systems is quoting for a scope they invented. Price the estate, not the company.
In This Article
The same three words described two very different programmes
The cheapest bid was pricing two source systems, one fact table and a Power BI report on top. The most expensive was pricing every plant, every legal entity, a governed semantic layer, historical reprocessing and two years of support. The same three words described both.
Any vendor who gives you a single figure before asking about your source systems is quoting for a scope they have invented. The figure will be defensible — for the scope in their head. It will not be defensible for yours. What follows is the cost structure, the variables that move each line, and a method to price your own estate.
Why can't anyone give you a straight price?
A data lakehouse has no standard scope. The term covers a two-source first slice delivered in six weeks and a group-wide platform spanning multiple ERPs, plants and legal entities. Cost is driven almost entirely by implementation effort, and effort is driven by source count, source accessibility, data quality and the number of contested metric definitions — none of which a vendor knows before discovery.
The second reason is structural: in a lakehouse programme the software line is small and the labour line is large, which inverts the buying instinct of anyone who has bought ERP. With SAP or Business Central, licence cost anchors the negotiation. Here there is no anchor — the platform can be provisioned in an afternoon; the twelve weeks that follow are where the money goes. The third reason is that a large share of the true cost never appears on any supplier invoice: your people's time, the parallel running period, and the year-two remodelling when the business reorganises are all real costs carried internally.
What are the cost lines in a data lakehouse programme?
Seven: platform and compute, storage, implementation, internal business time, parallel running during migration, ongoing run cost, and periodic costs. Two of the surprising ones first. Platform and compute is usually the smallest of the major lines in year one, not the largest — your first slice runs a handful of pipelines against a few million rows overnight, which is not a demanding workload. The genuine complication is that F64 is the threshold at which Power BI Free licences can view content, so a business with 300 report-readers often buys F64 for that licensing rule and receives the compute as a by-product.
Storage is negligible for structured data. Delta Parquet on a decade of ERP, MES and WMS records for a mid-size manufacturer routinely lands in the low terabytes or less, and OneLake storage is billed separately from capacity at a pay-as-you-go rate per GB. Two caveats: soft-deleted data bills at the same rate as active data, so retention policy has a cost; and read and write transactions do consume capacity units, so chatty small-file access is charged even though the bytes are cheap.
Implementation — the dominant line, and what actually moves it
This is where the programme cost lives, and four variables move it far more than data volume does:
- Source count and access difficulty. Each system is a discrete integration with its own authentication and failure modes. A cloud ERP with a documented API is straightforward; a SAP S/4HANA extract through a BW layer, an on-premises Epicor behind a gateway, an MES with no supported export, or a SCADA historian needing an OPC-UA bridge — each can cost more than three easy sources combined. Ask of every system: accessible, or gated?
- Data quality. Duplicate customer masters, three plants coding the same material differently, missing timestamps on production events. Cleansing is invisible in a proposal and unavoidable in delivery.
- Contested metric definitions. The line vendors never price. If Sales, Finance and Operations define OTIF differently, the programme absorbs the cost of resolving it. Count the metrics where two people would give different numbers today — that count is a cost driver.
- Undocumented business logic. The rebate calculation, the freight allocation, the scrap classification living in one analyst's spreadsheet. Reconstructing it depends on that person's availability, not the engineering.
A single gated source — SAP via BW, an MES with no export, a SCADA historian needing an OPC-UA bridge — can exceed the effort of three accessible ones. The gated-source count, not the total, predicts the effort.
Internal time — the line nobody budgets
Your people are not spectators. Discovery needs subject-matter time from ERP owners, plant leads and finance. Definition workshops need decision-makers, not delegates. UAT needs the people who will actually use the reports. Estimate it honestly: for each source system, name the one person who understands it, and ask how many hours a week they can genuinely give for the duration. For each contested metric, name the person who can rule on the definition and get their diary commitment in writing. Then apply the multiplier every operations leader knows — the named person is usually also running month-end.
Parallel running is a real line with a defined end: during migration you run old reporting and the new platform side by side until the numbers agree. Budget it as a defined period with exit criteria — the new number and old number agree to an agreed tolerance for a full accounting period, and the owner signs. Without stated exit criteria this line runs indefinitely, and it is the single most common source of budget overrun I see.
Ongoing run cost and periodic costs
Year two is not free. It contains the platform subscription, pipeline monitoring and failure response, schema-change maintenance when the ERP is upgraded, semantic model changes as the business asks new questions, and governance administration in Purview. Above all it contains a person who owns it — internal engineer, shared responsibility, or managed service. Estates without a named owner do not save that cost; they defer it and pay it as a rebuild.
Two periodic costs recur and neither appears in a year-one business case. Remodelling when the business changes — an acquisition, a new plant, an ERP upgrade or a reorganised chart of accounts forces changes to the dimensional model. Reprocessing history — when a definition changes or an upstream bug is found, restating several years is compute-heavy, careful work. Budget a modest annual allowance for both. A platform that cannot absorb a change without a project is a platform that will be worked around.
Sizing by estate complexity, not company size
For a mid-size manufacturer, cost tracks estate complexity rather than revenue. Two manufacturers with the same revenue can sit a factor of three apart on implementation effort. Size the estate, not the company.
| Complexity factor | Low-effort shape | High-effort shape |
|---|---|---|
| Source system count | 2–4 for the first slice | 8+, several needed at once |
| Source accessibility | Cloud ERP with documented API | Gated: SAP via BW, on-prem Epicor, MES with no export, SCADA via OPC-UA |
| Transaction volume | Millions of rows/year | Hundreds of millions; high-frequency sensor data |
| Legal entities and plants | One entity, one plant, one CoA | Multiple entities, several plants, inconsistent master data |
| Regulatory scope | Internal management reporting only | Statutory, ESG, e-invoicing or audited reporting |
Judge your own estate against that table before reading any proposal. If a bid does not distinguish between an accessible source and a gated one, it has not been costed — it has been guessed. The commercial wrapper matters too: fixed-scope pricing transfers estimation risk to the supplier and forces discovery to be real; time and materials keeps flexibility and keeps the risk with you.
The false economies
The cheapest bid — a bid materially below the others is usually pricing a narrower scope, a less experienced team, or an assumption that your sources are accessible. Ask what it excludes; if the answer is vague, it excludes the hard parts. Skipping discovery — the cheapest phase and the one most often cut to make a number work; skipping it does not remove the discovery, it moves it into build at build rates, where each finding is a change request. Skipping parallel running — cutting the overlap saves a few weeks and costs you the business's trust in the numbers, and a lakehouse nobody believes is a total loss.
Buying capacity too small — undersized capacity produces throttling, which the business experiences as slow or failed reports at exactly the moment adoption is being decided; capacity is one of the few lines you can adjust in weeks, but adoption, once lost, takes a year to recover. And buying the platform before defining the problem — the most expensive false economy of all.
Where this breaks: what a cost structure does not tell you
A cost structure prices the build. It does not price the organisational change around it, cannot predict source-system surprises, and will not tell you whether the programme returns anything. No estimate survives the first gated source — until someone has authenticated against the MES or historian and pulled a sample, the effort for that source is a guess, so sequence the riskiest extraction first, in week one. Cost structure is not a business case — knowing what the programme costs tells you nothing about what it returns; the return comes from decisions changing, and if nobody can name the decision that changes, the answer to "how much" is "too much", at any price.
The model assumes the data is worth landing — a cheap first slice built without governance, ownership and definition discipline produces a data swamp at low cost, which is not a saving. And internal time estimates are the least reliable part — people are optimistic about their own availability, particularly at quarter-end, so assume the real figure is meaningfully higher than the one your experts volunteer.
What to do first
A one-week exercise that produces a defensible budget bracket:
- List every source system, and mark each accessible or gated. Accessible means someone in your building can pull a sample this week. Count the gated ones separately — that number, not the total, predicts the effort
- Count the contested metrics. Ask three people to define OTIF, OEE and scrap rate independently. Every disagreement is a workshop, a decision and a rebuild of one measure
- Name the person behind each source and each definition, and get their weekly hours in writing. If a name cannot be produced, note it as a risk, not an assumption
- Decide the first decision, not the first dataset. One question, owned by one person, that would change what they do this month. Scope the first slice to that and nothing more
- Ask every bidder to price against this list. Any proposal that does not vary its number when you change the gated-source count has not been costed
Do those five and the order-of-magnitude spread in your quotes collapses to something you can defend to a board. We scope these programmes as a first slice with a fixed shape — and we will tell you when the honest answer is that your estate needs three months of source access work before anyone should be quoting a platform.
The order-of-magnitude spread in lakehouse quotes is not dishonesty — it is undefined scope. Mark your gated sources, count your contested metrics, name the first decision, and the number becomes defensible. Book 30 minutes with Amit — no slides, no pitch deck, no obligation to proceed — a straight read on which of your sources are gated and what that means for the number you take to your board.
Free Assessment
Where does your operation sit on the data maturity curve?
8 questions. 3 minutes. You get a scored breakdown across data infrastructure, analytics readiness, and automation potential — with a specific next step for your industry.