Skip to main content
Data Platform

What a Good Enterprise Data Solution Architecture Looks Like

A good architecture is not a diagram of components. It is a chain of decisions — requirements, then design, then trade-offs — where every box can be defended by why it is there and what the alternative would have cost.

Amit Kumar Singh - Technology Consulting Partner at MyData Insights

Technology Consulting Partner · MyData Insights

14+ years in industrial data · Former Accenture & EY · India, GCC, SEA

29 September 2026 · 11 min read

The bottom line

A good enterprise data architecture starts from business requirements, not components. The reference shape is a layered lakehouse: sources → ingestion (batch, CDC or streaming per source and latency) → Bronze (raw) → Silver (cleansed, conformed) → Gold (business-ready), with a governed semantic layer feeding BI and AI/agent consumers, and a governance plane (Unity Catalog on Databricks, or OneLake plus Microsoft Purview on Fabric) spanning everything for security, lineage and audit. What makes it good is not the boxes but the decisions behind them: batch vs streaming, managed vs external tables, compute isolation, centralised vs domain governance, cost, reliability/DR and security — each chosen for a stated reason with the trade-off named. Design requirements → architecture → components → security → performance → cost → reliability → trade-offs.

Architecture Is Decisions, Not Boxes

Ask a data engineer for an architecture and you often get a list of components: a lakehouse, some pipelines, a BI tool. Ask a solutions architect and you get a chain of decisions: here are the requirements, here is the design that meets them, here is why these components and not the alternatives, and here are the security, scalability, reliability and cost trade-offs I accepted.

That is the difference between a diagram and an architecture. A good enterprise data architecture can defend every box — why it is there, what belongs in it, who can access it, and what choosing it costs versus the alternative. The boxes are often unsurprising; the quality is in the reasoning.

So a good architecture starts not with Databricks or Fabric but with the business: what decisions must the data support, at what latency, for whom, under what security and compliance constraints, at what cost. Only then does it map those requirements to components.

A diagram lists components; an architecture defends decisions. Every box should have a reason — why it is there, what belongs in it, who can access it, and what the alternative would have cost.

The Layered Lakehouse

The reference shape for an enterprise estate is a layered lakehouse. Sources — ERP, CRM, MES, APIs, files — feed an ingestion layer, where the method is chosen per source and latency need: batch for periodic loads, change data capture for transactional databases that need to stay current, streaming where near-real-time matters, and file auto-loading for landing zones.

Data then flows through medallion layers. Bronze holds raw data exactly as ingested. Silver holds cleansed, deduplicated, conformed data with business logic applied. Gold holds business-ready, aggregated, modelled data — the layer BI and AI actually consume. Delta Lake underpins all three, giving ACID transactions, schema enforcement and evolution, time travel, and MERGE for upserts and CDC.

The architect’s questions about each layer are not "what is Bronze?" but "why this layer, what belongs there, who can access it, and what are the performance, governance and cost implications?" A CDC-fed Silver layer with MERGE, checkpointing and idempotency for exactly-once processing is a design decision with real trade-offs, not a default.

Sources → ingestion (batch/CDC/streaming per need) → Bronze (raw) → Silver (conformed) → Gold (business-ready), on Delta Lake. The SA question is never "what is Bronze?" but "why this layer, who accesses it, and at what cost?"

Governance and the Semantic Layer

Two things span the whole architecture. The first is a governance plane: on Databricks that is Unity Catalog; on Microsoft Fabric it is OneLake with Microsoft Purview and the OneLake catalog. It provides one permission model, row- and column-level security, lineage and audit across every workspace and layer. Governance is not a box at the side of the diagram — it wraps everything, because access and traceability apply at every layer.

The second is a governed semantic layer on top of Gold: one agreed definition of each measure — revenue, OEE, OTIF, margin. This is what makes the Gold layer business-ready rather than just aggregated, and it is the difference between reports that agree and reports that contradict. It is also the prerequisite for trustworthy AI: an agent or a natural-language tool reasons over the semantic layer, so ambiguous definitions there produce confidently wrong answers everywhere.

A good architecture treats both as first-class, designed in from the start — not bolted on after the pipelines work. Governance and semantics are what turn a data platform into an intelligence platform.

Consumers: BI and AI

The Gold and semantic layer feeds two kinds of consumer. Traditional BI — Power BI dashboards and reports for analysts and executives — and AI/agent consumers: natural-language analytics (Power BI Copilot, a Fabric Data Agent, or Databricks Genie) and operational AI agents that reason and act. Both read from the same governed, business-defined layer, which is why the semantic work pays off twice.

Take a common requirement: CRM, ERP and operational data, Power BI dashboards, and 5,000 users asking questions in natural language. The architecture routes both Power BI and the natural-language interface to the Gold/semantic layer, over the medallion stack, with the governance plane enforcing who sees what. The users get dashboards and plain-English answers from one governed source, not two parallel stacks.

The design discipline is to serve every consumer from the same governed foundation, not to build a separate pipeline per tool. One Gold layer, one semantic definition, many consumers — that is what keeps the numbers consistent across BI, natural language and agents.

One Gold and semantic layer, many consumers — BI, natural-language analytics and AI agents all read the same governed definitions. Serve every consumer from one foundation, not a separate pipeline per tool.

The Trade-offs That Matter

What makes an architecture defensible is the trade-offs named alongside it. Batch versus streaming: streaming meets a five-minute latency need but costs more and adds complexity; nightly batch is simpler and cheaper where the business can wait. Managed versus external tables: Databricks-owned lifecycle versus existing cloud storage kept under your control. Compute: separate ETL, BI and ML workloads so a heavy Spark job does not starve interactive queries — SQL warehouses for BI, job compute for ETL, isolation between them.

Governance: centralised versus domain-oriented catalogs — one metastore for consistency, or federated catalogs for domain autonomy. Cost: right-sized compute, autoscaling, auto-termination, reserved capacity for steady workloads, and monitoring so a doubled bill can be traced to a workload rather than guessed at. Reliability and DR: RPO/RTO, backups, and a recovery plan proportionate to how critical the platform is.

And security throughout: identity and SSO, groups and service principals, private networking, encryption and key management, PII protection, least privilege and audit. A good architecture does not present these as afterthoughts; it states the choice and the reason for each, because that is what a real design is.

So What — the Framework

Design in this order, every time: business requirement → data and workload requirement → architecture → component selection → security and governance → performance and scalability → cost → reliability and DR → trade-offs → why this design. Lead with requirements, close with the reasoning; the components sit in the middle, chosen to serve, not to impress.

So the answer to "design an enterprise data platform" is not "I’ll use a lakehouse, a catalog and a BI tool." It is: "the business needs near-real-time analytics for 5,000 users under these security constraints, so I would use CDC ingestion, a medallion lakehouse on Delta Lake, a governed semantic layer for BI and natural-language analytics, a governance plane for access and lineage, separated compute for ETL and BI, and least-privilege security — with the final shape depending on latency, workload, security and cost." That is architecture.

The reference shape is stable; the right instance depends on the estate. We build Microsoft-first — Fabric, OneLake, Power BI, Purview — and deliver on Databricks with Unity Catalog and Delta where a client’s estate runs there. The platform follows the requirements; the discipline of requirements → design → trade-offs does not change.

Requirement → data/workload → architecture → components → security → performance → cost → reliability → trade-offs → why. Lead with requirements, close with reasoning; the components sit in the middle, chosen to serve.

If you are designing or reviewing an enterprise data platform and want the decisions pressure-tested — not just a diagram — that is a conversation worth having. 30 minutes with Amit on your requirements, the architecture that fits, and the trade-offs behind each choice. No slides. No pitch deck. No obligation to proceed.

Free Assessment

Where does your operation sit on the data maturity curve?

8 questions. 3 minutes. You get a scored breakdown across data infrastructure, analytics readiness, and automation potential — with a specific next step for your industry.

Data PlatformArchitectureMicrosoft FabricDatabricksLakehouseData Governance

Your Data · Our Technology · Our Automation

Get practical insights every fortnight

Amit writes about Microsoft Fabric, Power BI, AI in operations, and digital transformation for manufacturing and supply chain leaders. Practitioner perspective - no fluff, no vendor spin.

No spam. Unsubscribe any time. Also on Substack.

FAQ

Common questions

What does a good enterprise data architecture look like?

A layered lakehouse driven by business requirements: sources → ingestion (batch, CDC or streaming per source and latency) → Bronze (raw) → Silver (cleansed, conformed) → Gold (business-ready) on Delta Lake, with a governed semantic layer feeding BI and AI/agent consumers, and a governance plane (Unity Catalog, or OneLake plus Microsoft Purview) spanning everything for security, lineage and audit. What makes it good is the decisions behind the boxes, each chosen for a stated reason with its trade-off named.

How does a solutions architect approach a data platform design?

By starting with requirements, not components. The framework is: business requirement → data and workload requirement → architecture → component selection → security and governance → performance and scalability → cost → reliability and DR → trade-offs → why this design. A data engineer asks "how do I build the pipeline?"; an architect asks "how should I design the platform, why these components, and what are the security, scalability, reliability and cost trade-offs?"

What are Bronze, Silver and Gold layers?

They are the medallion architecture. Bronze holds raw data exactly as ingested; Silver holds cleansed, deduplicated and conformed data with business logic applied; Gold holds business-ready, aggregated and modelled data that BI and AI consume. The architect’s question about each is not "what is it?" but "why this layer, what belongs there, who can access it, and what are the performance, governance and cost implications?"

Should BI and AI read from the same layer?

Yes — both should read from the same governed Gold and semantic layer, not separate pipelines. Power BI dashboards, natural-language analytics (Copilot, Fabric Data Agent, Databricks Genie) and operational AI agents all consuming one governed set of definitions is what keeps numbers consistent across tools. It also means the semantic work — one agreed definition of each measure — pays off across BI and AI at once.

Which platform is best for an enterprise data architecture?

It depends on the estate and requirements, not a default. The reference architecture — layered lakehouse, governed semantic layer, governance plane, BI and AI consumers — is the same on either stack. We build Microsoft-first (Microsoft Fabric, OneLake, Power BI, Purview) and deliver on Databricks with Unity Catalog and Delta Lake where a client’s estate already runs there. The platform follows the requirements; the discipline of requirements → design → trade-offs does not change.

Related FAQs

Questions operations leaders ask

Continue Reading

Related Articles

Data Platform

Microsoft Fabric vs a Legacy BI Stack (SSIS + SSAS + Power BI): The Migration Case

The most common estate I walk into is not a mess. It is an on-premises SQL Server, a set of SSIS packages built between 2014 and 2019, one or two SSAS cubes, and Power BI bolted on the front. It runs. Finance closes on it. The reason I get called is a symptom — the person who wrote the packages left, the overnight batch now finishes at 07:20 and the plant meeting is at 07:30. "It is old" is not a business case.

16 min read

Data Platform

Microsoft Fabric vs SAP Datasphere: Which One Do You Actually Need

The SAP account team says the analytics answer is SAP Datasphere, because that is where the business semantics already live. Two weeks later the Microsoft team says Fabric, because that is where Power BI, the MES extracts and the 3PL feeds already live. Both are internally consistent, and neither mentions the other except to dismiss it. The IT Head is asked to pick, and picks badly — because the two products solve different halves of one problem.

16 min read

Data Platform

The Hidden Costs of a Microsoft Fabric Migration Nobody Tells You About

The awkward conversation happens in month five, not month one. The platform works. The first three reports are live. Then the finance business partner circulates the actual run-rate against the approved business case, and the number is 30–50% over — not because the partner overran, but because six or seven cost lines were never in the case at all. I sell Fabric implementations. This names the costs my own proposals have to cover.

15 min read

Want to see how MDI solves this in your industry? Explore industry solutions

Is this the challenge you're facing?

Book a 30-minute call. We'll look at your specific operation and tell you what's achievable - plainly and without slides.