The bottom line
A good enterprise data architecture starts from business requirements, not components. The reference shape is a layered lakehouse: sources → ingestion (batch, CDC or streaming per source and latency) → Bronze (raw) → Silver (cleansed, conformed) → Gold (business-ready), with a governed semantic layer feeding BI and AI/agent consumers, and a governance plane (Unity Catalog on Databricks, or OneLake plus Microsoft Purview on Fabric) spanning everything for security, lineage and audit. What makes it good is not the boxes but the decisions behind them: batch vs streaming, managed vs external tables, compute isolation, centralised vs domain governance, cost, reliability/DR and security — each chosen for a stated reason with the trade-off named. Design requirements → architecture → components → security → performance → cost → reliability → trade-offs.
In This Article
Architecture Is Decisions, Not Boxes
Ask a data engineer for an architecture and you often get a list of components: a lakehouse, some pipelines, a BI tool. Ask a solutions architect and you get a chain of decisions: here are the requirements, here is the design that meets them, here is why these components and not the alternatives, and here are the security, scalability, reliability and cost trade-offs I accepted.
That is the difference between a diagram and an architecture. A good enterprise data architecture can defend every box — why it is there, what belongs in it, who can access it, and what choosing it costs versus the alternative. The boxes are often unsurprising; the quality is in the reasoning.
So a good architecture starts not with Databricks or Fabric but with the business: what decisions must the data support, at what latency, for whom, under what security and compliance constraints, at what cost. Only then does it map those requirements to components.
A diagram lists components; an architecture defends decisions. Every box should have a reason — why it is there, what belongs in it, who can access it, and what the alternative would have cost.
The Layered Lakehouse
The reference shape for an enterprise estate is a layered lakehouse. Sources — ERP, CRM, MES, APIs, files — feed an ingestion layer, where the method is chosen per source and latency need: batch for periodic loads, change data capture for transactional databases that need to stay current, streaming where near-real-time matters, and file auto-loading for landing zones.
Data then flows through medallion layers. Bronze holds raw data exactly as ingested. Silver holds cleansed, deduplicated, conformed data with business logic applied. Gold holds business-ready, aggregated, modelled data — the layer BI and AI actually consume. Delta Lake underpins all three, giving ACID transactions, schema enforcement and evolution, time travel, and MERGE for upserts and CDC.
The architect’s questions about each layer are not "what is Bronze?" but "why this layer, what belongs there, who can access it, and what are the performance, governance and cost implications?" A CDC-fed Silver layer with MERGE, checkpointing and idempotency for exactly-once processing is a design decision with real trade-offs, not a default.
Sources → ingestion (batch/CDC/streaming per need) → Bronze (raw) → Silver (conformed) → Gold (business-ready), on Delta Lake. The SA question is never "what is Bronze?" but "why this layer, who accesses it, and at what cost?"
Governance and the Semantic Layer
Two things span the whole architecture. The first is a governance plane: on Databricks that is Unity Catalog; on Microsoft Fabric it is OneLake with Microsoft Purview and the OneLake catalog. It provides one permission model, row- and column-level security, lineage and audit across every workspace and layer. Governance is not a box at the side of the diagram — it wraps everything, because access and traceability apply at every layer.
The second is a governed semantic layer on top of Gold: one agreed definition of each measure — revenue, OEE, OTIF, margin. This is what makes the Gold layer business-ready rather than just aggregated, and it is the difference between reports that agree and reports that contradict. It is also the prerequisite for trustworthy AI: an agent or a natural-language tool reasons over the semantic layer, so ambiguous definitions there produce confidently wrong answers everywhere.
A good architecture treats both as first-class, designed in from the start — not bolted on after the pipelines work. Governance and semantics are what turn a data platform into an intelligence platform.
Consumers: BI and AI
The Gold and semantic layer feeds two kinds of consumer. Traditional BI — Power BI dashboards and reports for analysts and executives — and AI/agent consumers: natural-language analytics (Power BI Copilot, a Fabric Data Agent, or Databricks Genie) and operational AI agents that reason and act. Both read from the same governed, business-defined layer, which is why the semantic work pays off twice.
Take a common requirement: CRM, ERP and operational data, Power BI dashboards, and 5,000 users asking questions in natural language. The architecture routes both Power BI and the natural-language interface to the Gold/semantic layer, over the medallion stack, with the governance plane enforcing who sees what. The users get dashboards and plain-English answers from one governed source, not two parallel stacks.
The design discipline is to serve every consumer from the same governed foundation, not to build a separate pipeline per tool. One Gold layer, one semantic definition, many consumers — that is what keeps the numbers consistent across BI, natural language and agents.
One Gold and semantic layer, many consumers — BI, natural-language analytics and AI agents all read the same governed definitions. Serve every consumer from one foundation, not a separate pipeline per tool.
The Trade-offs That Matter
What makes an architecture defensible is the trade-offs named alongside it. Batch versus streaming: streaming meets a five-minute latency need but costs more and adds complexity; nightly batch is simpler and cheaper where the business can wait. Managed versus external tables: Databricks-owned lifecycle versus existing cloud storage kept under your control. Compute: separate ETL, BI and ML workloads so a heavy Spark job does not starve interactive queries — SQL warehouses for BI, job compute for ETL, isolation between them.
Governance: centralised versus domain-oriented catalogs — one metastore for consistency, or federated catalogs for domain autonomy. Cost: right-sized compute, autoscaling, auto-termination, reserved capacity for steady workloads, and monitoring so a doubled bill can be traced to a workload rather than guessed at. Reliability and DR: RPO/RTO, backups, and a recovery plan proportionate to how critical the platform is.
And security throughout: identity and SSO, groups and service principals, private networking, encryption and key management, PII protection, least privilege and audit. A good architecture does not present these as afterthoughts; it states the choice and the reason for each, because that is what a real design is.
So What — the Framework
Design in this order, every time: business requirement → data and workload requirement → architecture → component selection → security and governance → performance and scalability → cost → reliability and DR → trade-offs → why this design. Lead with requirements, close with the reasoning; the components sit in the middle, chosen to serve, not to impress.
So the answer to "design an enterprise data platform" is not "I’ll use a lakehouse, a catalog and a BI tool." It is: "the business needs near-real-time analytics for 5,000 users under these security constraints, so I would use CDC ingestion, a medallion lakehouse on Delta Lake, a governed semantic layer for BI and natural-language analytics, a governance plane for access and lineage, separated compute for ETL and BI, and least-privilege security — with the final shape depending on latency, workload, security and cost." That is architecture.
The reference shape is stable; the right instance depends on the estate. We build Microsoft-first — Fabric, OneLake, Power BI, Purview — and deliver on Databricks with Unity Catalog and Delta where a client’s estate runs there. The platform follows the requirements; the discipline of requirements → design → trade-offs does not change.
Requirement → data/workload → architecture → components → security → performance → cost → reliability → trade-offs → why. Lead with requirements, close with reasoning; the components sit in the middle, chosen to serve.
If you are designing or reviewing an enterprise data platform and want the decisions pressure-tested — not just a diagram — that is a conversation worth having. 30 minutes with Amit on your requirements, the architecture that fits, and the trade-offs behind each choice. No slides. No pitch deck. No obligation to proceed.
Free Assessment
Where does your operation sit on the data maturity curve?
8 questions. 3 minutes. You get a scored breakdown across data infrastructure, analytics readiness, and automation potential — with a specific next step for your industry.