Skip to main content
Data Architecture

The Enterprise Microsoft Fabric + AI Architecture: From OneLake to Fabric Data Agent and MCP

The difference between an AI agent that helps and one that answers confidently and wrongly is entirely in the architecture beneath it. Here is the full stack — OneLake to Fabric IQ to the Data Agent to MCP — and why you never wire an external model straight to your database.

Amit Kumar Singh - Technology Consulting Partner at MyData Insights

Technology Consulting Partner · MyData Insights

14+ years in industrial data · Former Accenture & EY · India, GCC, SEA

30 September 2026 · 13 min read

The bottom line

A trustworthy AI-over-your-data architecture on Microsoft Fabric is a stack, not a shortcut. Sources feed OneLake and a medallion lakehouse (Bronze→Silver→Gold) on Delta Lake; a governed semantic model defines each measure once; a Fabric IQ ontology models the business entities and relationships; a Fabric Data Agent lets users ask questions in plain English over that governed context; and MCP (Model Context Protocol) exposes governed capabilities to external agents like Claude through an authenticated, authorised gateway — never direct database access. Microsoft Purview and the OneLake catalog govern every layer; Entra-to-RLS/OLS security wraps it; Git and DEV→QA→UAT→PROD CI/CD deploys it. The reason agents hallucinate or leak is almost always a missing layer, not a weak model — you fix it by building the stack, not by prompting harder.

A Stack, Not a Shortcut

When an AI agent over company data gives a confident, wrong answer, the instinct is to blame the model or rewrite the prompt. It is almost always neither. It is a missing layer in the architecture beneath the agent — no governed Gold layer, no semantic definitions, no ontology, or no access control — and no amount of prompting fixes a foundation problem.

A trustworthy AI-over-your-data architecture on Microsoft Fabric is a stack, and every layer earns its place. From the bottom: OneLake and a medallion lakehouse; a governed semantic model; a Fabric IQ ontology; a Fabric Data Agent; and, for external agents, an MCP gateway. Governance and security wrap all of it, and CI/CD deploys it.

The contrast that matters most is at the top. The wrong architecture is Claude — or any external model — wired straight to a database with credentials and free rein to run SQL. The right architecture routes it through a governed interface that decides what it can do. Everything below exists to make that top layer safe.

Agents do not hallucinate because the model is weak — they hallucinate because a layer is missing. Fix the stack, not the prompt.

OneLake and the Medallion Foundation

It starts, as always, with the data foundation. Sources — SAP, Dynamics 365, MES, IoT, SQL Server, APIs — feed OneLake through the right ingestion method per source: batch, change data capture, streaming or file auto-loading. Data flows through the medallion layers on Delta Lake: Bronze holds it raw, Silver cleanses and conforms it, Gold makes it business-ready.

This is the unglamorous layer that decides everything above it. An agent reasoning over an un-curated estate is reasoning over chaos, and it will answer accordingly. The Gold layer — modelled, deduplicated, business-ready Delta tables — is the first thing that makes AI answers trustworthy, and there is no skipping it.

Nothing about AI changes this order. Unify the data first; the intelligence layers come second. Trying to put an agent on fragmented data produces a confident interface to unreliable answers — worse than no agent, because people believe it.

Semantic Model and Ontology

On top of Gold sit two layers of meaning. The semantic model defines each measure once — revenue, OEE, OTIF, margin — so that every consumer, human or agent, uses the same definition. This is what stops two people getting two answers to the same question, and it is the difference between a semantic model and a pile of tables.

Above that, a Fabric IQ ontology models the business itself: the entities (order, customer, plant, shipment), the relationships between them, and the measures that describe them. Where a semantic model is optimised for reporting over a dataset, the ontology is a broader, cross-domain model designed to be reasoned over by agents and operational processes as well as analytics.

This is the layer that turns a data platform into something an agent can reason across. Point an agent at raw tables and it guesses how the business fits together, inconsistently; give it an ontology and it reasons over your actual business meaning. The ontology is not a reporting upgrade — it is the shared context that makes agentic operations trustworthy.

The semantic model defines each measure once; the Fabric IQ ontology models the business entities and relationships. Together they are the meaning an agent reasons over — without them, it guesses.

The Fabric Data Agent

The Fabric Data Agent is the natural-language interface over that governed context. A user asks "what were UAE sales last quarter?" and the agent interprets the question against the semantic model and ontology and returns the answer — no SQL, no analyst queue. But it is only trustworthy because of what sits beneath it.

Keeping it honest takes deliberate design: ground it only on approved, governed sources; give it clear instructions on what it can and cannot answer; use the semantic model’s definitions; and enforce that it honours the asking user’s permissions, so it never answers using data they should not see. Then validate it before anyone relies on it — build a suite of golden questions with known-correct answers and confirm the agent returns them, exactly as you would test any production system.

An agent that has not been validated against known answers is not ready, however fluent it sounds. Fluency is not accuracy, and the gap between them is where trust is lost on the first wrong answer in a real meeting.

MCP and External Agents

When an external AI — Claude, say — needs to work with your Fabric data, the Model Context Protocol (MCP) is how it does so safely. MCP is a standard way for a model to discover and invoke tools and data capabilities exposed by an MCP server. The architecture is Claude → MCP client → MCP server → an approved capability → Fabric — with authentication, authorisation and audit at the boundary.

The contrast is the whole point. The safe pattern is Claude → MCP → authentication and authorisation → Fabric Data Agent → ontology and semantic model → OneLake. The unsafe pattern is Claude → database credentials → arbitrary SQL. The first exposes a bounded set of governed operations that honour your security; the second hands an external model the keys and hopes for the best.

So you secure the MCP server as you would any enterprise gateway: authentication, authorisation, least privilege, tool allow-listing, input validation and audit logging. An external agent should reach a curated set of governed capabilities, never the raw database — and it should never be able to run SQL you did not sanction. That boundary is what makes external AI on enterprise data defensible.

Right: Claude → MCP → auth → Fabric Data Agent → ontology → OneLake. Wrong: Claude → database credentials → SQL. Expose governed capabilities through an authenticated gateway, never the raw database.

Governance, Security and CI/CD

Three things wrap the whole stack. Governance: the OneLake catalog and Microsoft Purview provide discovery, classification, lineage and endorsements across every layer, so agents ground on trusted, classified sources rather than the nearest table. Governance is not a box at the side — it spans the stack.

Security: Entra groups → workspace roles → OneLake security → row- and object-level security, least privilege throughout. Because the Data Agent and any MCP-exposed capability inherit that security, wider AI access never becomes wider data exposure — the access model designed for reports is exactly what protects the agents.

And delivery: Git integration and DEV→QA→UAT→PROD deployment pipelines, so the whole architecture — including agents and their validation suites — is promoted through tested, reversible releases rather than changed live. An agent going to production should pass the same gates as any other system: architecture and dependency checks, data and functional tests, security tests, UAT with golden questions, and sign-off with a rollback plan.

So What — the Full Picture

The full enterprise picture reads bottom to top: sources → OneLake → Bronze→Silver→Gold on Delta Lake → semantic model → Fabric IQ ontology → Fabric Data Agent → MCP → external agents like Claude, with Purview and the OneLake catalog governing every layer, Entra-to-RLS/OLS securing it, and Git and CI/CD deploying it. Every layer has a job, and the trustworthiness of the top depends on the discipline of the bottom.

So the answer to "can we put AI on our data?" is not "yes, connect the model". It is "yes — once the Gold layer is curated, the measures are defined once, the ontology models the business, the agent is grounded and validated, external access goes through a governed MCP gateway, and governance and security wrap all of it". That is the architecture. Anything shorter is the confident-and-wrong version.

We build Microsoft-first by default, and the same shape maps to a Databricks estate — Unity Catalog governing, Databricks Genie as the natural-language interface, the same MCP boundary for external agents. The platform follows the estate; the principle does not: ground agents in governed, defined, secured data, and never give an external model direct database access.

Sources → OneLake → medallion → semantic model → Fabric IQ ontology → Data Agent → MCP → external agents. Governance, security and CI/CD across all of it. The trust at the top is built from the discipline at the bottom.

If leadership wants AI over your Fabric data and you want it trustworthy — grounded, governed and safe for external agents — that is exactly the architecture to get right first. 30 minutes with Amit on your stack, from OneLake to the agent layer and the MCP boundary. No slides. No pitch deck. No obligation to proceed.

Free Assessment

Where does your operation sit on the data maturity curve?

8 questions. 3 minutes. You get a scored breakdown across data infrastructure, analytics readiness, and automation potential — with a specific next step for your industry.

Data ArchitectureMicrosoft FabricFabric IQAI & AutomationMCPFabric Data Agent

Your Data · Our Technology · Our Automation

Get practical insights every fortnight

Amit writes about Microsoft Fabric, Power BI, AI in operations, and digital transformation for manufacturing and supply chain leaders. Practitioner perspective - no fluff, no vendor spin.

No spam. Unsubscribe any time. Also on Substack.

FAQ

Common questions

What does an enterprise Microsoft Fabric plus AI architecture look like?

A layered stack: sources feed OneLake and a medallion lakehouse (Bronze→Silver→Gold) on Delta Lake; a governed semantic model defines each measure once; a Fabric IQ ontology models the business entities and relationships; a Fabric Data Agent provides natural-language access over that governed context; and MCP exposes governed capabilities to external agents through an authenticated gateway. Microsoft Purview and the OneLake catalog govern every layer, Entra-to-RLS/OLS security wraps it, and Git with DEV→QA→UAT→PROD CI/CD deploys it.

Why do AI agents give confident but wrong answers over company data?

Almost always because of a missing architecture layer, not a weak model. No curated Gold layer, no semantic definitions, no ontology, or no access control — any of these leaves the agent reasoning over chaos, and it answers accordingly. Prompting harder does not fix a foundation problem. The fix is to build the stack: a governed Gold layer, one definition of each measure, an ontology of the business, and a grounded, validated agent.

What is MCP and why use it with Fabric?

MCP — the Model Context Protocol — is a standard way for an AI model to discover and invoke tools and data capabilities exposed by an MCP server. With Fabric, it lets an external agent like Claude work with your data through a governed interface: Claude → MCP client → MCP server → an approved capability → Fabric, with authentication, authorisation and audit at the boundary. It exposes a bounded set of governed operations rather than raw database access, which is what makes external AI on enterprise data defensible.

Should you connect an external AI model directly to your database?

No. Handing an external model database credentials and the ability to run arbitrary SQL gives away control of what it can do. The safe pattern routes it through an abstraction layer — Claude → MCP → authentication and authorisation → Fabric Data Agent → ontology and semantic model → OneLake — so it reaches only a curated set of governed capabilities that honour your security. Secure the MCP server with authentication, authorisation, least privilege, tool allow-listing, input validation and audit logging.

How do you stop a Fabric Data Agent from hallucinating or leaking data?

Ground it only on approved, governed sources; give it clear instructions on what it can and cannot answer; use the semantic model’s definitions and a Fabric IQ ontology for business meaning; enforce that it honours the asking user’s row- and object-level security; and validate it before rollout against a suite of golden questions with known-correct answers. An agent that inherits proper security cannot answer using data the user should not see, and one validated against known answers is far less likely to be confidently wrong.

Related FAQs

Questions operations leaders ask

Is this the challenge you're facing?

Book a 30-minute call. We'll look at your specific operation and tell you what's achievable - plainly and without slides.