Skip to main content
Data Governance

How the OneLake Catalog Makes Copilot and AI Agents Trustworthy

The reason AI over your data gives confident wrong answers is rarely the model. It is that the model grounded on the wrong data. The OneLake catalog is the unglamorous governance layer that decides what your agents reason from.

Amit Kumar Singh - Technology Consulting Partner at MyData Insights

Technology Consulting Partner · MyData Insights

14+ years in industrial data · Former Accenture & EY · India, GCC, SEA

28 September 2026 · 8 min read

The bottom line

AI over your data gives confident wrong answers mostly because it grounded on the wrong data, not because the model is weak. The OneLake catalog is the governance layer that fixes this: it makes trusted data discoverable, marks the authoritative sources with endorsements, surfaces sensitivity so agents respect confidential data, and is queryable by agents through public APIs and the Fabric MCP. Ground Copilot and Fabric Data Agents on certified, classified catalog items and their answers become far more reliable; leave them to reason over an ungoverned estate and they ground on whatever is nearest. Governance stopped being a back-office compliance task the moment agents started grounding on your data — it became the thing that decides whether the answers can be trusted.

It Is Not the Model, It Is the Grounding

When Copilot or an AI agent gives a confident wrong answer about your business, the instinct is to blame the model. Usually the model is fine. The problem is that it grounded on the wrong data — an out-of-date copy, an ambiguous table, a dataset with a different definition of the measure than the one the user meant. The reasoning was sound; the source was wrong.

This reframes what makes AI over your data trustworthy. It is not primarily a modelling problem or a prompting problem; it is a data governance problem. The question that decides answer quality is "what did the agent ground on", and that is exactly what a catalog governs.

The OneLake catalog is the unglamorous layer that answers that question well. It is not the exciting part of an AI project, but it is the part that decides whether the exciting part can be trusted.

When AI gives a confident wrong answer, the model is usually fine — it grounded on the wrong data. Trustworthy AI over your data is a governance problem, and the question that decides answer quality is "what did the agent ground on".

Making Trusted Data Findable

An agent, like a person, can only use data it can find. In an ungoverned estate, the trusted dataset and three stale copies all look the same, so the agent may ground on any of them. The OneLake catalog makes the estate discoverable — tags, descriptions, ownership, search — so the right data is findable rather than lost among duplicates.

Discoverability is the precondition for grounding well. When the catalog is curated so the current, owned datasets are clearly described and easy to locate, both people and agents that query the catalog can reach the right source. A catalog nobody has curated is just a list, and an agent grounding on an uncurated list is grounding on chance.

Crucially, the catalog is queryable by agents themselves — through public APIs and the Fabric MCP — so an agent can discover the right data programmatically rather than being hard-wired to a fixed source. That is what lets grounding adapt as the estate changes.

Grounding on Endorsed Sources

Discoverability gets the agent to the data; endorsements tell it which data to trust. Certified and promoted endorsements in the catalog mark the authoritative sources, so an agent configured to prefer endorsed items grounds on the dataset the organisation stands behind rather than the nearest match. This is the single most effective lever on agent answer quality.

The practice is to ground your Copilot and Fabric Data Agents on certified, endorsed catalog items as the trusted set, rather than pointing them at the whole estate. The endorsement scheme — certification controlled tightly, promotion used generously — becomes the trust signal the agents consume, which is why a well-run endorsement practice is now part of AI enablement, not just governance tidiness.

Without this, an agent has no way to distinguish authoritative from incidental data and grounds on whatever it finds, which is the mechanism behind most confident wrong answers. Endorsements are how you give the agent the judgement it otherwise lacks.

Discoverability gets the agent to the data; endorsements tell it which data to trust. Grounding agents on certified, endorsed catalog items is the single most effective lever on answer quality.

Respecting Sensitivity

Trustworthy is not only about correctness; it is also about not exposing what should stay confidential. An agent that answers a question using data the asker should not see is a governance failure even if the answer is right. The OneLake catalog, with Microsoft Purview sensitivity classification flowing onto Fabric data, surfaces which items are sensitive so agents respect those boundaries.

Combined with row-level and object-level security on the underlying data, this means an agent honours the same access rules as your reports — a user only gets answers from data they are entitled to. Grounding on classified, secured catalog items is what keeps wider AI access from becoming wider data exposure.

This is the part that lets governance and security teams say yes to AI over company data. An agent grounded on an endorsed, classified, access-controlled set is one they can trust not to leak, which is often the difference between an AI initiative that ships and one that stalls in review.

So What — Governance Is Now an AI Problem

The OneLake catalog makes Copilot and AI agents trustworthy in three connected ways: it makes trusted data discoverable so agents can find the right source, endorsed so they know which to trust, and classified so they respect what is confidential — and it is queryable by the agents themselves. Ground your agents on the curated, endorsed, classified catalog and their answers become reliable; leave them on an ungoverned estate and they ground on whatever is nearest.

The wider point is that governance changed jobs. It used to be a back-office compliance exercise that produced reports nobody read. The moment agents started grounding on your data, governance became the thing that decides whether the answers can be trusted — a direct input to AI quality, not a cost centre. That is a reason to invest in the catalog now, not after the first embarrassing wrong answer.

So if you are deploying Copilot or Fabric Data Agents, the catalog work is not a separate governance project to do later — it is part of the AI project itself, because it is what makes the AI trustworthy. Foundation first, as ever: discoverable, endorsed, classified data is the foundation, and the agent is the visible payoff of getting it right.

Governance changed jobs: the moment agents ground on your data, it became the thing that decides whether the answers can be trusted. The catalog work is part of the AI project, not a separate governance task for later.

If you are deploying Copilot or Fabric Data Agents and worried about confident wrong answers, the fix is usually the data they ground on, not the model. 30 minutes with Amit on setting up the OneLake catalog so your agents reason from endorsed, classified, trusted data. No slides. No pitch deck. No obligation to proceed.

Free Assessment

Where does your operation sit on the data maturity curve?

8 questions. 3 minutes. You get a scored breakdown across data infrastructure, analytics readiness, and automation potential — with a specific next step for your industry.

Data GovernanceOneLake CatalogCopilotFabric Data AgentMicrosoft FabricAI & Automation

Your Data · Our Technology · Our Automation

Get practical insights every fortnight

Amit writes about Microsoft Fabric, Power BI, AI in operations, and digital transformation for manufacturing and supply chain leaders. Practitioner perspective - no fluff, no vendor spin.

No spam. Unsubscribe any time. Also on Substack.

FAQ

Common questions

Why does Copilot give wrong answers about our data?

Usually because it grounded on the wrong data, not because the model is weak — an out-of-date copy, an ambiguous table, or a dataset with a different definition of the measure than the user meant. The reasoning is sound but the source is wrong. This makes trustworthy AI over your data primarily a governance problem: the question that decides answer quality is what the agent grounded on, which is exactly what the OneLake catalog governs.

How does the OneLake catalog make AI agents more trustworthy?

In three connected ways: it makes trusted data discoverable through search, tags and descriptions so agents can find the right source; it marks authoritative sources with certified and promoted endorsements so agents ground on data the organisation stands behind rather than the nearest match; and it surfaces Purview sensitivity so agents respect confidential data. It is also queryable by agents through public APIs and the Fabric MCP, so they discover the right data programmatically.

Should agents ground on all our data or a curated set?

A curated, endorsed set. Pointing an agent at the whole estate means it grounds on whatever it finds, including stale copies and ambiguous tables, which is the mechanism behind most confident wrong answers. Ground Copilot and Fabric Data Agents on certified, endorsed, classified catalog items as the trusted set, and their answer quality improves markedly. The endorsement scheme becomes the trust signal the agents consume.

How do we stop an AI agent from exposing sensitive data?

Ground it on classified, secured catalog items and enforce the underlying row-level and object-level security so the agent honours the same access rules as your reports — a user only gets answers from data they are entitled to. Microsoft Purview sensitivity classification flowing onto Fabric data, surfaced in the OneLake catalog, lets agents respect confidentiality boundaries. An agent grounded on an endorsed, classified, access-controlled set is one governance teams can trust not to leak.

Related FAQs

Questions operations leaders ask

Is this the challenge you're facing?

Book a 30-minute call. We'll look at your specific operation and tell you what's achievable - plainly and without slides.