Data Catalog vs Metadata Layer vs Semantic Layer: Where Governance Actually Lives

Intergalactic Data Labs
—
—
14 min read
Updated October 2026
A data catalog tells you what data exists and who owns it. A metadata layer carries schemas, lineage, tags, and policies to the engines that enforce them. A semantic layer defines what business metrics mean and how to compute them. Governance lives across all three, and in 2026 catalog and warehouse vendors are pulling each layer closer to AI agents.
Layer | Main job | What it governs | Typical owner | Example tools |
|---|---|---|---|---|
Data catalog | Find and document data | Ownership, glossary, certification, audit | Governance team, CDO | Atlan, Collibra, Alation, DataHub, Purview |
Metadata layer | Collect and push metadata | Lineage, tags, access and masking policies | Platform engineering | Unity Catalog, Snowflake Horizon, OpenLineage |
Semantic layer | Define and serve metrics | Metric logic, joins, entity definitions | Analytics engineering | dbt Semantic Layer, Snowflake semantic views, Databricks metric views, LookML |
Context layer | Serve all of the above to AI | What an agent may see and how it reads it | Data and AI platform teams | Atlan MCP server, DataHub MCP, dbt MCP server |
What is the difference between a data catalog, a metadata layer, and a semantic layer?
The difference is the question each one answers. A catalog answers where the data is and who is responsible for it. A metadata layer answers how that knowledge reaches every engine and tool. A semantic layer answers what a number means and how to calculate it the same way everywhere.
Atlan puts it plainly in its own comparison. It calls the catalog a system of record for what exists and who owns it, and says it is "documentation, not execution" (Atlan). The same page calls the semantic layer the place where a metric like revenue gets its calculation. That split holds up across vendors, even as product names change.
The metadata layer is the least visible of the three. It is rarely sold as its own product. Instead it shows up as the open catalog inside your platform, the lineage feed between tools, and the policy engine that turns a sensitivity tag into a masked column.
Where does governance actually live?
Governance lives in three places at once, and each place handles a different kind of rule. Access rules are enforced where queries run. Meaning is enforced where metrics are defined. Accountability is recorded where people search and certify data. The table below maps common governance tasks to the layer that owns them.
Governance task | Where it is enforced | Example mechanism |
|---|---|---|
Mask a PII column | Engine or metadata layer | Snowflake masking policy |
Restrict rows by region | Engine or metadata layer | Row access policy |
Define active customer | Semantic layer | Metric or entity definition |
Keep revenue consistent across tools | Semantic layer | One metric served to every BI tool |
Name an owner and certify a table | Catalog | Stewardship workflow |
Prove where sensitive data flows | Catalog plus lineage | Column-level lineage |
Limit what an AI agent can query | Context or semantic layer | Governed metrics over MCP |
The practical lesson is that a policy written in one layer is only as good as its reach into the others. A catalog can label a column as sensitive, but the warehouse has to mask it. Snowflake, for example, enforces dynamic masking in the engine, and that feature requires Enterprise Edition or higher (Snowflake docs).
What does a data catalog do in 2026?
A data catalog is still the searchable inventory of an organization's data, with owners, glossary terms, classifications, and lineage attached. What changed in 2026 is the pitch. Most major catalogs now describe themselves as the layer that feeds trusted context to AI agents, not only to human analysts.
Atlan now leads with "The Context Layer for AI" and ships an MCP server that serves certified context to agents (Atlan). Collibra calls itself an enterprise AI control plane and in June 2025 bought Raito to add data access governance (Collibra, Collibra press release). Alation now brands its platform the Alation Intelligence Operating System, with catalog, lineage, ontologies, and agents (Alation).
Where catalogs are strong
Catalogs are best at discovery, stewardship, and audit evidence. They give compliance teams one place to show where sensitive data lives and who approved access. They also give analysts a starting point that is better than asking in a chat channel which table is correct.
Where catalogs stop
A catalog describes meaning in words, but it does not compute anything. A glossary entry for net revenue does not stop two dashboards from calculating it differently. That gap is why catalogs increasingly sync semantic models from dbt, Snowflake, or Databricks instead of trying to own metric logic themselves.
What is a metadata layer?
A metadata layer is the infrastructure that collects technical and business metadata and moves it to the systems that act on it. It holds schemas, lineage, tags, and policies, and it pushes them into query engines, BI tools, and pipelines. In most 2026 stacks it is the platform's own catalog plus open standards.
Two platform catalogs now do most of this work. Snowflake Horizon Catalog offers column-level lineage across Snowflake, external databases, BI tools, and OpenLineage feeds, and enforces masking and row policies across Iceberg REST compatible engines (Snowflake docs). Unity Catalog is open source under Apache 2.0 as an LF AI and Data sandbox project, and governs tables, files, functions, and AI assets (Unity Catalog).
Why the metadata layer decides whether governance works
The metadata layer is where a tag becomes an action. When a steward marks a column as confidential, the metadata layer is what makes that label change a masking rule, a lineage alert, or an agent permission. Without it, governance becomes a manual task repeated in every tool.
Why open formats matter here
Open table formats and open catalog APIs let several engines read the same governed data. Horizon exposes the Iceberg REST Catalog API so Spark, Flink, and Trino can query Snowflake managed tables (Snowflake docs). That shift moves metadata out of single tools and into a shared layer that the company controls.
What is a semantic layer?
A semantic layer defines business metrics, dimensions, and entities once and serves them to every consumer. It turns tables and joins into terms like revenue, churn, and active subscription. When the definition changes, every dashboard, notebook, and AI agent picks up the change. Every major data platform now ships one.
The dbt Semantic Layer is powered by MetricFlow and lets AI tools query governed metrics through the dbt MCP server (dbt docs). MetricFlow itself is open source under Apache 2.0 (dbt Labs). Snowflake semantic views store facts, dimensions, and metrics as schema objects that Cortex Agents can query (Snowflake docs).
Databricks metric views are Unity Catalog objects that separate measures from dimensions, and they can be queried from SQL, Genie, and external tools like Power BI, Tableau, and Sigma (Databricks docs). Looker continues to use LookML as its semantic modeling language (Google Cloud).
How semantic definitions move between tools
Portability used to be the weak point. The Open Semantic Interchange, a shared YAML format for metrics and dimensions, entered the Apache Incubator on July 10, 2026 and was renamed Apache Ossie (Apache Ossie). The project says it grew from 17 launch partners to more than 50 organizations, and the spec did not change with the rename.
Semantic layer versus ontology
A semantic layer computes metrics. An ontology defines what concepts are and how they relate, often in standards like OWL and RDF. Many teams now use both, with the ontology describing entities such as customer and contract, and the semantic layer computing numbers over them. Our guide to how ontology powers AI analytics covers that split in more detail.
Which tools cover which layer?
Most products now span more than one layer, so the useful question is what each does best. The table below lists one real strength and one documented limit for each, based on vendor documentation checked in October 2026.
Tool | Main layer | Strength | Documented limit |
|---|---|---|---|
Atlan | Catalog and context | MCP server serves certified context to agents | Pricing is quote only |
Collibra | Catalog and governance | Long track record in governance and access control | Access enforcement added via 2025 Raito deal |
Alation | Catalog | Catalog, lineage, and ontologies in one platform | Pricing is quote only |
DataHub | Catalog and metadata | Open source core with MCP support | Monitoring and access workflows are Cloud only |
Microsoft Purview | Catalog | Data products and glossary tied to Microsoft estate | Unified Catalog needs the enterprise version |
Snowflake Horizon | Metadata | Lineage and policies across Iceberg engines | Masking needs Enterprise Edition |
Unity Catalog | Metadata and semantic | Open source core, metric views in Databricks | Open source project is still sandbox stage |
dbt Semantic Layer | Semantic | Apache 2.0 MetricFlow, MCP server | Hosted service needs a paid dbt plan |
Snowflake semantic views | Semantic | Native objects used by Cortex Agents | Definitions live inside Snowflake |
Looker LookML | Semantic | Mature modeling language for BI | Managed MCP server is still in preview |
Sources for the table are Atlan pricing, Alation pricing, DataHub Cloud, DataHub, Microsoft Learn, Unity Catalog, dbt docs, and Looker MCP docs. For a ranked view of governance and lineage products, see our data governance, lineage, and catalog tools guide.
Is the context layer a fourth layer?
The context layer is better understood as a delivery pattern than a new layer. It packages catalog metadata, semantic definitions, and policy, then serves them to an AI agent when the agent asks a question. Atlan describes it this way and argues the tiers are "permanent and co-existing" rather than stages that replace each other (Atlan).
The practical change is the interface. Catalogs, semantic layers, and warehouses now expose MCP servers, so an agent can look up a certified table, read a governed metric, and respect access rules in one session. DataHub, for example, promotes native MCP integration with agents such as Claude, Cursor, and Genie (DataHub).
Why does the semantic layer matter for AI agents?
AI agents answer business questions more accurately when they query defined metrics instead of raw tables. Without definitions, a model has to guess joins, filters, and which column means revenue. Several published tests show the size of that gap, though each uses a different setup.
Study | Setup | Result |
|---|---|---|
data.world, 2023 | GPT-4 on insurance questions | 16% raw SQL vs 54% knowledge graph |
Academic study, April 2026 | Explicit business semantics added | 17 to 23 point gain across models |
dbt Labs, 2026 | 11 questions, vendor benchmark | 98 to 100% semantic layer vs 84 to 90% text to SQL |
The sources are the 2023 benchmark paper, the April 2026 study, and the dbt benchmark post. Treat vendor numbers with care, since each vendor picks its own questions. The direction is consistent, and it matches why we compare these approaches in RAG vs knowledge graph vs semantic layer.
The stakes are real. Gartner predicts that over 40 percent of agentic AI projects will be canceled by the end of 2027, citing costs, unclear value, and weak risk controls (Gartner). Governed definitions and enforced access rules address two of those failure modes directly.
Which layer should you build first?
Build first where the pain is sharpest, then connect the rest through shared metadata. There is no single correct order. The sequence depends on whether the problem is finding data, controlling it, or agreeing on what it means.
If people cannot find or trust data
Start with a catalog and engine-level policies. Inventory the critical sources, name owners, and tag sensitive columns. Make sure those tags actually drive masking and row rules in the warehouse, or the catalog becomes a document that nobody enforces.
If dashboards or agents disagree on numbers
Start with a semantic layer for the ten to twenty metrics leadership watches most. Define them in a format that more than one tool can read, such as MetricFlow YAML or a platform semantic view with an Ossie export path. Then point BI tools and agents at those definitions instead of raw tables.
If AI agents are the main consumer
Start with the semantic layer and the access rules together. An agent needs both a correct definition of the metric and a clear limit on what it may read. Our overview of data governance for AI agents covers the policy side.
Who should own the business definitions?
The company should own them. Metric logic, entity definitions, and ontologies are a written record of how the business works, and that knowledge is one of its most valuable assets. If it lives only inside one BI tool or one SaaS catalog, changing tools means rebuilding it from memory.
Keeping definitions in open formats on infrastructure you control avoids that trap. MetricFlow is Apache 2.0, Unity Catalog's open source core is Apache 2.0, and Ossie gives metrics a shared YAML shape. A catalog can then index those definitions without becoming the only copy.
Some teams build this layer with a partner on their existing stack. Galaxy is one such partner, using internal tooling that speeds up semantic layer, ontology, and entity resolution work. It builds in the open, starting with Filament, and plans to open source more of that tooling.
Frequently asked questions
What is the difference between a data catalog and a semantic layer?
A data catalog records what data exists, where it lives, who owns it, and how sensitive it is. A semantic layer defines how business metrics such as revenue or churn are calculated and serves those definitions to BI tools and AI agents. The catalog answers where the data is. The semantic layer answers what the number means and how to compute it.
Is a metadata layer the same as a data catalog?
Not quite. The metadata layer is the technical plumbing that collects schemas, lineage, tags, and policies and pushes them to the engines that enforce them. The catalog is the application people use to search, document, and steward that metadata. Many products bundle both, so the difference shows up in where a policy is actually enforced, not in the product name.
Where does data governance actually live?
It lives in three places at once. Access rules such as masking and row filters are enforced in the engine or open catalog. Business meaning, such as the definition of active customer, lives in the semantic layer. Ownership, certification, and audit evidence live in the catalog. A governance program fails when one of the three is missing or disconnected.
Can a data catalog replace a semantic layer?
No. A catalog can hold a glossary entry that describes revenue in plain words, but it does not compute the metric or guarantee every tool computes it the same way. A semantic layer stores the executable definition and answers queries against it. Some catalogs now sync semantic models from dbt or warehouses, which links the two without merging them.
What is a context layer and is it a new category?
Context layer is the 2026 label several catalog vendors use for serving catalog metadata, semantic definitions, and policy to AI agents at query time, often through an MCP server. Atlan describes it as an activation tier on top of the catalog and semantic layer. It is a new delivery pattern for existing layers rather than a replacement for them.
Do I need a semantic layer if I already use dbt?
dbt already includes one. The dbt Semantic Layer is powered by MetricFlow, which is open source under Apache 2.0, and it serves metrics to BI tools and to AI tools through the dbt MCP server. The managed Semantic Layer service requires a paid dbt platform plan. Teams on Snowflake or Databricks can also use semantic views or metric views.
What is Apache Ossie?
Apache Ossie is the new name of the Open Semantic Interchange, a vendor-neutral YAML format for metrics and dimensions that started at Snowflake in 2025. It entered the Apache Incubator on July 10, 2026, with more than 50 participating organizations. The goal is to let a metric defined once move between warehouses, BI tools, catalogs, and AI agents.
Which layer should a company build first?
Start with the pain. If people cannot find data or prove where sensitive fields are, start with a catalog and engine-level policies. If dashboards disagree or AI agents give wrong numbers, start with a semantic layer for the ten to twenty metrics that matter most. Most companies end up running all three, connected through shared metadata.
Do semantic layers make AI agents more accurate?
Published tests point that way. A 2023 benchmark found GPT-4 answered 16 percent of enterprise questions correctly over raw SQL and 54 percent over a knowledge graph. An April 2026 study found explicit business semantics added 17 to 23 points across current models. A dbt vendor benchmark reported 98 to 100 percent accuracy with its semantic layer.
Who should own the business definitions?
The company should own them, in open formats it controls. Metric definitions, entity models, and ontologies capture how the business actually works. If they live only inside one BI tool or one SaaS catalog, switching tools means rebuilding that knowledge. Keeping them in formats such as MetricFlow YAML or Ossie keeps them portable across engines and agents.
More articles
Questions
Answered
FAQ
What does Galaxy do?
What is Filament?
What is enterprise context management?
What does working with Galaxy look like?
How do you handle security and compliance?
Why does Galaxy build in the open?
Company
Copyright © 2026 Galaxy. All rights reserved.