A single governed source of truth for every AI agent and platform: Introducing Collibra’s governed context compiler
As organizations scale AI agents, the context those systems rely on determine whether they succeed or fail. A developer asking Snowflake Cortex about "monthly active users" may get a different answer than a colleague asking Databricks Genie the same question, because no single governed definition exists to ground either response. According to Gartner, 60% of AI projects will be abandoned due to poor data readiness, and missing and fragmented context sits at the heart of that failure. The organizations that crack this problem share a common trait: they stop treating governed context as something that lives in a catalog and start treating it as something that actively flows to every agent, platform and tool that needs it.
What’s new: Collibra’s governed context compiler
With Collibra’s governed context compiler, organizations can close the gap between governed context and the AI systems and platforms that depend on it, delivering trusted, structured context to any downstream consumer, in the exact formats they require, without custom engineering for every integration. The governed context compiler is a rule-based extraction engine that pulls governed context including semantic models, governed measures, business terms, data quality signals, policies, ownership and more, out of the Collibra Knowledge Graph and delivers it in the structured format a downstream consumer (agent or human) requires. It is deterministic by design: pre-configured context specifications produce precise, repeatable results every time, making it suitable not just for one-off data engineering tasks but for production AI agents and automated pipelines that need to trust what they receive.
The governed context compiler works in two stages: context specification configuration, where teams define what to extract and how to structure it, and context delivery, where that governed context is made available through multiple channels. Through a visual, no-code UI, teams configure exactly how their Collibra assets map to target schemas like Snowflake Semantic Views, Databricks Metric Views, Open Semantic Interchange (OSI) or Open Data Contract Standard (ODCS). Collibra ships out-of-the-box context specifications for the most common formats so teams have a tested starting point, and custom formats are supported for proprietary schemas or emerging standards. Once configured, governed context flows to downstream consumers via REST API, MCP tools, YAML export, creating a reusable, composable foundation that scales across the entire data estate. And via Collibra’s Edge integrations, users can also ingest existing metric and semantic layer definitions from platforms like Snowflake, Databricks, Power BI and LookML into Collibra, so organizations can govern, enrich and maintain them as the authoritative source.
How the governed context compiler helps
The challenge of governed context isn't just about having the right information. It's about getting it to the right places, in the right format, reliably. Context including metric definitions, semantic models, business terms, data ownership, policies, data quality signals is all valuable, but only if the platforms and AI systems that depend on it can actually consume it. In practice, that context ends up scattered: defined inconsistently across Snowflake, Databricks, dbt and BI tools, locked in a catalog that downstream systems can't easily query, or simply absent. Every new platform or AI agent that needs business context requires another round of manual engineering to package and deliver it, and every manual step is an opportunity for drift. Meanwhile, the number of consumers demanding governed context keeps growing. AI agents, agentic frameworks, data platforms and emerging standards like OSI and ODCS all need context in different formats. As AI adoption accelerates, the problem compounds: more agents, more platforms, more context that needs to be consistent and trustworthy at scale.
Problems the governed context compiler solves:
- Context drift across platforms: Business metrics, definitions, relationships, synonyms, etc. end up inconsistently defined across Snowflake, Databricks, Power BI and other tools, each platform maintaining its own version with no governed source behind any of them. The result is unreliable AI outputs and eroding trust in data across the organization.
- AI hallucinations from ambiguous metadata: AI agents querying ungoverned data generate incorrect or inconsistent answers because definitions, relationships and synonyms are ambiguous, missing or inconsistent across platforms
- Manual, error-prone context packaging: Data engineers manually translate business definitions into platform-specific YAML formats: a slow, repetitive process where every translation is an opportunity for inconsistency and every upstream change triggers a round of manual rework.
- No flexibility for diverse operating models: Every organization structures its governed context differently: Different asset types, naming conventions, domain hierarchies and operating models. At the same time, target schemas are proliferating:Snowflake Semantic Views, Databricks Metric Views, OSI, ODCS and proprietary formats each expect a different output structure. A rigid, one-size-fits-all export satisfies neither side of that equation.
- Governed definitions and platform implementations stay disconnected: When governed context does exist, it typically lives separately from its technical implementations in Snowflake, Databricks or dbt, for instance, built and maintained independently with no link between them. When either side changes, the other doesn't know.
- Governance investment doesn't scale with AI adoption: As organizations deploy more AI agents and automated pipelines, the demand for trusted, machine-readable context grows faster than teams can manually supply it. Without a mechanism to deliver governed context programmatically and at scale, governance becomes a bottleneck rather than an enabler.
How the governed context compiler works
At the core of the governed context compiler is a mapping engine powered by Collibra's Neo4j-based Knowledge Graph, which enables fast, reliable traversal of complex asset relationships and returns a complete, structured context package in seconds. Users create context specifications: named configurations that define how to traverse the Knowledge Graph, what content to extract at each step and how to map that content to a target schema. Through the no-code UI, teams configure which assets to include, how they relate to one another and how they should map to the target format. Teams configure a context specification once and execute it against any qualifying asset without reconfiguration; when the configuration needs to change, they update it in the UI without touching any downstream integrations.
Collibra ships with out-of-the-box context specifications for Snowflake, Databricks, ODCS and OSI built around four core asset types in the Collibra context layer:
- Semantic models use data entities and data attributes to provide a stable abstraction layer over physical tables, including entity relationships and cardinality that give AI agents the navigation logic needed to join data correctly.
- Governed measures capture business intent, plain-English calculation rules and approved synonyms, going beyond standard technical metadata to codify what a metric means before specifying how it's computed.
- Business terms enrich data assets with additional business context, linking semantic assets to the broader governance layer of ownership, policy and definitions.
- Data products bundle physical data assets with their associated governed measures and business context, making them a natural starting point for extracting a complete, consumption-ready context package.
Inside the no-code visual UI, users can map relationships and Knowledge Graph traversals while instantly testing and validating the resulting YAML output on the right.
Once configured, governed context is delivered through multiple consumption paths:
- REST API: Provides programmatic access to governed context including metric definitions, calculation rules, approved synonyms and relationships, so data engineers can retrieve exactly what they need to accelerate platform development.
- MCP tools: Exposes the governed context compiler as a callable tool within the Collibra MCP Server, enabling AI assistants and agents to fetch governed, structured context at runtime via a standard protocol.
- YAML export: Generates valid, ready-to-deploy configuration files for Snowflake Semantic Views, Databricks Metric Views, ODCS, OSI or custom formats directly from governed Collibra definitions.
Inside the no-code visual UI, users can map relationships and Knowledge Graph traversals while instantly testing and validating the resulting YAML output on the right.
Users can also pull existing metric and semantic layer definitions from Snowflake, Databricks, Power BI, LookML and other sources directly into Collibra. This gives governance teams visibility into what's already implemented across the data stack, enabling them to govern those definitions in Collibra, enrich them with business context and identify drift between what's approved and what's actually deployed.
Teams can manage their pre-configured context specifications, mapping core assets to out-of-the-box or custom target schemas.
Why you should be excited
The governed context compiler redefines how governed context flows through the modern data stack, from Collibra outward to every platform, tool and AI agent that depends on it.
- AI and analytics teams: Reduce hallucinations caused by ambiguous metadata. AI agents can access governed context at runtime via MCP: business definitions, relationships, approved synonyms and calculation logic, helping to ensure every answer is grounded in a single, verified source of truth.
- Data Governance Managers: Define your organization's semantic standards once: metric names, definitions, approved synonyms, ownership, and have them flow automatically to every downstream consumer. When definitions change, update once in Collibra; every platform drawing from that context specification reflects the change without requiring manual re-integration work.
- Data Stewards: Own the authoritative definition of your metrics without needing to know how they're implemented in Snowflake or Databricks. The governed context compiler separates business intent from technical implementation: you define what "Gross margin" means and who owns it, and data engineers handle how it's computed in each platform.
- Data Engineers: Stop manually packaging metric definitions for each platform. Retrieve governed, machine-readable context via REST API or export a YAML file and build Snowflake Semantic Views or Databricks Metric Views directly from approved business definitions, cutting reconciliation work and eliminating rework when definitions change upstream.
- Data Leaders: Further establish Collibra as the context governance layer for the modern data stack: the place from which every AI agent, data engineer and BI tool draws its context, and the place to which platform implementations are reconciled. This is the foundation that makes scaling AI responsibly possible.
Use cases
Below are three scenarios illustrating how the governed context compiler works in practice:
Grounding AI agents with the right context for the right data product: A Databricks Genie agent answers natural-language questions about sales performance, but different data products in the organization define "ARR" differently, and the agent has no way to know which definition applies when. A data steward configures a context specification for the relevant sales data product that explicitly maps to the correct governed ARR measure, including its calculation rules, approved synonyms and the data attributes it depends on. When Genie queries that data product via MCP, it retrieves the context specification and gets exactly the right definition for that context, rather than a guess across competing versions.
Building metric views from a governed source: A data engineering team is standing up a new Snowflake Semantic View for a financial reporting data product. Instead of defining metric logic from scratch, they call Collibra's REST API with the relevant data product as the starting point. The governed context compiler runs the configured context specification, traversing the Knowledge Graph to return "Net revenue" and "Cost of goods sold", complete with plain-English definitions, approved synonyms, calculation context and the underlying data attributes. Collibra generates a valid Snowflake YAML file. The team deploys it knowing the logic is aligned with the business definition from day one, not reverse-engineered after the fact.
Bringing existing platform implementations under governance: A governance team discovers that the metric definitions powering their Snowflake Semantic Views were built by data engineers without reference to any official business definition, and no one is sure whether they're still accurate. Using metrics ingestion, the team pulls those existing definitions directly into Collibra, where stewards can review them, reconcile them against approved business definitions and fill in missing context. From that point, the governed version in Collibra becomes the authoritative source, and a context specification ensures every future consumer, whether an AI agent, a BI tool or a new platform integration, draws from it.
Key takeaways about the governed context compiler
The governed context compiler further positions Collibra as the context governance layer for the modern data stack, not just a catalog where metadata lives, but the engine from which every AI agent, data engineer and BI tool draws its semantic context. By separating business definitions from technical implementation, organizations can govern context once and propagate it everywhere: across Snowflake, Databricks and other platforms and AI systems, through REST API, MCP or YAML export, without custom scripts or backend deployments. And with metrics ingestion closing the loop between governed definitions and platform implementations, the result is a data estate where drift becomes structurally difficult, AI outputs are grounded in verified context and data teams spend their time on high-value work rather than manual reconciliation.
Where to learn more about the governed context compiler
Apply to the Private Preview today: Governed context compiler
Keep up with the latest from Collibra
I would like to get updates about the latest Collibra content, events and more.
Thanks for signing up
You'll begin receiving educational materials and invitations to network with our community soon.