• datapro.news
  • Posts
  • Meaning Before Models: The Semantic Foundation Enterprise AI Needs

Meaning Before Models: The Semantic Foundation Enterprise AI Needs

THIS WEEK: Why enterprise AI needs a semantic foundation, and how to build one on top of your data platform.

Dear Reader…

Enterprise AI rarely fails loudly. It fails with a clean chart, a plausible number and no warning that the number is wrong. When that happens, the model usually did its job. It simply had no reliable way to know what the business means by its own words.

That gap is now widely recognized. Gartner, in a February 2025 analysis, predicted that through 2026 organizations will abandon 60% of AI projects that lack AI-ready data. It also found that 63% of organizations are unsure they have the right data management practices for AI.

This article looks at the fix from the ground up: how business meaning is defined, where it should live in a modern data platform, and how it reaches an AI agent at the moment the agent answers a question.

Why a valid query can still be wrong

Ask three systems how many active customers you have, and you may get three answers. Sales counts recent logins. Finance counts accounts with an open balance. An older warehouse relies on a status flag that nobody has updated in years. Each number is correct by its own rules.

An AI agent sees only tables and columns. Nothing in the schema says which definition applies to the question being asked. So the agent picks one, or blends them, and returns a confident result. The SQL runs, the answer looks reasonable, and no one is told a choice was made.

Research shows how much this matters. In a peer-reviewed benchmark on an enterprise insurance database, Sequeda, Allemang and Jacob found GPT-4 answered 16% of questions correctly against the raw schema. With an ontology describing the business concepts, accuracy rose to 54%.

A 2026 study by researchers at Cube tested three current frontier models on retail data. Adding a short document of business definitions lifted accuracy by 17 to 23 percentage points for every model. Just as telling, the models performed about equally within each condition. The business context mattered more than the choice of model.

Writing meaning down

The cure is to make business meaning explicit and machine-readable. Practitioners usually describe this in three layers, each adding something the one before it cannot express.

Layer

Purpose

Example

Taxonomy

A shared vocabulary, organized as a hierarchy

Category → Sub-category → SKU

Ontology

Concepts, the relationships between them, and the rules that govern them

A Customer places an Order; an Order contains Products

Semantic model

The link between those concepts and the actual tables, columns and calculations

"Active customer" = this rule, applied to these sources

The first two are well established. W3C standards such as SKOS for taxonomies and OWL for ontologies have existed for more than a decade. The third layer is where most enterprises fall short, because it has to reconcile what each source system actually stores.

For a clear, practical walkthrough of all three layers, see Ignition's guide From Taxonomy to Trusted AI. It works through the "active customer" problem step by step and shows how each layer maps onto a Data Vault model.

Data Vault as the integration backbone

Meaning has to rest on integrated, trustworthy data. In lakehouse terms, that integration happens in the silver layer, and Databricks' own reference architecture names Data Vault as a common way to model it.

Data Vault, set out by Dan Linstedt and Michael Olschimke, organizes data into three structures. Hubs hold the business keys of core concepts, such as a customer number. Links record relationships between concepts, such as which customer placed which order. Satellites hold the descriptive details and their full history, usually kept separate per source system.

This structure echoes an ontology's concepts and relationships, which makes it a natural host for business meaning. Two distinctions keep that comparison honest:

  • Raw versus business. The Raw Vault stores data the way each source delivers it, so conflicting definitions remain visible and traceable. The Business Vault is where those conflicts are resolved and a single rule for "active customer" is applied.

  • Storage versus serving. A vault's many narrow tables are built for integration and audit, not for answering questions. Agents are better served by marts and a semantic layer built on top of it.

In short, Data Vault keeps the evidence and applies the rules. It is the foundation for meaning, not the place where an agent should look for it.

From model to agent

A definition only helps if the agent can read it at the moment it answers. That is the job of the semantic layer: one place where metrics, dimensions and business rules are defined, and from which every consumer draws the same answer. When a definition changes there, it changes everywhere.

Two trends make this layer more reachable for AI:

  • Open interfaces. The Model Context Protocol, an open standard for connecting AI applications to data and tools, lets a semantic layer or catalog expose its definitions to any compatible agent. The specification also requires user consent and access controls, which matters for governed data.

  • Agent-aware catalogs. Metadata platforms are adding interfaces built for agents, so an agent can retrieve the definition, lineage and owner of a dataset along with the data itself.

The principle is simple. Give the agent the governed definition that fits the question, rather than raw tables or an unfiltered pile of documentation.

For a hands-on look at how this works in practice, see Scalefree's Semantic Layer Guide. It explains how centrally defined metrics sit on top of a Data Vault–based warehouse, so dashboards and AI agents alike use the same definitions.

Analytics on Live Data Without Leaving Postgres

When analytics on Postgres slows down, most teams add a second database. Then they manage pipelines, sync lag, and drift forever. TimescaleDB extends Postgres instead. Analytics run on live data, in the database you already have. No pipeline. No migration. No new query language.

Automation, with people in the loop

Building this foundation by hand is slow, and slow work is where definitions drift. Automation now helps in two ways. Pattern-based generators turn an agreed model into consistent loading code. Newer AI-assisted tools go further, profiling source data and proposing business keys and relationships.

The second kind needs care. A tool can infer structure from data, but it cannot infer a business rule that was never written down. Treat its suggestions as drafts: domain experts confirm them, and automated tests guard what reaches production. The goal is machines proposing and people deciding, not the other way round.

Where to start

The evidence points one way: AI answers improve most when business meaning is written down and handed to the model. A practical path looks like this:

  1. Pick one high-value concept, such as "active customer," and list every definition in use today.

  2. Integrate the sources traceably, so every conflict stays visible.

  3. Agree on one rule and apply it in a business layer.

  4. Publish it in a semantic layer that dashboards and agents both use.

  5. Expand concept by concept, with experts signing off each definition.

To go deeper, the two guides referenced above cover the ground from both ends. Ignition's From Taxonomy to Trusted AI explains the layers of meaning and how they map onto Data Vault. Scalefree's Semantic Layer guide shows how to deliver those definitions to every tool that needs them.

That’s a wrap for this week
Happy Engineering Data Pro’s