← All projects
Personal Project · Case Study

Semantic Engine Governed context for AI agents

AI tools are only as accurate as their understanding of your data. Semantic Engine gives them a human-approved map of what every field means, so agents can answer questions about structured business data safely and accurately.

Enterprise AI MCP Data Governance BigQuery 0→1

Also explored under the working name Webster.

AI can’t reliably use data it doesn’t understand

Most organizations have data scattered across many systems, with names and definitions that only a few people really understand. Centralizing it into one warehouse is a big project, and building purpose-built datamarts for each reporting need only adds to the fragmentation.

That’s a problem for AI. Pointed at a raw schema, an LLM guesses what status or region means, writes queries that look right and aren’t, and has no way to explain where a number came from. Analysts, ML models, governance teams, and AI agents all need one shared source of truth about what the data is.

I saw this firsthand building an enterprise analytics and AI platform across 16 operating companies. The tools that solve it, like Collibra and Alation, are heavy and expensive. Mid-market teams and startups have nothing lightweight, and nothing built for LLMs from the start.

AI drafts it. People approve it. Agents use it.

01

Connect, don’t migrate

Point it at a BigQuery project, scoped to specific datasets and tables. It reads the schema in place, and nothing is copied out of the warehouse. A connector model is built for other warehouses.

02

Draft every field

An LLM reads each table’s schema and a sample of values, then drafts a plain-English definition per column. Every draft lands in a review queue as pending. Nothing is trusted yet. The design target is cutting manual authoring by more than 70%.

03

A human signs off

Triage in bulk in Scan mode, or go field by field in a keyboard-driven Focus mode: approve, reject, skip, next. Every decision is logged with who made it and why.

04

Agents stop guessing

Approved definitions power a natural-language Ask view and an MCP server that any AI tool can use. Every answer cites the definition it stood on.

Two things keep the dictionary honest over time. If a column’s type or nullability changes after approval, it flips to “needs re-review” instead of staying trusted. And a feature called Semantic Terms groups the same field name across tables and flags when their definitions disagree, so you catch drift in meaning, not just drift in schema.

Semantic Engine review queue: AI-drafted field definitions listed with category and status, a detail panel to approve or reject with a required reason, and bulk actions
The review queue. AI-drafted definitions wait for a human decision, and every approval or rejection needs a reason. Demo data.

A dictionary people can browse, a map agents can follow

Semantic Engine dictionary: a source, dataset, table, and field tree with approval progress per table, and a detail panel showing a field's approved definition, category, and reviewer
The dictionary. Browse source, dataset, table, and field, with governed coverage shown at every level. Each field shows its approved definition and who approved it.

Questions are rarely asked in column names. Someone asks about cost by region, not cost_line_items.actual_amount grouped by projects.region. The ontology layer models the business concepts themselves (what each one means, where it lives in the warehouse, and what it rolls up into) so an agent can translate between the two. The hierarchy is a graph, not a tree: a project rolls up through both its region and its client.

Semantic Engine ontology map: Enterprise rolls up through Client and Region to Project, which contains Submittal and Cost Line Item, each grounded in a table and governed
The ontology map. Business concepts, their warehouse grounding, and how they roll up. Dashed means the concept isn’t stored in any table. Demo data.

Mid-market teams building on AI

The target customer is a mid-market company or startup with a real data warehouse, a growing list of AI initiatives, and no appetite for an enterprise governance rollout. The same governed dictionary serves four needs: accurate LLM data questions, a lightweight data-governance record, a shared vocabulary for analytics and ML, and, later, an API that retrieves data by business concept instead of physical schema.

What I chose, and what I held back

Nothing is served until a person says so

LLMs draft every definition, but only approved ones are ever served. The guardrail is enforced in the database query, not just in the prompt, so a clever question can’t talk its way around it. Every approve and reject requires a reason, and bulk actions state exactly what they skipped and why.

Decline instead of guess

When there is no approved definition for what someone asks about, the engine says so rather than inventing one. Each miss is recorded as a “context gap” that ranks the review queue by real demand, so the fields people actually need get reviewed first.

Drift reopens the review

Physical schema facts are stored separately from human annotations, so a re-scan never overwrites reviewed work. A field whose type or nullability changes goes back in front of a person before it’s served again.

Nothing is silently deleted

Archiving a source or excluding a table hides it from the dictionary but never erases the review history underneath. Bring it back and every past approval, rejection, and reason is still there.

Scope stays in the customer’s hands

Table-level selection means the governed dictionary contains only what the team chose to include, not everything a connected warehouse happens to expose.

Consistency across tables, not just within one

The same field name can mean different things on different tables. Semantic Terms surfaces that automatically, instead of letting each table’s reviewer quietly approve a different answer.

A control plane with no lock-in

Phase 1 describes data and points to it. It doesn’t move or copy it. Sources sit behind a connector interface and LLMs behind a provider interface (Claude and Gemini both work, chosen at deploy time), so neither the warehouse nor the model is a dependency.

Scoped the risky feature, didn’t build it

Letting an AI write and run live SQL is the obvious next step. I specified it in full, including cost caps, read-only access, SQL safety checks, and audit logging, and deliberately held it back until the governed-context layer is proven.

One governed context layer, any AI tool

An MCP server exposes the approved dictionary to whatever agent or IDE a team already uses, with no separate integration per tool and nothing unapproved leaking through. It has five tools:

list_sourcesSee what’s connected and how much of it is governed (approved vs. total fields) before asking about any of it.
get_source_contextPull the full approved shape of a source: every table and field a human has signed off on.
get_table_contextThe same, scoped to one table.
search_fieldsResolve a business term to the field it lives on. Searches approved fields only, so a search never surfaces something no one has reviewed.
report_context_gapWhen an agent can’t find something, it says so. Every miss feeds the review queue’s demand ranking.
Claude Desktop chat using the Semantic Engine tools: it answers what submittal_status means from the approved definition, then, asked about variance_amount, reports there is no approved definition and labels its inference as a guess
An agent in Claude Desktop using the MCP server. The first answer comes from the approved definition. For the unapproved field, it says there is no governed definition, labels its inference as a guess, and offers to log a context gap. Demo data.

A working prototype, deployed

Phase 1 is built and running on Google Cloud behind authenticated access. I wrote the PRD, the technical design for the MCP server, and the sprint plan, and built it against a construction-industry dataset. It covers source connection, AI-assisted definition drafting, the review queue, the dictionary and relationship map, an Ask experience grounded only in approved definitions, and an MCP server that works with Claude Desktop.

Next: a UI for managing relationships between fields, grouping fields into business categories, governed direct querying, and selective materialization for fast, agent-friendly retrieval.

Next.js TypeScript PostgreSQL Model Context Protocol Claude Gemini Google Cloud Run BigQuery