AI tools are only as accurate as their understanding of your data. Semantic Engine gives them a human-approved map of what every field means, so agents can answer questions about structured business data safely and accurately.
Also explored under the working name Webster.
Most organizations have data scattered across many systems, with names and definitions that only a few people really understand. Centralizing it into one warehouse is a big project, and building purpose-built datamarts for each reporting need only adds to the fragmentation.
That’s a problem for AI. Pointed at a raw schema, an LLM guesses what status or region means, writes queries that look right and aren’t, and has no way to explain where a number came from. Analysts, ML models, governance teams, and AI agents all need one shared source of truth about what the data is.
I saw this firsthand building an enterprise analytics and AI platform across 16 operating companies. The tools that solve it, like Collibra and Alation, are heavy and expensive. Mid-market teams and startups have nothing lightweight, and nothing built for LLMs from the start.
Point it at a BigQuery project, scoped to specific datasets and tables. It reads the schema in place, and nothing is copied out of the warehouse. A connector model is built for other warehouses.
An LLM reads each table’s schema and a sample of values, then drafts a plain-English definition per column. Every draft lands in a review queue as pending. Nothing is trusted yet. The design target is cutting manual authoring by more than 70%.
Triage in bulk in Scan mode, or go field by field in a keyboard-driven Focus mode: approve, reject, skip, next. Every decision is logged with who made it and why.
Approved definitions power a natural-language Ask view and an MCP server that any AI tool can use. Every answer cites the definition it stood on.
Two things keep the dictionary honest over time. If a column’s type or nullability changes after approval, it flips to “needs re-review” instead of staying trusted. And a feature called Semantic Terms groups the same field name across tables and flags when their definitions disagree, so you catch drift in meaning, not just drift in schema.
Questions are rarely asked in column names. Someone asks about cost by region, not cost_line_items.actual_amount grouped by projects.region. The ontology layer models the business concepts themselves (what each one means, where it lives in the warehouse, and what it rolls up into) so an agent can translate between the two. The hierarchy is a graph, not a tree: a project rolls up through both its region and its client.
The target customer is a mid-market company or startup with a real data warehouse, a growing list of AI initiatives, and no appetite for an enterprise governance rollout. The same governed dictionary serves four needs: accurate LLM data questions, a lightweight data-governance record, a shared vocabulary for analytics and ML, and, later, an API that retrieves data by business concept instead of physical schema.
An MCP server exposes the approved dictionary to whatever agent or IDE a team already uses, with no separate integration per tool and nothing unapproved leaking through. It has five tools:
list_sourcesSee what’s connected and how much of it is governed (approved vs. total fields) before asking about any of it.get_source_contextPull the full approved shape of a source: every table and field a human has signed off on.get_table_contextThe same, scoped to one table.search_fieldsResolve a business term to the field it lives on. Searches approved fields only, so a search never surfaces something no one has reviewed.report_context_gapWhen an agent can’t find something, it says so. Every miss feeds the review queue’s demand ranking.
Phase 1 is built and running on Google Cloud behind authenticated access. I wrote the PRD, the technical design for the MCP server, and the sprint plan, and built it against a construction-industry dataset. It covers source connection, AI-assisted definition drafting, the review queue, the dictionary and relationship map, an Ask experience grounded only in approved definitions, and an MCP server that works with Claude Desktop.
Next: a UI for managing relationships between fields, grouping fields into business categories, governed direct querying, and selective materialization for fast, agent-friendly retrieval.