AI ON BUSINESS DATA

How can AI answer questions about business data?

The model does not read your database. It turns your question into a structured query, runs it against indexed data, and writes up the result. The numbers come from the index; the words come from the model.

Anyone PILLAR 7 min read
Your LLM provideryour account · your key · outside
↑ the information required for the request↓ intent · field choice · wording
Your infrastructure · AI assistant
  1. UnderstandWhat is being asked, and about what."Orders by region, this month"
  2. MapPick the collection, fields, filters and facets. No SQL.collection orders · facet region · period this month
  3. RunThe index searches and counts. Exact numbers, production untouched.North 395 · West 247 · East 168 · South 142
  4. ShapeThe result picks its format.Leaderboard
  5. SummarizeThe model writes it up and suggests what to ask next."North leads order volume by a wide margin"
Indexed collectionsorders · products · customers · invoices
The numbers come from the index. The words come from the model.

The question people actually have

"Does it read all my data?" No. The model never sees the orders table. What it reads is the question, a description of what data exists - the names of collections and their fields, with a note on what each means - and, once the index has answered, the result it has to write up. It does not read every record, because it does not need to and because, as the rest of this article explains, that is not how a correct number is produced.

The division of labour is the whole design. The model is good with language: understanding what was asked, choosing the fields that answer it, writing a clear sentence about the result. The index is good with numbers: finding, filtering, counting and totalling exactly. Each does its part, and neither does the other's.

Five steps

Take one question and carry it through. Someone types: orders by region, this month.

1. Understand. The model reads the question and works out what is being asked: a count, broken down by one dimension, over one period. It also resolves the things people leave implicit - "this month" is a date range, "orders" is a collection, "by region" is a grouping.

2. Map. The question becomes a structured request against a vocabulary the assistant was given: the orders collection, a filter on order date for the current month, a facet on the region field. No SQL is written. The model chooses from collections, fields, filters and facets that exist, and nothing else is available to it.

3. Run. The request goes to the index, which does the searching and counting: North 395 · West 247 · East 168 · South 142. These numbers are computed by the index over the indexed records, exactly, the same way a dashboard widget would get them. The model has not touched them.

4. Shape. The result picks its format. Four regions ranked by a count is a leaderboard. Had the question been "orders this month" it would be a KPI tile; "by region and month" a pivot; "show me the urgent ones" a table.

5. Summarize and suggest. The model writes the sentence - North leads order volume by a wide margin - names the rest, and proposes what to ask next: which stores in North declined, or the orders behind the number. The words come from the model; every number in them came from step 3.

What the model is good at, and what it must not do

Good at: naming the right fields from a loosely worded question, resolving "this month" and "top" and "compared with last year", choosing a format, and writing a summary a colleague would write. That is language and intent, and it is what large models are for.

Must not: compute. A model asked to add four numbers will usually get it right, and sometimes will not, and there is no way to tell which from the outside. So it never adds, averages or compares; the index does, and the model reports. Must not: fill a gap with a plausible value. If the question cannot be mapped to the data that exists, the right answer is to say so, or to ask, not to produce a number that looks like the others. The guardrail is structural: the model can only choose from the vocabulary it was given, and the numbers only ever come from the index, so an invented value has nowhere to come from.

Why an index rather than the database

Three reasons, and the production database is the third. Structured. An index organizes records by field, with a known schema per collection, which is exactly the kind of vocabulary a model can be constrained to. Fast. Counts and totals over a filtered set are facets, computed beside the results in one request, so "orders by region" is one call in milliseconds, and the follow-up is another. Facets are aggregation for free. Safe. The index is a copy. Nothing the assistant does can change a record, and nothing it asks adds load to the system that is taking orders. The production database is never on the path, which also means no model ever holds a database credential.

How the answer gets its shape

ORDERS952vs last periodURGENT14vs last periodREJECTED36vs last period
KPI · one number and its change
REGIONJANFEBMARTOTAL North18412289395 West5082115247 East415176168
Pivot · two dimensions and totals
01020304 NorthWestEastSouth 395247168142
Leaderboard · ranked
+ HEATMAP, FLOW
Chart · trend, comparison, share
ORDERREGIONSTATUS #48213NorthUrgent #48214WestRejected #48215NorthShipped
Table · the records behind the number
Five shapes. The question decides which one the answer takes.

Five shapes, chosen by the question. A KPI for one number with its change: 952 orders, 14 urgent, 36 rejected, each against the previous period. A pivot splits North's 395 into 184, 122 and 89 by month, West's 247 into 50, 82 and 115, East's 168 into 41, 51 and 76, with the totals beside them. A leaderboard for a ranking. A chart for a trend, a comparison or a share. A table for the records themselves. Around the shape, when the findings call for it: a summary in plain language, a short list of what you could do about it, and the next questions worth asking, each one a click away. The format rules are the same ones a person would use, and Choosing the right chart for business data sets them out.

Follow-ups

The suggested next questions are derived from what came back, not from a script. North led, so which stores in North declined the most? is a reasonable next step; rejected orders appeared in the result, so show the orders behind this is another. Each suggestion is a question the assistant can actually answer against the same collections, so clicking one runs the five steps again with the context carried forward. A conversation about data is a chain of these, and the chain is what a dashboard cannot offer.

Bring your own LLM

The model is not part of the product. You choose the provider, buy usage directly from them, and configure the key in the assistant. The provider account, the usage cost and the model choice stay with you, and no LLM cost is bundled into the assistant's pricing. If a better or cheaper model appears next year, the key changes and nothing else does.

What crosses to the provider

The honest statement is narrower than "nothing" and wider than "the question". The assistant runs inside your infrastructure, with the indexed collections beside it. The LLM provider is outside, under your account and key. Between them run two rails. Outward: only the information required for the request - the question, and what the model needs in order to map and then write up the answer; the assistant can send relevant metadata rather than exposing the full dataset. Inward: the model's intent, its field choices and its wording. The collections themselves never cross, and the index is never queried by the provider.

YOUR INFRASTRUCTUREQUESTIONAI AssistantMAPS · RUNS · WRITES UPORDERSPRODUCTSCOLLECTIONS NEVER CROSSWHAT THE REQUEST NEEDSINTENT · FIELDS · WORDINGYour LLM providerYOUR ACCOUNT · YOUR KEYANSWERQUESTIONYOUR INFRASTRUCTUREAI AssistantMAPS · RUNS · WRITES UPORDERSPRODUCTSCOLLECTIONS NEVER CROSSANSWER↓ WHAT THE REQUEST NEEDS↑ INTENT · FIELDS · WORDINGYour LLM providerYOUR ACCOUNT · YOUR KEY · OUTSIDE
Two rails to the provider. The collections stay inside.

What exactly is sent per request depends on the question and on how the assistant is configured, and your provider's own terms govern what happens to it on their side. The per-product picture - what leaves your environment for Engine, Insights, the AI Assistant and the MCP Server - is in Data flow map: what leaves your environment, per product.

Under the hood

Intent, plan, facets, format, summary

Internally the five steps are a short pipeline. An intent stage turns the question into a structured description: measure, dimensions, filters, period, and the shape being asked for. A query plan stage maps that description onto one collection and its fields, producing a request the index understands - filters, facets, sorting, paging - rather than free text or SQL. The index runs it and returns records and facets, which is where every number comes from. A formatter picks KPI, pivot, leaderboard, chart or table from the plan's shape. A summarizer writes the sentence, the suggested actions and the next questions from the returned result, with the numbers passed through as-is.

The guardrails sit at the plan stage. The model's choices are drawn from a constrained vocabulary - the collections and fields that exist, with their types - and a plan that references something outside it is rejected before anything runs. There is no free SQL to inject into and no credential to misuse, because the only thing the plan can become is an index request. That constraint is what makes the numbers exact and the production database untouchable, and it is why Three ways AI answers from your data treats indexed retrieval as a different approach from text-to-SQL rather than a variant of it.

See it on real data.

The demo instance runs dashboards, data grids and the AI Assistant on real business data. No sign-up.