AI ON BUSINESS DATA

What the LLM sees

Not your database, not your index, not your collections. The question, the names it needs to map it, and the result it writes up. That is the list.

Leadership Engineering EXPLAINER 4 min read
YOUR INFRASTRUCTUREQUESTIONAI AssistantMAPS · RUNS · WRITES UPORDERSPRODUCTSCOLLECTIONS NEVER CROSSWHAT THE REQUEST NEEDSINTENT · FIELDS · WORDINGYour LLM providerYOUR ACCOUNT · YOUR KEYANSWERQUESTIONYOUR INFRASTRUCTUREAI AssistantMAPS · RUNS · WRITES UPORDERSPRODUCTSCOLLECTIONS NEVER CROSSANSWER↓ WHAT THE REQUEST NEEDS↑ INTENT · FIELDS · WORDINGYour LLM providerYOUR ACCOUNT · YOUR KEY · OUTSIDE
Two rails to the provider. The collections stay inside.

The question behind the question

"What does the model see?" is the first thing a security review asks about an assistant over business data, and it deserves a precise answer rather than a reassuring one. The precise answer is a short list, and this page is that list. It describes the shape of the design - what has to cross to the model for the assistant to work, and what never needs to - and it is the shape to hold any vendor to.

What it sees

The question, as typed. There is no way to understand a question without reading it.

The vocabulary. The names of the collections - orders, products, customers, invoices - and of their fields, with a line describing each: what region means, that status takes one of five values, that order date is a date. This is metadata, not data. It is what lets the model turn "orders by region this month" into a request against the orders collection, faceted by region, filtered to the month, without guessing field names.

The result it is writing up. When the index has answered - North 395 · West 247 · East 168 · South 142 - the rows and counts that make up that answer are given to the model so it can write the summary: North leads order volume by a wide margin. The result is the same thing the user is about to see on screen. The model sees the answer to this question, not the records the answer was counted from.

The conversation so far, so that "now only North" can be understood as a narrowing of the previous question.

What the model sees
  • The question, as typed
  • The names and descriptions of collections and fields, so it can map the question
  • The result it is asked to write up: the rows and counts that answer this question
  • The conversation so far, for follow-ups
What it never sees
  • The production database
  • The index, or any collection as a whole
  • Records the question did not retrieve
  • Access keys, credentials, the schedule
  • Anything from a question another user asked
The question, the vocabulary, the answer. Not the data.

What it never sees

The production database, because the assistant never reads it: it reads an indexed copy. The index, because the model never searches; it asks, and the index searches. Any collection as a whole, because no request returns one. Records the question did not retrieve, because only the result of the request is written up. Credentials, keys, refresh schedules, access lists, because none of them is part of any request. And nothing from a conversation another user had, because each conversation carries only its own context.

ItemSent to the model?Why
The questionYesIt has to understand what is being asked
Collection and field names, with descriptionsYesThe vocabulary, not the dataIt maps the question onto fields that exist
The result for this questionYesRows and counts, as returnedIt writes the summary from what came back
The conversation so farYesFollow-ups depend on it
The whole collectionNoThe index does the searching; the model never scans
The production databaseNoIt is never on the path
Records outside the resultNoOnly what the request retrieved is written up
Keys, schedules, access listsNoNot part of any request

Why the result has to cross

A reasonable question is whether the model could be kept away from the numbers altogether, receiving only the shape of the answer and having the figures filled in afterwards. It could, for a template. It cannot, for a summary: a sentence like North far ahead with 395 orders, followed by West with 247 is written by reading the result, and a model that has not read the result cannot write it. So the honest statement is the one on this page: the retrieved result for the question crosses, the dataset does not. The design choice that keeps that bounded is the size of the result - a facet with four regions is four numbers, not 952 orders - and the fact that the request, not the model, decides what is retrieved.

Under whose account

What crosses, crosses to a provider you chose, under your own account and your own key, and is handled under your agreement with that provider, not under a vendor's. That matters because the provider's terms are the terms that govern what it may retain or train on, and because an account you hold is an account you can cap, audit and close. Bring your own LLM covers the arrangement; Data flow map places this flow beside the three others.

How to verify it

Ask the vendor for the list on this page, in this form: what is sent, what is not, and under whose account. Ask what bounds the size of a result. Ask whether the model can issue a query of its own or only choose from a vocabulary, because a model that can write SQL can retrieve whatever the SQL retrieves. And ask whether the same questions have the same answers for every product in the platform, because they usually do not, and the honest vendor will say so.

See it on real data.

The demo instance runs dashboards, data grids and the AI Assistant on real business data. No sign-up.