DATA & INDEXING
What is faceted search?
Filters that come with counts. The user sees what is there before clicking, and the numbers come back in the same request as the results.
- Denim jacket
- Rain jacket
- + 5 more
- Polo302
- Levis144
- Nike85
- Small412
- Medium88
- Large31
- 0–25120
- 25–50245
- 50–100166
- Field facets: counts per value - brands and sizes each total 531
- Range facet: the same 531 records in price bands
- Stat facet: the average price over the filtered set
Facets vs filters
Everyone has used the sidebar. Search for a jacket and the left column reads Brand: Polo (302), Levis (144), Nike (85). The numbers are the facet. A plain filter would list the brands and let you click one; a facet tells you, before you click, how many results each click would leave. That turns a guess into a choice, and it is why the pattern is on every catalog and every listing screen that handles a large set.
The difference matters more than it looks. A filter is a condition. A facet is a condition plus a count for each of its values, and counting is work: it has to be done across every record that matches the current search, for every value of the field, every time the search changes. Where that counting happens decides whether the screen is fast.
Four kinds
Field facets count records per value of a field: brand, size, status, region. In the products example, brand gives Polo 302 · Levis 144 · Nike 85 and size gives Small 412 · Medium 88 · Large 31.
Range facets count records in buckets of a numeric, currency or date field: price 0–25 (120), 25–50 (245), 50–100 (166); orders this week, last week, earlier; invoices by amount band.
Stat facets compute a number over the filtered set rather than a count per value: sum, average, minimum, maximum, median. The average price of the 531 matching products is 42.60. The total value of this month's rejected orders is a stat facet too.
Nested facets put a facet inside a facet. Ask for orders by region and, within each region, by category, and each level carries its own counts and totals: North 395 orders, of which Clothing 212, Footwear 118, Accessories 65. Any kind can nest in any other, to whatever depth the screen needs.
- RegionLevel 14 values · 952 orders
- North395
- Level 2Clothing212
- Footwear118
- Accessories65
- West247 · 3 categories
- East168 · 3 categories
- South142 · 3 categories
- North395
Facets are reporting
Say the same thing in SQL and it is SELECT region, COUNT(*) … GROUP BY region, then another GROUP BY for category within each region, then a SUM for the totals. A facet request is a GROUP BY with totals, answered by the index instead of the database, on the set the user has already narrowed. That is most of what a dashboard widget is: a pivot is a nested facet with two dimensions, a leaderboard is a field facet sorted by its count or sum, a KPI tile is a stat facet with a comparison. The word "aggregation" describes all of them, and once you see facets as aggregation it stops being surprising that a search layer can answer reporting questions.
Consistency
Every facet is computed on the same filtered set as the results. Search for jaket in Clothing and the brands total 531, the sizes total 531 and the price bands total 531, because all three were counted over the same 531 records, in the same pass, at the same moment. There is no second query that ran a half-second later against data that had moved, and no sidebar count that disagrees with the result count. Across the four regions - North 395 · West 247 · East 168 · South 142 - orders total 952 whichever way they are sliced. The numbers agree with each other because they come from one place.
The one-request screen
Put it together and the catalog page is one request. The lead visual is the products request annotated: text jaket, a category filter, facets on brand and size, a price range facet, an average-price stat facet, and typo tolerance of one edit. The response carries the 531 matching records, the field facets for the sidebar, the price bands, the average, and the suggestions. The sidebar, the result list, the total and the summary line are all rendered from one response, 624 ms for the whole screen, which is why the screen appears all at once instead of assembling itself in stages. A report screen is the same request with different facets and no text.
How counting stays cheap
The search step produces the set of matching record ids. For a field facet the index then reads that field's doc values - the field's values stored together in record order - for exactly those ids, and tallies them. That is a pass over one column, touching only the matching records, with no row to assemble and no table to scan. Range facets do the same with bucket boundaries; stat facets accumulate a sum, a minimum or a running median over the same column.
In SQL the equivalent is a row scan. The database has to find the matching rows, read the grouping column out of each, and sort or hash them into groups, once per facet, because each GROUP BY is its own query. The index has the ids already and the column already laid out for this, which is the whole difference, and it is why facets come back in the same response as the results instead of one query each.
Nesting is a tree of buckets. The region facet produces four buckets; inside each, the category facet runs over just that bucket's ids and produces three more; totals are summed up the tree. Each level is the same column pass over a smaller set, so the cost of a nested facet is close to the cost of its levels added together, not multiplied.