mirror of
https://github.com/qdrant/qdrant.git
synced 2026-10-03 11:27:37 -05:00
* feat(facet): add sampling strategy for high-cardinality fields For approximate facet queries the current per-segment implementation walks every distinct value in the field index — O(unique_values_count) even when the user asks for a tiny top-K. On UUID-style fields with millions of unique values this dominates the request latency even after #9208 capped the cross-shard payload. This commit adds a parallel sampling strategy that runs in O(limit) instead of O(unique_values_count): 1. Phase 1 — iterative novelty sampling. Stream point IDs in random order (filtered if requested), look up each point's value via the facet index, and collect distinct values into a candidate set until `limit * 10` (min 1000) candidates have been gathered. Uses a batch size of 32 to amortise the inner `for_points_values` call, and bails out early after 128 consecutive empty batches when the long tail is too thin to keep finding novel values. 2. Phase 2 — exact-count post-pass. For each candidate value, compose `field == value` with the user filter and count via the payload index. This guarantees the returned counts are exact (matching the semantics of the full-scan path); only the *set* of returned values is approximate. The two strategies live side by side; `SegmentReadView::approximate_facet` picks between them per-request based on `unique_values_count > limit * FACET_FULL_SCAN_FACTOR` (FACTOR = 4). Below that, the existing scan path runs unchanged — it'd visit most of the index either way, and the post-pass adds no value. The Monte-Carlo simulation behind this design (see thread context for Zipf-distributed fields with cardinality up to 10^5 in ~1000 samples, and trivially-correct results on UUID-style fields where every value has count 1. Adds a new `unique_values_count` method on the `FacetIndex` trait (implemented for `MapIndex`, `ReadOnlyMapIndex`, `BoolIndex`, `ReadOnlyBoolIndex`, and the `FacetIndexEnum` dispatcher) so the strategy switch can run without touching the index. Co-authored-by: Cursor <cursoragent@cursor.com> * [AI] simplify, use single file [AI] better selection of filtering approach fmt [AI] simplify, use single file * manual simplification * [AI] implement candidate-based lookups [AI] 🧹 * precollect filter into bitmap * avoid sampling with restrictive filter * fix rebase + clippy * refactor tests * no duplicate values in map index * polish comments --------- Co-authored-by: root <111755117+qdrant-cloud-bot@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>