* override `for_values_map` for `OnDiskMapIndex`
* refactor and apply to filter as id iterator
* add comment
---------
Co-authored-by: generall <andrey@vasnetsov.com>
* feat: feature-flagged segment manifest for LocalShard
Add an on-disk segment manifest (`segments/manifest.json`) that lists a
shard's segments and their state, so out-of-process readers (e.g. a
read-only follower, possibly over object storage) can discover segments
without scanning the filesystem. Gated by the new `write_segment_manifest`
feature flag (off by default).
shard: define the structure + helpers (`SegmentsManifest`,
`SegmentManifestState`, `from_segment_holder`) in a new `segment_manifest`
module, plus the `SEGMENT_MANIFEST_FILE` constant and path helper. The
manifest is a flat `{ "<uuid>": "<state>" }` map; only `active` is written
today, with `under_construction`/`retiring` defined so the format can be
extended without breaking compatibility.
collection: LocalShard owns the writing logic. The manifest is persisted
via `SaveOnDisk<SegmentsManifest>`, initialized from the live segment set
on load/build and refreshed by the optimization worker whenever the
segment set changes (the helper re-derives from the holder and no-ops when
unchanged). No changes to `lib/shard/src/optimize.rs` internals.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Emj8TFxdtrf3K32eWhGgor
* feat: make append_only_mutations a proper feature flag
Replace the debug-only `QDRANT_APPEND_ONLY_MUTATIONS=1` env-var escape
hatch in segment construction with a `FeatureFlags::append_only_mutations`
flag, so it works in release builds and is configurable like the other
flags (config / `QDRANT__FEATURE_FLAGS__APPEND_ONLY_MUTATIONS`).
Deliberately left out of `FeatureFlags::all()`: it changes mutation
semantics and `all` is enabled in dev and e2e configs, so it stays
explicit opt-in.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Emj8TFxdtrf3K32eWhGgor
* upd openapi schema
* feat: register segments in manifest via must-use token, holder builder
Wire segment-manifest maintenance through the segment lifecycle so a
newly created segment is registered as soon as it exists on disk, and
construction can't silently skip it:
- build_segment now returns a #[must_use] NewSegmentToken carrying the
new segment UUID; the lint forces callers to register or drop it.
- SegmentHolder owns the manifest and reconciles it on sync; new
segments are registered ASAP (even before being added to the holder)
via the token, before they can receive writes.
- SegmentHolderBuilder is the only way to obtain a shard's holder; its
build() wires up the manifest, so it can't be forgotten. init/set
manifest helpers are now private / test-only.
- Optimization registers the optimized segment before dropping the
superseded segments' data; deletion is intentionally lenient.
- Document the consistency assumptions on SegmentsManifest: it is a
superset-biased view that may list not-yet-finalized or already-deleted
segments, which readers must tolerate.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add two-phase config reload to `ReadOnlySegment`, alongside the existing
live-reload:
- `config_reload_diff(&self, fs)` re-reads the on-disk config, diffs it
against the in-memory config, and eagerly loads every new or changed
component (dense/sparse vector storages, indexes, quantized vectors and
payload field indexes) under a shared `&self` borrow, so the segment
keeps serving reads while loading.
- `apply_config_reload(&mut self, diff)` installs the pre-loaded
components and drops removed ones under `&mut self` — a cheap swap with
no I/O, so the exclusive borrow is held only briefly.
A vector or field whose config changed is reloaded (drop + load). The
payload-index half lives on `ReadOnlyStructPayloadIndex` with the same
diff/apply split, plus register/unregister helpers for its `has_vector`
storage map.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`version()` is the segment's update version — only meaningful for the
mutable/storage path. Every real caller already reaches it through
`StorageSegmentEntry` (flush), `SegmentEntry` (update), or a concrete
`Segment`/`ProxySegment`; none use it through the read-only base trait.
Move the declaration down to `StorageSegmentEntry` and relocate the
`Segment`/`ProxySegment` impls accordingly. `ReadOnlySegment` no longer
needs it, so drop the method and the write-only `version`/`initial_version`
fields it carried (`live_reload` never refreshed `version`, and neither
field was ever read).
This leaves `ReadSegmentEntry` a clean read-only surface and stops a
read-only segment from exposing a meaningless (stale) version.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: immutable field index live_reload
* feat: read-only segment live_reload
* fix: linter
* fix: linter
* fix: review comments
* fix: linter
* feat: ReadOnlyVectorData::live_reload covering all components
Extract the per-vector reload into a dedicated method that destructures
`ReadOnlyVectorData` so storage, index and quantized vectors are all
covered. Adding a field without reloading it won't compile, which guards
against silently skipping a component. The segment orchestrator now just
delegates per named vector.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: keep id-tracker delta pending until all reloads succeed
ReadOnlySegment::live_reload drained the id-tracker delta (which advances
tracker state and cannot be replayed) and only then ran the fallible
payload/vector reloads. On error the delta was lost, so un-updated
components drifted out of sync permanently.
Accumulate the delta into a new `pending_reload` field and clear it only
once every component has reloaded successfully. On a later reload the
tracker's fresh delta is folded in via `LiveReloadResult::merge` and the
union is replayed, so a partial failure self-heals. Offsets are monotonic,
so the only merge conflict is an inserted-but-unapplied offset later
deleted: it is dropped from `inserted` and kept in `deleted` so a
partially-applied component drops it on replay.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(facet): add sampling strategy for high-cardinality fields
For approximate facet queries the current per-segment implementation
walks every distinct value in the field index — O(unique_values_count)
even when the user asks for a tiny top-K. On UUID-style fields with
millions of unique values this dominates the request latency even
after #9208 capped the cross-shard payload.
This commit adds a parallel sampling strategy that runs in O(limit)
instead of O(unique_values_count):
1. Phase 1 — iterative novelty sampling. Stream point IDs in random
order (filtered if requested), look up each point's value via the
facet index, and collect distinct values into a candidate set until
`limit * 10` (min 1000) candidates have been gathered. Uses a
batch size of 32 to amortise the inner `for_points_values` call,
and bails out early after 128 consecutive empty batches when the
long tail is too thin to keep finding novel values.
2. Phase 2 — exact-count post-pass. For each candidate value, compose
`field == value` with the user filter and count via the payload
index. This guarantees the returned counts are exact (matching the
semantics of the full-scan path); only the *set* of returned values
is approximate.
The two strategies live side by side; `SegmentReadView::approximate_facet`
picks between them per-request based on
`unique_values_count > limit * FACET_FULL_SCAN_FACTOR` (FACTOR = 4).
Below that, the existing scan path runs unchanged — it'd visit most of
the index either way, and the post-pass adds no value.
The Monte-Carlo simulation behind this design (see thread context for
Zipf-distributed fields with cardinality up to 10^5 in ~1000 samples,
and trivially-correct results on UUID-style fields where every value
has count 1.
Adds a new `unique_values_count` method on the `FacetIndex` trait
(implemented for `MapIndex`, `ReadOnlyMapIndex`, `BoolIndex`,
`ReadOnlyBoolIndex`, and the `FacetIndexEnum` dispatcher) so the
strategy switch can run without touching the index.
Co-authored-by: Cursor <cursoragent@cursor.com>
* [AI] simplify, use single file
[AI] better selection of filtering approach
fmt
[AI] simplify, use single file
* manual simplification
* [AI] implement candidate-based lookups
[AI] 🧹
* precollect filter into bitmap
* avoid sampling with restrictive filter
* fix rebase + clippy
* refactor tests
* no duplicate values in map index
* polish comments
---------
Co-authored-by: root <111755117+qdrant-cloud-bot@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
`GridstoreReader`)
- Make the entrypoint of payload storage choose the `Populate` variant
- Mutable payload indexes now populate gridstore blockingly before using
it to load
- Propagate `Populate` into `flags` module
- Add `UniversalRead::populate_auto` to know whether a backend chooses
to populate or not when `Populate::Auto`
* feat(bm25): add explicit Disabled stemmer; deprecate language hack
Adds a `Disabled` variant to `StemmingAlgorithm` (`stemmer: {"type": "none"}`)
so stemming can be turned off explicitly in both the main engine and Edge,
instead of relying on the undocumented `language: "none"` footgun that
silently disabled both stemming and stopwords.
For language-neutral text processing the supported setup is now:
1. set the stemmer to disabled, and
2. configure an empty stopword set.
The main engine still tolerates unsupported languages (so existing
`language: "none"` configs keep working on upgrade) but now logs a
deprecation warning pointing users to the explicit setup. Edge continues
to reject unsupported languages, and now has a real way to disable stemming.
Refs: https://github.com/qdrant/qdrant/issues/9289
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(edge-py): handle Disabled stemmer in python bindings; fix openapi schema
- Handle the new StemmingAlgorithm::Disabled variant in the qdrant-edge-py
bindings (FromPyObject/IntoPyObject/Repr) and add a DisabledStemmer pyclass
plus its .pyi stub entry.
- Match generator output for the StemmingAlgorithm OpenAPI schema (plain $ref
in anyOf) so docs/redoc/master/openapi.json stays consistent.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(openapi): regenerate StemmingAlgorithm schema with generator output
Ran tools/generate_openapi_models.sh so docs/redoc/master/openapi.json
exactly matches generator output: DisabledStemmerParams/NoStemmer are placed
after SnowballLanguage, and the StemmingAlgorithm anyOf entry is a plain $ref
(the schema2openapi step flattens the allOf+description wrapper).
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(test): avoid wildcard enum match arm in bm25 sparse_len helper
clippy --all-targets flags `other => panic!()` as wildcard_enum_match_arm;
match the Dense/MultiDense variants explicitly instead.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: issues
* fix: log::warn as call once
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Daniel Boros <dancixx@gmail.com>
* Buffer multi-dense offsets store to prevent reload corruption
The appendable multi-dense storage flushes its `vectors` and `offsets`
chunked stores independently. Each flusher snapshots its `status.len` at
creation time but msyncs chunk bytes at execution time. A re-upsert that
grows a point past its reserved capacity rewrites the point's `offsets`
entry in place to a freshly-appended row region; if that lands in a
flush's creation-to-execution window, the relocated entry becomes durable
while the `vectors` store's recorded length still predates the rows it
references. On reload that point is unreadable and a WAL append reuses the
rows, clobbering another point.
Wrap the offsets store in a write-back buffer (`BufferedOffsets`) so the
durable offsets can never reference rows beyond the durable `vectors`
length. Offset writes stage in a pending overlay and only land in the
durable store while a flush executes; the flusher snapshots the pending
set at creation time, so any write after that stays buffered for the next
flush. Both flushers snapshot at the same instant, yielding a consistent
durable cut. Rows written after the cut are unreferenced garbage the next
append overwrites. This prevents the skew at the source instead of
patching it on reload, and also closes the residual offsets-length smear.
The buffer follows the Gridstore flusher convention: pending writes live
inside the single lock-guarded store, reads consult the overlay then the
durable bytes under one lock, and the durable msync runs after releasing
the write lock so it never stalls concurrent reads.
Enable the "m" multivector in the collection model test now that its
reload divergence is resolved.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Narrow offsets flush write-lock scope
Build and sort the pending snapshot before taking the write lock; only the
apply + reconcile need it. Shortens the lock hold so concurrent reads block
less. Addresses review feedback.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* perf: use Entry API to avoid redundant map double-lookups
Replace get_mut/contains_key followed by insert with the entry API
across several maps, collapsing two hash lookups into one. Limited to
sites where the key is Copy or already owned and moved, so no extra
key clone is added to any hot path.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* More Entry API usage in mutable_geo_index
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: xzfc <xzfcpw@gmail.com>
* Add DynConditionChecker type alias
* Add DynConditionChecker type alias (point_scorer.rs)
* Move trait ConditionChecker to `common`
Reason: will be used both in `segment` and `sparse` crates.
The 6 turbo vector storage model tests are flagged SLOW by nextest on
Windows CI (>60s, some >240s):
- vector_storage::turbo::tests::turbo_model_test_random_ops_{dot,cosine}
- vector_storage::turbo::multi::tests::turbo_multi_model_test_random_ops_{dot,cosine}
- vector_storage::turbo::test::congruent_random_ops_{dot,cosine}
These are dim x seed x ops sweeps; Windows runners are several times slower.
Lower SEEDS_PER_CELL on Windows only via #[cfg(windows)] so the runs stay
within the slow-timeout, while keeping full coverage on Linux/macOS.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat: add open for read-only sparse vector index + enum sparse dispatcher
* fix: universal-IO loads for read-only sparse index open
* refactor: rename load_via/open_via to load_universal/open_universal
* refactor: drop StorageVersion::load in favor of load_universal
The regular-IO `load` duplicated `load_universal` over plain `File` IO.
Remove it and route all callers through `load_universal(&MmapFs, ..)`.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(sparse): decouple inverted index from concrete storage; split segment_constructor_base (#9461)
Make the read-only index enum generic over storage `S` (no concrete MmapFile),
and remove construction callbacks from the sparse index open paths.
- InvertedIndex is now a pure read/search trait: `open`/`from_ram_index`/
`type Fs` moved off the trait to inherent methods on each concrete index
type, so construction no longer requires `S::Fs: Default`.
- SparseVectorIndex open split into a generic `plan` (load-vs-build decision +
RAM-index build) and generic `finish` (assembly); callers do the concrete
per-type construction, so no construction callbacks are needed.
- ReadOnlySparseVectorIndex::open takes the already-constructed inverted index
and caller-loaded config instead of a `load_inverted_index` callback.
- VectorIndexReadEnum is generic over `S: UniversalRead`; sparse mmap variants
hold `InvertedIndexCompressedMmap<_, S>` rather than a concrete `MmapFile`.
- Split the 1182-line segment_constructor_base.rs into a module (paths,
vector_storage, payload_storage, id_tracker, vector_index,
sparse_vector_index, create_segment, segment, legacy_state); the sparse
dispatcher's match arms collapse into three per-family helpers.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat: LiveReload (no-op) dispatch for read-only vector index enum (#9436)
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: live_reload for read-only immutable id tracker + enum dispatch
* deleted should be already sorted
---------
Co-authored-by: Luis Cossío <luis.cossio@outlook.com>
* feat: add open to read-only HNSW index
* fix: dedicated universal-IO load for read-only graph
* fix: report actual graph residency in read-only is_on_disk
* refactor: rename load_via to load_universal
* refactor: unify graph links loading on universal IO
Replace the dual GraphLinks load paths (regular mmap + universal IO) with a
single universal-IO loader, and the `Mmap` GraphLinksEnum variant with a
`Universal` variant that keeps a type-erased `UniversalRead` handle alive
behind `Box<dyn GraphLinksStorage>` (UniversalRead is not object-safe).
- GraphLinksEnum::from_storage picks the variant from UniversalKind:
borrowable (mmap-backed) kinds stay `Universal`, others are materialized
into `Ram`, so the borrowability invariant in GraphLinksStorage::bytes
holds by construction.
- Loading takes a `Populate` parameter, derived from `hnsw_config.on_disk`:
on_disk -> Populate::No (lazy), otherwise Populate::Blocking.
- is_on_disk is reported from config again instead of the enum variant.
- Split graph_links.rs into a graph_links/ module (format, vectors, storage,
links, tests) with a relationship diagram in mod.rs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor: drop LoadOption in favor of a Populate parameter
After unifying the link load paths on universal IO, LoadOption only encoded a
populate choice with a fs that was always MmapFs. Replace it with a `Populate`
argument to GraphLayers::load, removing the enum, its constructors, the
load_links helper, and the unused generic backend.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A DropIndex (or vector name deletion) landing between flusher capture
and execution drops the component's storage, making its captured
flusher return Cancelled. Aborting the rest of the flush sequence at
that point leaves the components flushed so far durably ahead of
payload storage and point versions. WAL replay then re-derives
filter-based operations through the too-new field index and silently
skips points whose payload still needs the operation re-applied,
losing the update.
Skip the dropped storage and keep flushing instead. The drop itself is
a versioned operation that WAL replay re-applies, so skipping it is
safe.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
`delete_vector_name` removes a vector's storage/index directories, but that
removal is best-effort: it logs a warning and continues on failure (and a crash
mid-delete leaves them too), so the directories can still be on disk after a
delete.
When the same name is later recreated, `create_vector_name` reopened whatever
was on disk: `open_vector_storage` reuses existing data and
`prefill_deleted_entries` only pads missing entries, so a point that was never
re-upserted silently regained its old vector after reload.
Remove any leftover storage/index directories in `create_vector_name_impl`
before opening, so a recreated name always starts empty regardless of whether
the earlier delete fully removed them.
Adds deterministic regression tests (dense + sparse) that reproduce the
resurrection by restoring the storage directory after the delete, standing in
for leftover files the delete did not remove.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat: ReadOnlySegment read view + ReadSegmentEntry impl
Add the read_only segment module: ReadOnlySegment / ReadOnlyVectorData,
a `with_view` builder mirroring `Segment::with_view` (with a
`ReadOnlySegmentReadViewFor` alias and a `VectorDataRead` impl for
`ReadOnlyVectorData`), and a `ReadSegmentEntry` impl that delegates to the
shared `SegmentReadView`. A read-only segment is never appendable, so
`is_appendable()` is `false` and the view builders receive `false`.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Merge pull request #9400
* fix: pass multi_vector_config in multi read-only live_reload test
* feat: make VectorIndexReadEnum usable for ReadSegment
* feat: use ReadOnlyStructPayloadIndex in ReadOnlySegment
Replace the placeholder in-memory StructPayloadIndex with the
storage-generic ReadOnlyStructPayloadIndex<S>, mirroring how the
read-only HNSW index already wires its payload index. The segment
read view now nests the read-only payload-index view generics
(ReadOnlyPayloadStorage / ReadOnlyIdTrackerEnum / VectorStorageReadEnum
/ ReadOnlyFieldIndex).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Daniel Boros <56868953+dancixx@users.noreply.github.com>
A plain upsert could fail with a spurious "No point with id ... found"
(HTTP 404) when it raced a `prevent_unoptimized` optimization. This
surfaced as flaky CI failures of
`test_shard_transfer_includes_deferred_points[snapshot]`.
Root cause: `apply_points_with_conditional_move` reads the source point's
vectors and payload before relocating it into an appendable segment. Those
reads used the default `DeferredBehavior::VisibleOnly` accessors
(`all_vectors`/`payload`). When the source point is deferred — invisible to
ordinary reads, e.g. a point whose internal id is beyond the deferred
threshold under `prevent_unoptimized` while an optimization wraps its
segment in a proxy — `VisibleOnly` cannot resolve it: `payload`/`vector`
raise `PointIdError`, which propagates out as the user-facing 404 on a plain
upsert that internally takes the copy-on-write move path.
Fix: add deferred-aware read accessors (`vector_with_behavior`,
`all_vectors_with_behavior`, `payload_with_behavior`) to `ReadSegmentEntry`,
implemented on both `Segment` and `ProxySegment` (the proxy forwards the
behavior to its wrapped segment). The existing accessors delegate with
`VisibleOnly`, preserving current behavior. The CoW move path now reads with
`WithDeferred`, so deferred points are relocated with their real data
instead of failing.
Adds a deterministic regression test asserting `VisibleOnly` hides a
deferred point while `WithDeferred` resolves its real vectors/payload.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat: live_reload for read-only quantized vectors
* Propagate live_reload params through quantized chunked storage
Thread fs, deleted_points, new_points and hw_counter through the full
quantized live_reload chain instead of synthesizing empty deltas and a
disposable hardware counter at the leaf storages.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: use LiveReload trait in chunked mmap
* fix: use LiveReload trait
* fix: compiler errors
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: live_reload for read-only sparse vector storage
* fix: pass multi_vector_config in multi read-only live_reload test
* feat: live_reload dispatch for VectorStorageReadEnum