mirror of
https://github.com/qdrant/qdrant.git
synced 2026-10-03 03:17:43 -05:00
0e37dfd29036fc9c00a38cf604621bdf695d3971
2019
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0e37dfd290 |
ci(windows): skip IO-heavy tests that aren't OS-specific (#9188)
* ci(windows): skip IO-heavy tests that aren't OS-specific On the Windows CI runner, several tests are 3-25x slower than on Ubuntu purely due to slow filesystem IO. These tests exercise platform-agnostic logic (optimizer, snapshot, WAL recovery, dedup, deferred points) and are fully covered by the Linux and macOS jobs. Mark them with `#[cfg_attr(target_os = "windows", ignore = "...")]` so: - Windows CI skips them and finishes faster. - They're still listed and runnable via `cargo test -- --ignored` on Windows for local debugging. Based on JUnit timings from CI run 26462785436 (PR #8827), this should save ~5 minutes wall-clock on the Windows job, taking it closer to the ~13min Ubuntu and ~9min macOS jobs (currently 20m24s). Tests affected: - lib/wal: check_wal, check_last_index, check_clear, check_reopen, check_truncate, check_prefix_truncate, test_prefix_truncate_parametric - lib/edge/optimize: full tests module - lib/segment deferred-point tests: read_operations, dense_segment_combinations, sparse, facets - lib/collection: snapshot_test, points_dedup, wal_recovery, collection_test::test_ordered_read_api, snapshot_recovery_test Co-authored-by: Cursor <cursoragent@cursor.com> * revert(ci/windows): keep WAL and WAL-recovery tests on Windows Reviewer correctly pointed out that WAL is mmap-backed and has substantial Windows-specific code paths: - Different segment allocation (fs4 vs rustix::ftruncate) - Windows-specific delete_windows() with mmap-drop + retry loop - Windows-specific sync_all() because directory fsync is unavailable - Windows-specific lock proxy file (directories aren't lockable) So those tests genuinely need Windows coverage. Reverted skips for: - lib/wal/src/lib.rs: all check_* tests and test_prefix_truncate_parametric - lib/collection/src/tests/wal_recovery_test.rs: all three tests Still skipped on Windows (no OS-specific code in their production paths): - lib/edge/optimize.rs (no cfg(windows) in source) - lib/segment deferred-point tests (segment/ has no cfg(windows)) - lib/collection snapshot/dedup tests (collection/ has no cfg(windows)) - lib/collection integration snapshot_recovery + ordered_read_api Co-authored-by: Cursor <cursoragent@cursor.com> * revert(ci/windows): keep collection integration and snapshot_test Per reviewer request, keep running these on Windows: - lib/collection/tests/integration/* (snapshot_recovery_test, collection_test::test_ordered_read_api) - lib/collection/src/tests/snapshot_test.rs These exercise higher-level collection/snapshot behavior that benefits from cross-platform validation. Remaining Windows skips (production code has no cfg(windows) branches): - lib/edge/src/optimize.rs: 14 optimizer tests - lib/segment/src/segment/tests/mod.rs: 4 deferred-point tests - lib/collection/src/tests/points_dedup.rs: 2 dedup tests Co-authored-by: Cursor <cursoragent@cursor.com> * ci(windows): also skip HNSW/quantization integration tests Per reviewer, also skip these segment integration test modules on Windows: - hnsw_quantized_search_test::* (25 tests) - multivector_filtrable_hnsw_test::* (rstest cases) - multivector_quantization_test::* (rstest cases) - byte_storage_quantization_test::* (rstest cases) - payload_index_test::test_struct_payload_index_nested_fields These exercise pure HNSW/quantization correctness on top of standard segment IO that is already covered by tests we keep running on Windows. Adds ~930s of sequential time to the Windows skip list, bringing the expected wall-clock saving from ~3 min to ~8-10 min. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
24823310d8 |
[UIO] Use UniversalRead::reopen (#9127)
* use universal reopen * remove unused fs args |
||
|
|
bd8a6ae655 |
feat: sync->async bridge (#9093)
* feat: add io_bridge * fix: cached dispatcher * feat: add open_with_handle * fix: naming * fix: wording * fix: rebase io_bridge onto split BorrowedReadPipeline/OwnedReadPipeline * fix: linter * feat: add S3 backend * chore: remove io_bridge * fix: linter * fix: linter * fix: read handle * chore: simplify S3Source * fix: s3 test * feat: support multi runtime * fix: clippy errors * fix: review comments * feat: add io design * feat: add S3 backend * chore: fix docs * fix: dev changes * chore: add some docs * chore: remove explicit type * feat: add new methods * [WIP] review refactor * fmt * fix: bytes alignment * fix: linter * feat: remove Bytes * fix: ci/cd * fix: tests * fix: tests * refactor: simplify io_bridge pipeline to Handle-based dispatch Replace the BridgeRuntime worker thread + request channel + boxed BridgeRequest with direct tokio Handle usage: - BridgeRuntime is now just an Arc<Runtime>; schedule() spawns the read future via Handle::spawn instead of routing it through a dispatcher thread. Removes BridgeRequest and the now-unreachable S3RuntimeShutDown error variant. - Guard against a panicking read task hanging wait() forever: the spawned task catches unwinds and converts them into a TaskPanicked error reply, so every scheduled slot is always answered. - Encapsulate slot bookkeeping in PendingSlots, exposing only the needed operations instead of a public map + counter. - Split the grown pipeline.rs into a pipeline/ module (slots / inner / borrowed / owned), de-duplicating the shared read-future construction. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: move pipeline buffer ownership into the read future Instead of the pipeline owning the destination Vec<T> in a slot map and the future writing through a SendBytePtr raw pointer, let the future allocate the buffer itself and return it through BridgeResponse. The buffer crosses the worker-thread boundary as a normal move via the reply channel, wrapped in a SendableVec<T> newtype that asserts Send for T: bytemuck::Pod only. This removes the entire unsafe SendBytePtr apparatus from the pipeline: no raw pointer, no unsafe fn, no per-call-site unsafe blocks, no heap-stability invariants. The only remaining unsafe in the crate is one bounded `unsafe impl<T: Pod> Send for SendableVec<T>` with a trivially true invariant (Pod types are plain bytes). PendingSlots collapses back to PendingSlots<U>: slots no longer carry buffers. AlignedBufWriter::from_raw_bytes (used only by the SendBytePtr path) and its test are removed. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: drop SendableVec wrapper now that Item: Send UniversalRead's element type is now bound to `Item` (`Pod + Send`), so the io_bridge pipeline no longer needs a hand-rolled `Send` wrapper around `Vec<T>` to ship buffers through the reply channel. Replace `SendableVec<T>` with `Vec<T>` end-to-end and tighten the local impls to `T: Item`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat: implement UniversalRead::reopen for BlobFile BlobFile has no cached file metadata or mapping — `len()` queries the object store fresh on each call — so reopen is a no-op, matching the io_uring impl. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat: add new fs impl for Blob * fix: is_in_ram_or_mmap for S3 --------- Co-authored-by: generall <andrey@vasnetsov.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
4b0d4ab76c |
[UIO] Split UniversalReadFileOps into filesystem + file traits (#9151)
* [UIO] Split UniversalReadFileOps into filesystem + file traits
`UniversalReadFileOps` is now an instance-based trait describing a
filesystem handle (list/exists with `&self`, plus `from_context`). A new
`UniversalReadFs: UniversalReadFileOps` subtrait adds the
`open(&self, path, options) -> Self::File` capability with `type File:
UniversalRead`. `UniversalRead` no longer extends `UniversalReadFileOps`
and is purely a file-handle trait.
This separates "filesystem instance" from "file handle". Backends that
need per-instance configuration (S3 bucket name + credentials, mmap
default advice, io_uring runtime, block-cache controller `Arc`) gain a
typed home in `Self::ContextConfig`, and `list_files`/`exists`/`open`
become `&self` methods on the filesystem handle.
Concrete filesystem handles introduced for the three existing backends:
- `MmapFs` — unit struct; `ContextConfig = ()`; produces `MmapFile`.
- `IoUringFs` — carries `prevent_caching`; `ContextConfig =
IoUringConfigContext`; produces `IoUringFile`. `IoUringConfigContext`
becomes the construction input rather than a per-open argument.
- `BlockCacheFs` — carries `Arc<CacheController>`; `ContextConfig =
BlockCacheConfigContext`; produces `CachedSlice`.
`TConfigContext` (universal builder methods) is kept so generic-over-`Fs`
code can still set cross-backend knobs via
`Fs::ContextConfig::default().with_prevent_caching(true)`.
Wrappers (`ReadOnly<S>`, `TypedStorage<S, T>`, `StoredStruct<S, T>`),
higher-level storages (`StoredBitSlice<S>`, `UniversalHashMap<K, V, S>`)
keep `<S: UniversalRead>` parameterization but their `open(...)`
constructors now grow a `fs: &Fs` argument bound by
`Fs: UniversalReadFs<File = S>`. `read_json_via` becomes
`read_json_via(fs: &Fs, path)`.
All test code and benches in `common` updated to construct
`MmapFs`/`IoUringFs` inline as needed. `common` compiles cleanly with
tests and benches. `gridstore`, `segment`, and `tonic` caller updates
are in flight in subsequent commits.
* WIP: gridstore + segment caller sweep (partial)
Threads `fs: &Fs` through gridstore's `BitmaskGaps`, `Bitmask`, `Pages`,
and `Gridstore::new`/`open`/`create_new_page`. Most segment callers
have `OpenOptions { extra: ... }` removed and `S::open(path, opts, ctx)`
sites updated mechanically but the trait change is not yet propagated.
Does NOT compile yet. Tracker still has static `S::open` calls, segment
generic constructors (`MmapInvertedIndex<S>`, `UniversalMapIndex`,
`StoredGeoMapIndex`, etc.) still call `S::open`/`S::list_files`/`S::exists`
statically — they need an `fs: &Fs` parameter added. Tonic API
`StorageReadService<S>` also unconverted.
Committed as branch checkpoint; cascade continues in subsequent work.
* gridstore: thread `fs: &Fs` through Bitmask, BitmaskGaps, Pages, Tracker
Per the new `UniversalReadFs` shape, every constructor/method that opens
files takes an `fs: &Fs` parameter. Gridstore is currently mmap-only,
so the top-level `Gridstore` / `GridstoreReader::open` callers in the
crate pass `&MmapFs` inline. Tests do the same.
gridstore lib + tests now compile cleanly. Segment + tonic cascade
still pending.
* WIP: segment caller sweep — dynamic_stored_flags first
* WIP: segment cascade - id_tracker partial
* common benches: update to new UniversalReadFs::open shape (clippy clean)
* segment flags: thread Fs through BufferedDynamicFlags / Bitvec / Roaring
DynamicStoredFlags::set_len now takes `fs: &Fs`. BufferedDynamicFlags
stores an `Arc<Fs>` so the flusher closure can call `set_len` on resize.
BitvecFlags and RoaringFlags expose a new `Fs` type parameter and the
flag tests now pass `Fs::default()` (MmapFs/IoUringFs via duplicate_item).
Concrete consumers (bool/null index, mmap dense/multi/sparse storages)
pin `Fs = MmapFs` and pass `&MmapFs` to inner opens.
* segment: thread Fs through field-index lifecycle methods
Apply the new UniversalReadFs::open shape across:
- full_text_index (MmapInvertedIndex, MmapFullTextIndex, UniversalPostings)
- geo_index (StoredGeoMapIndex build/open + tests + builders)
- numeric_index lifecycle (UniversalNumericIndex build/open)
- map_index lifecycle (UniversalMapIndex build/open)
- stored_point_to_values (open / from_iter)
Concrete consumers pin Fs = MmapFs and pass &MmapFs inline; generic
open paths thread `fs: &Fs` where Fs: UniversalReadFs<File = S>.
* segment: thread Fs through chunked vectors and id-tracker callers
- ChunkedVectors gains `Fs` generic so add_chunk can call create_chunk
after open. ChunkedVectorsRead/load_config and chunks::{read_chunks,
create_chunk} take `fs: &Fs`. Concrete callers (dense / multi-dense /
sparse / quantized) pass MmapFs inline.
- DenseVectorStorageImpl stores `fs: Fs` so `update_from` can reopen
ImmutableDenseVectors. ImmutableDenseVectors::open takes `fs: &Fs`.
- VectorStorageEnum DenseUring* variants thread IoUringFs alongside
IoUringFile.
- segment_builder + segment_constructor_base pass MmapFs to
ImmutableIdTracker::{new, open}.
- QuantizedStorage::from_file takes `fs: &Fs`; quantized_vectors callers
pass &MmapFs.
* segment: apply nightly rustfmt after Fs refactor
cargo +nightly fmt --all over the segment crate after the
UniversalReadFs cascade. No semantic changes.
* segment: thread Fs through benches and id-tracker tests
Update the dynamic-mmap-flags and buffered-update-bitslice benches to
the new UniversalReadFs::open shape (pass `&MmapFs`). Update the
immutable-id-tracker test suite to forward `&MmapFs` to
`from_in_memory_tracker` / `open`.
* uio: pin Fs via UniversalRead::Fs assoc type; per-call OpenExtra
Two design changes that fall out of the per-instance Fs refactor:
1. Bidirectional Fs ↔ File pinning. `UniversalRead::Fs:
UniversalReadFs<File = Self>` lets generic-over-`S` code refer to
`S::Fs` directly instead of carrying an extra `<Fs: UniversalReadFs<File = S>>`
generic param. `ReadOnly<S>` wraps a file but has no natural
filesystem; a phantom `ReadOnlyFs<S::Fs>` satisfies the constraint
while inherent `ReadOnly::open` keeps taking `&S::Fs` directly.
2. `prevent_caching` moves from filesystem-instance state to per-call
`UniversalReadFs::OpenExtra: Default`. Was previously a knob on
`IoUringConfigContext` / `IoUringFs`, conflating "how this fs is
built" with "how this file is opened." Now `IoUringFs::OpenExtra =
IoUringOpenExtra { prevent_caching }`; mmap and block-cache use `()`.
`IoUringConfigContext` is gone, `TConfigContext` slims to a `Default`
marker.
Tonic StorageReadService holds `Arc<S::Fs>` (was `PhantomData<S>`); its
`new()` builds via `S::Fs::from_context(default)` and the spawn_blocking
closures clone the Arc to call instance methods.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* uio: cfg-gate IoUringOpenExtra import for non-linux builds
The IoUringOpenExtra reexport from `universal_io` is gated on
`target_os = "linux"`. The previous commit left an unconditional import
in `persisted_hashmap/tests.rs`, breaking macOS/Windows CI.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* uio: migrate simple_disk_cache to per-instance Fs API
PR #9097 (merged into dev concurrently with this branch) introduced a
`DiskCache<R: UniversalRead>` using the pre-refactor trait shape:
file-handle-as-filesystem (`R::open`, `R::list_files`), an
`OpenOptionsExtra` field on `OpenOptions`, and trait methods without
`&self`. The Fs-instance refactor on this branch removed all three.
Reshape `simple_disk_cache` to match the new design without breaking the
lazy mirror semantics:
- New `DiskCacheFs<R>` is the filesystem handle. Holds a clone of the
remote `R::Fs`; `list_files`/`exists` delegate; `from_context`
forwards to the inner Fs context. `open` constructs a `DiskCache<R>`
via the global `DiskCacheConfig`.
- `DiskCache<R>` now stores `remote_fs: R::Fs` + `remote_extra:
<R::Fs as UniversalReadFs>::OpenExtra`, so lazy remote opens go
through `self.remote_fs.open(path, options, extra)` instead of the
removed `R::open`. `open_with_config` takes the remote Fs + extra
explicitly (no more hard-coded `prevent_caching: true`; callers pass
the appropriate `OpenExtra`).
- `UniversalRead for DiskCache<R>` now declares `type Fs =
DiskCacheFs<R>` (no more `open` method on the file trait).
- Propagate the necessary bounds (`R::Fs: Clone`,
`<R::Fs as UniversalReadFs>::OpenExtra: Clone`,
`R::OwnedReadPipeline<u8, Range<u32>>: Send`) through `pipeline.rs`
free functions and impl blocks that reach into `DiskCache::remote` /
`local_state`.
- Drop the now-removed `extra: _` destructure in `LocalState::new`.
- Tests construct the remote Fs via `R::Fs::from_context(Default::default())`
and exercise `DiskCache::open_with_config`. 17 simple_disk_cache
tests pass; the 3 `empty_read_does_not_materialize_local_file`
failures pre-exist on dev (verified) and are unrelated.
`mold -run cargo clippy --all-targets` clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* style: nightly fmt on simple_disk_cache migration
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* style: nightly fmt for local_state imports after rebase
Co-authored-by: Cursor <cursoragent@cursor.com>
* uio: split DiskCacheFs into its own module
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* uio: replace TConfigContext with OpenExtra trait; move DiskCacheConfig onto DiskCacheFs
- Drop the empty TConfigContext marker; ContextConfig is now unconstrained
so backends can require explicit construction.
- Add OpenExtra trait with with_prevent_caching for backend-agnostic
per-call knobs; impl for () (no-op) and IoUringOpenExtra.
- DiskCacheFs now carries Arc<DiskCacheConfig> via the new
DiskCacheFsContext<C>; the prefill flow moves from the deleted
open_with_config into DiskCacheFs::open so Populate::Blocking /
PreferBackground work through the trait API.
- Remove the DiskCacheConfig global; callers must construct the context
explicitly.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* style: nightly fmt after OpenExtra refactor
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs: fix spelling — Implementors → Implementers
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: possible panic insetad of error propagation
* fix: missing cfg annotation
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Cursor Agent <agent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Daniel Boros <dancixx@gmail.com>
|
||
|
|
cf085d31db |
build(deps): bump geohash from 0.13.1 to 0.13.2 (#9169)
Bumps [geohash](https://github.com/georust/geohash.rs) from 0.13.1 to 0.13.2. - [Release notes](https://github.com/georust/geohash.rs/releases) - [Commits](https://github.com/georust/geohash.rs/compare/v0.13.1...v0.13.2) --- updated-dependencies: - dependency-name: geohash dependency-version: 0.13.2 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
117a2a93fd |
[UIO] Impl DiskCache::reopen (#9143)
|
||
|
|
a6bf261bc4 |
Merge pull request #9034
* Add `iter_batch` method to `EncodedStorage` and `EncodedVectors` * Remove `EncodedVectors::for_each_in_batch` method * Add `for_each_in_multi_batch` and `score_vector_max_similarity` metho… * Refactor `score_stored_batch` for quantized multi-vector scorers... * Remove `QuantizedMultivectorStorage::score_multi` * Implement `score_points_batch_mmap` and `score_points_batch_uring`... * Implement runtime routing between mmap and io_uring batch scoring met… * review: rename + comments |
||
|
|
840f6aebf7 |
Append-only mutable segment: groundwork (filter fix + point_id refactor) (#9156)
* refactor: pass point_id into handle_point_version_and_failure Drop the external_id reverse lookup used for error_status correlation and take the external point_id as an explicit argument instead. All callers already have it in scope, and decoupling it from op_point_offset keeps error correlation correct in flows where the old internal_id is tombstoned before the recovery check runs. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: filter_deferred_and_deleted also consults the deleted bitslice Field-index primary-clause iterators (and the analogous plain/sparse vector-index paths) are routed through PointMappingsRefEnum to apply the deferred-threshold cutoff. The old `filter_deferred` only applied that threshold, so any soft-deleted internal id sitting below the cutoff slipped through whenever its field-index posting was still live. This was fine while the only source of mid-range tombstones was the deferred-tail design, but it breaks the moment a tombstone can land anywhere in the id range. Add an unconditional deleted-bitslice check (single bit test per element) and rename to `filter_deferred_and_deleted` so the contract is visible at the call site. The Either split is preserved so the no-threshold path still avoids the cutoff comparison. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * style: cargo +nightly fmt Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * style: collapse error-status recovery match arm Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
9b9e1034d7 |
[UIO] Require T: Send on UniversalRead (#9147)
Introduce a `pub trait Item: bytemuck::Pod + Send` marker (mirroring the existing `UserData` pattern) and use it in place of `T: bytemuck::Pod` on the read side of `UniversalRead`, its pipeline traits, and all impls/wrappers. Some implementations buffer `Vec<MaybeUninit<T>>` that may be transferred across threads in future backends; tightening the bound makes that explicit at the trait level instead of leaving callers to add `+ Send` ad-hoc. Write paths keep the looser `bytemuck::Pod` bound since they only borrow `&[T]`. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
a57df81e1f |
sparse: use zerocopy (#9140)
* sparse: use zerocopy for mmap reads * loaders::Csr: use zerocopy |
||
|
|
98ca1dc0c5 |
Clear id tracker page cache after building a segment (#9137)
The segment builder evicts the page cache of vector storage, quantized vectors, payload, payload index and the vector index after a build to avoid cache pollution, but the id tracker was never cleared. Its on-disk files (mappings, versions, deleted bitslice) are written during the build and stay resident in the page cache, so after each optimization the id tracker files linger as cache even though they are meant to be on-disk only (expected_cache_bytes == 0). Add `IdTracker::clear_cache` (default no-op) plus a `clear_cache_if_on_disk` policy wrapper, and implement the eviction for the mutable and immutable trackers: - ImmutableIdTracker pages out its two mmap-backed storages via madvise and drops the RAM-loaded mappings file via fadvise(DONTNEED). - MutableIdTracker drops its append-only log files via fadvise(DONTNEED). The builder calls `clear_cache_if_on_disk`, mirroring the payload index. The id tracker has no on-disk mode yet, so this always clears for now; a TODO marks where to gate it once an on-disk mode exists. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
1c22e8cd90 |
fix: reject source-superset schema mismatch at SegmentBuilder (#9110)
* fix: serialize optimization proxy install against shard updates execute_optimization captures `target_config` from the optimizer's frozen config and then wraps source segments in proxies. Between those two points, `CollectionUpdater::update` can apply a `CreateVectorName(V)` to the source segments via `apply_segments`, leaving the optimizer with sources that have V but a target_config that does not. The optimization then produces a merged segment without V, and a follow-up optimization (running with the refreshed config that includes V) fails to use that segment as a source: "Cannot update from other segment because it is missing vector name X". Close the race by extending the scope of the existing `LockedSegmentHolder::acquire_updates_lock` to cover the proxy install window. `CollectionUpdater::update` already takes this lock before processing any shard update, so concurrent writers wait until proxies are in place — at which point further mutations hit the proxies (recorded as intent and propagated to the merged segment in `finish_optimization`) instead of the originals. The guard is dropped right after proxy install so the slow build phase does not extend it. Tests: - Three `SegmentBuilder::update` tests document the precondition the lock now guarantees: with a target schema that adds a named vector the source lacks, update errors with "missing vector name X". Quantized and mixed-source variants exercise the same error path. - `test_optimize_blocks_proxy_install_on_updates_lock` asserts the invariant directly: while the updates lock is held, proxies are not yet installed. Verified to fail when the new guard is removed (otherwise it passes because `finish_optimization` also takes the same lock, so a naive "did optimize finish?" check would not catch a missing proxy-install guard). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(optimize): finish_optimization lock order; drop redundant tests Address review of #9110: 1. `finish_optimization` was acquiring `upgradable_read` before `acquire_updates_lock`, while the new guard at the start of `execute_optimization` acquires them in the reverse order. With two optimizer threads in flight, thread A in `finish_optimization` could hold `upgradable_read` and wait on `updates_lock` while thread B at the top of `execute_optimization` held `updates_lock` and waited on `upgradable_read` (parking_lot allows only one upgradable reader), deadlocking. Swap `finish_optimization` to take `updates_lock` first so both halves agree. 2. Drop the quantized and mixed-source variants of the inverted `SegmentBuilder` unit test — all three asserted the same error path (the mismatch check fires before quantization training or per-source branching), so only one is useful as documentation of the precondition. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: reject source-superset schema mismatch at SegmentBuilder Drop the lock approach (deadlocked test_continuous_snapshot) and fix the bug at the merge layer instead. Snapshot's proxy_all_segments_and_apply acquires the segment_holder upgradable_read first and then takes acquire_updates_lock tactically inside the snapshot operation. The previous commits' lock-extension acquired updates_lock before upgradable_read, so a snapshot in flight and an optimization just entering execute_optimization could deadlock holding each other's required next lock. Snapshot cannot easily reverse its order — that would hold updates_lock for the entire snapshot duration, blocking all writes. Move the fix to where the actual harm happens: SegmentBuilder::update iterates the target's vector_data and silently drops source vectors that aren't in target. That silent drop is what produces the broken merged segment in the CreateVectorName-vs-optimizer race. Add a check that every source vector name is in the target schema; the optimization aborts cleanly on mismatch and the next round (with refreshed config) merges correctly. This is strictly stronger than the lock: the lock only closed the window where V arrived *during* the proxy-install region. The schema check catches both that window and the window where V's apply_segments completed before the optimizer's lock acquisition. Diff is contained to lib/segment; no locking changes, no cross-crate plumbing. test_continuous_snapshot passes again. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(segment_builder): use Cancelled instead of ServiceError for schema mismatch ServiceError flips the shard to RED status (via `report_optimizer_error` → `segments.optimizer_errors`) and stays sticky until the next `recreate_optimizers_blocking` clears it. That's the right shape for hardware/IO failures but wrong for the schema-mismatch case here, which is an expected, recoverable race outcome — the next optimizer round with a refreshed target_config merges the same originals cleanly. `Cancelled` is the variant the optimization worker treats as a recoverable cancellation: logged at debug, tracker marked Cancelled, no `report_optimizer_error` call, no RED status. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(segment_builder): also use Cancelled for the existing target-superset error The existing "missing vector name" check at the start of the merge loop also fires during a race — specifically the optimizer-vs-DeleteVectorName shape, where V is removed from originals before J wraps proxies but J's frozen target_config still has V. Like the new source-superset check, this is an expected, recoverable race outcome, so use Cancelled instead of ServiceError to avoid flipping the shard to RED. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
aa4bafdcc2 |
Restore pre-UIO-migration OpenOptions behavior (#9107)
* [ai] segment dynamic_stored_flags: need_sequential(true -> false) https://github.com/qdrant/qdrant/blob/v1.17.0/lib/segment/src/common/flags/dynamic_mmap_flags.rs#L115-L146 * [ai] segment immutable_id_tracker: need_sequential(true -> false) https://github.com/qdrant/qdrant/blob/v1.17.0/lib/segment/src/id_tracker/immutable_id_tracker.rs#L249-L262 * [ai] segment geo_index: need_sequential(true -> false) https://github.com/qdrant/qdrant/blob/v1.17.0/lib/segment/src/index/field_index/geo_index/mmap_geo_index.rs#L239 * [ai] segment geo_index: populate(Auto -> from(!is_on_disk)) https://github.com/qdrant/qdrant/blob/v1.17.0/lib/segment/src/index/field_index/geo_index/mmap_geo_index.rs#L239 * [ai] segment numeric_index: need_sequential(true -> false) https://github.com/qdrant/qdrant/blob/v1.17.0/lib/segment/src/index/field_index/numeric_index/mmap_numeric_index.rs#L171 * [ai] segment map_index: need_sequential(true -> false) https://github.com/qdrant/qdrant/blob/v1.17.0/lib/segment/src/index/field_index/map_index/mmap_map_index.rs#L68-L71 * [ai] segment map_index: populate(Auto -> from(!is_on_disk)) https://github.com/qdrant/qdrant/blob/v1.17.0/lib/segment/src/index/field_index/map_index/mmap_map_index.rs#L71 * [ai] segment full_text_index: need_sequential(true -> false) https://github.com/qdrant/qdrant/blob/v1.17.0/lib/segment/src/index/field_index/full_text_index/inverted_index/mmap_inverted_index/mod.rs#L117-L133 * [ai] segment full_text_index: populate(Auto -> from(populate)) https://github.com/qdrant/qdrant/blob/v1.17.0/lib/segment/src/index/field_index/full_text_index/inverted_index/mmap_inverted_index/mod.rs#L133 * [ai] gridstore bitmask: need_sequential(true -> false) https://github.com/qdrant/qdrant/blob/v1.17.0/lib/gridstore/src/bitmask/mod.rs#L120 * [ai] gridstore bitmask: advice(Global -> Random) https://github.com/qdrant/qdrant/blob/v1.17.0/lib/gridstore/src/bitmask/mod.rs#L40 * [ai] gridstore bitmask/gaps: advice(Global -> Normal) in create https://github.com/qdrant/qdrant/blob/v1.17.0/lib/gridstore/src/bitmask/gaps.rs#L103 * [ai] gridstore bitmask/gaps: populate(Blocking -> No) in open https://github.com/qdrant/qdrant/blob/v1.17.0/lib/gridstore/src/bitmask/gaps.rs#L119 * [ai] segment quantized_storage: need_sequential(true -> false) https://github.com/qdrant/qdrant/blob/v1.17.0/lib/segment/src/vector_storage/quantized/quantized_mmap_storage.rs#L32-L40 * [ai] gridstore pages: advice(Global -> Random) in attach_page https://github.com/qdrant/qdrant/blob/v1.17.0/lib/gridstore/src/page.rs#L38 * [ai] tonic storage_read_api: update OpenOptions for read-only handlers Rationale: - writeable: false — all 6 are read-only APIs; true would fail on read-only mounts. - need_sequential: false — opening a second mmap with MADV_SEQUENTIAL on Linux costs a VMA per call. For one-shot reads driven by a single gRPC request, the second mmap isn't justified — better to put the hint on the primary mmap via advice. - populate: No — these are one-shot reads, no point warming the page cache. - advice: Sequential for read_bytes_stream (streams the file in 1MB chunks) and read_whole (reads everything) — the kernel can read ahead. - advice: Normal for file_length (doesn't read content), read_bytes (single range), read_batch / read_multi (random ranges). |
||
|
|
9df8ae95e1 |
Make OpenOptions explicit (#9104)
* Refactor: make OpenOptions explicit * Refactor: remove unused OpenOptions::disk_parallel * Refactor: do not wrap `advice` in Option * Refactor: Add OpenOptionsExtra |
||
|
|
2b5cf80f66 | Warn on clippy::wildcard_enum_match_arm (#9096) | ||
|
|
37e9400c2e |
Introduce simple disk cache (#8792)
* [AI + manual] initial impl [AI] read_batch which actually batches manual nits [AI] better handling of local and remote paths manual refactor, respect open options don't delete local file dumbify read_batch we want to refactor it anyway simplify rename to `DiskCache` in `simple_disk_cache` module * refactor to use always use ReadPipeline pass meta to remote pipeline * nits * run tests for more Remotes * fix no more <T> in UniversalRead * fmt * chore(deps): unify roaring as workspace dep, move duplicate to dev-deps Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: generall <andrey@vasnetsov.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
03e15b5ed4 |
refactor: move deferred-point ownership into ID tracker (#9062)
* refactor: move deferred-point ownership into the ID tracker Re-implements the idea from #8512 against current `dev`. Deferred-point state (`deferred_internal_id` + `deferred_deleted_count`) moves out of `Segment.deferred_point_status` and the cached `SparseVectorIndex.deferred_internal_id` field into `PointMappings`, exposed through `IdTrackerRead`. The threshold is set once at `MutableIdTracker::open` time; `PointMappings::drop` now maintains the deleted counter inline (with double-delete protection), removing the manual increment in `delete_point_internal` and the `calculate_deleted_deferred_point_count` rescan. Read paths consume the threshold through the id tracker: - The segment read view drops the `deferred_point_status` field and `with_view` no longer threads it in; `read_view/{deferred,info}.rs` call `self.id_tracker.deferred_*()` directly. - `SparseVectorIndex` no longer stores its own copy and its `update_vector` / search debug-assert read from `self.id_tracker.borrow().deferred_internal_id()`. - `VectorQueryContext.deferred_internal_id` and the `SegmentQueryContext::get_vector_context` parameter are gone; the three downstream readers (`plain_vector_index`, sparse search, sparse `update_vector`) consult their own id tracker. `PointMappingsRefEnum` centralises the dispatch: - `iter_internal_with_behavior(DeferredBehavior)` replaces ad-hoc branches in `iter_filtered_points` impls. - `external_iter_cutoff(DeferredBehavior)` covers iterators sourced outside the mapping (field-index outputs in `struct_payload_index::iter_filtered_points`). - The internal `deferred_internal_id()` accessor is private; the raw threshold no longer leaks to consumers. - `iter_from_visible` / `iter_random_visible` read the mapping's own threshold; callers that previously passed `DeferredBehavior::apply(...)` now branch on `deferred_behavior.include_all_points()` (scroll / order_by) or simply drop the argument (sampling / facet). `PayloadIndexRead::query_points` drops the now-redundant `deferred_internal_id` parameter; `iter_filtered_points` takes `DeferredBehavior` directly so HNSW build/search can request `IncludeAll` while normal reads request `Exclude`. RocksDB-related parts of the original PR are skipped — that tracker is already gone from `dev`. Tests adapted: sites that mutated `segment.deferred_point_status` directly now construct a parallel non-deferred segment via `create_deferred_segment(..., 0)` for comparison; `test_deleted_deferred_point_count` reads counters through the id tracker. See `docs/plans/deferred-points-owned-by-id-tracker.md` for the design write-up. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(benches): drop stale deferred_internal_id arg from query_points calls The boolean / range / conditional bench files weren't built by `cargo test -p segment`, so they slipped through. `cargo clippy --workspace --all-targets` catches them. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: drop id_tracker / point_mappings args from iter_filtered_points Both impls already hold an id tracker on `self`: - `StructPayloadIndexReadView` carries `id_tracker: &'a I`, so `self.id_tracker.point_mappings()` borrows from `'a` and the lazy iterator chain keeps working unchanged. - `PlainPayloadIndex` carries `id_tracker: Arc<AtomicRefCell<...>>`, where the mapping borrow is local; collect into a `Vec` and return `into_iter()`. PlainPayloadIndex::iter_filtered_points has no direct callers — only `query_points` was using it — so eager collection is a non-issue. While here, take `self` by value on `iter_internal_visible`, `iter_from_visible`, `iter_random_visible`, `iter_internal_with_behavior`, and `external_iter_cutoff`. `PointMappingsRefEnum` is `Copy`; this matches the existing `iter_internal` / `iter_from` / `iter_random` shape and lets the iterator outlive a local `let point_mappings = ...;` binding. The HNSW `condition_points` helper drops its now-unused `id_tracker` parameter. All callers (sampling, scroll, order_by, facet ×2, hnsw build/search) just drop the two arguments. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore: ignore /docs/plans/ and untrack the previously-committed plan `docs/plans/` is a scratch directory for per-feature planning notes — not something we want under source control. Add it to `.gitignore` and drop the deferred-points plan that slipped into history; the design is captured in the PR description. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: replace external_iter_cutoff with filter_deferred iterator wrapper Instead of exposing a raw `Option<PointOffsetType>` cutoff that every caller has to apply with their own `.filter(...)`, give `PointMappingsRefEnum` an iterator wrapper: fn filter_deferred<I: Iterator<Item = PointOffsetType>>( self, iter: I, deferred_behavior: DeferredBehavior, ) -> impl Iterator<Item = PointOffsetType> It returns the iterator unchanged for `IncludeAll` (or when the mapping has no threshold) and otherwise wraps it in a cutoff `.filter`, dispatched via `itertools::Either` so the no-cutoff path stays allocation-free. The struct payload index's `iter_filtered_points` swaps its open-coded filter for a single `point_mappings.filter_deferred(...)` call. The deferred threshold no longer leaks out of `PointMappingsRefEnum`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: move deferred wrapping out of peek_top_all, gate it as test-only `BatchFilteredSearcher::peek_top_all` baked the deferred cutoff into its iterator construction, which was the last place outside `PointMappingsRefEnum` that knew about the threshold. Split the deleted-iteration concern out into a new accessor: fn iter_not_deleted(&self) -> impl Iterator<Item = PointOffsetType> + 'a It borrows `&'a BitSlice` directly (not via `&self`), so callers can chain `filter_deferred` and then move `self` into `peek_top_iter` without lifetime conflicts. Sparse + plain vector index call sites now do: let iter = id_tracker .point_mappings() .filter_deferred(searcher.iter_not_deleted(), DeferredBehavior::Exclude); searcher.peek_top_iter(iter, &is_stopped) leaving `BatchFilteredSearcher` completely ignorant of deferred state. With deferred handling lifted out, `peek_top_all` itself is now used only by tests (3 inline `#[cfg(test)] mod tests`, 1 integration test, 1 bench) — gate it under `#[cfg(feature = "testing")]` to match `new_for_test`. Production code goes through the `iter_not_deleted` + `filter_deferred` + `peek_top_iter` composition. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: optimize Segment::retrieve and thread user data through read_vectors Two interlocking changes that together collapse the per-point lookups and intermediate allocations in `Segment::retrieve` down to one external-to-internal pass. ## `IdTrackerRead::resolve_external_ids` (new default trait method) Single-pass translation of a `&[PointIdType]` slice into two parallel vectors `(Vec<PointIdType>, Vec<PointOffsetType>)`. Folds deferred filtering (compare offset against the threshold inline — no separate `point_is_deferred` lookup) and missing-id errors (eager `PointIdError`) into resolution. Lives on the trait so the deferred threshold never leaks out of the id tracker; the parallel-vector shape lets a future batched payload / vector fetcher consume `&offsets` straight without unzipping. The `appendable_flag` guard previously in `point_is_deferred` is gone: non-appendable trackers always carry `deferred_internal_id() == None` (set only via `MutableIdTracker::open`, guarded by the segment constructor), so the check was load-bearing nowhere. ## User-data threading through `read_vectors` `VectorStorageRead::read_vectors` now takes `IntoIterator<Item = (U, PointOffsetType)>` and yields `(U, PointOffsetType, CowVector)`. The user-data tag rides alongside each offset all the way through, so callers can map results back into a parallel array without keeping a separate `offset → ...` lookup table. - Default trait impl: one-line per-key loop. - Dense impl: `unzip()` into parallel `(Vec<U>, Vec<PointOffsetType>)` in a single pass — same allocation count as before, just U riding alongside. - Enum delegations (`VectorStorageEnum`, `VectorStorageReadEnum`) forward unchanged. - `for_each_in_batch` and below stay untouched. `SegmentReadView::vectors_by_offsets<U: Copy>` becomes a lazy filter chain — no parallel `Vec<(orig_idx, offset)>` allocation. The dead `SegmentReadView::read_vectors` helper is removed. ## `Segment::retrieve` end-to-end Per N points / V vectors / payload: | Operation | Before | After | |----------------------------|---------------------|-------| | `id_tracker.internal_id` | N × (1 + V + 1) | N | | `id_tracker.external_id` | N × V | 0 | | `point_is_deferred` | N (when applicable) | 0 | | `offset_to_id` HashMap | N entries | none | | `Vec` in `vectors_by_offsets` | 1 | 0 | The vectors stage passes the external id as `read_vectors`'s user data — the callback gets `id` directly without any index lookup. The payload stage uses `payload_by_offset` against the already-resolved offsets. The shape is also batch-friendly: swapping in a future `IdTrackerRead::batch_internal_id` or `payload_index.batch_get_payload` needs no changes outside the two call sites. ## Behavioural notes - Missing-id now errors eagerly inside resolution, instead of in the vectors stage (`WithVector::Bool(true)` / `Selector`) or payload stage (`with_payload.enable`). The previous `WithVector::Bool(false)` + no-payload path silently inserted an empty record; that is now also an error. None of the existing callers (search post-processing, external retrieve API, the deferred-points test on tests/mod.rs:1179) pass non-existent ids. - Added a per-payload `check_stopped`; the vectors stage already had `stop_if` on its iterator chain. - `vector_by_offset` (the single-element helper) passes `()` as the no-op user data. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * style: apply rustfmt to optimised retrieve / read_vectors paths Pre-push hook failure on the previous commit was rustfmt. Same content, formatted. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * do not error out on missing points in retrieve --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
7035440f4c |
Error propagation for sparse InvertedIndex (#9076)
* Remove unused InvertedIndexMmap / InvertedIndexImmutableRam * Error propagation for InvertedIndex trait |
||
|
|
ba710e1c50 | Remove unused InvertedIndexMmap / InvertedIndexImmutableRam (#9074) | ||
|
|
1d910011b6 |
Refactor MultivectorsOffsetsStorageMmap to use UniversalRead (#9071)
* Empty commit to open a PR * Refactor `MultivectorOffsetsStorageMmap` to use `MmapFile` instead of `MmapSlice` * simpler result types --------- Co-authored-by: generall <andrey@vasnetsov.com> |
||
|
|
6d96cbcb53 |
[UIO] Generic quantized storage (#9059)
* use generic in QuantizedMmapStorage * rename to QuantizedStorage * be explicit about S * rename builder to `QuantizedStorageBuilder` * rename file to `quantized_storage.rs` |
||
|
|
075dc0d247 |
accept Cow in quantization_preprocess (#9061)
|
||
|
|
6598a47933 | make quantized_vector_size method static (#9060) | ||
|
|
62594253c2 |
expose universal read in read only id tracker (#9057)
* Expose Universal Read in read-only ID tracker * fmt |
||
|
|
b04519bf25 | fix: skip deleted winners in for_each_unique_point (#9052) | ||
|
|
99355603f9 |
Fix match: {except: []} returning zero results with payload index (#9055)
* Add integration test for `match: {except: []}` with integer index
Regression test for https://github.com/qdrant/qdrant/issues/9050
An empty `except` list (NOT IN []) should always match all points that
have the field, both with and without a payload index. Currently the
integer (and keyword) indexed path incorrectly returns zero results
because:
1. serde deserializes `except: []` as `AnyVariants::Strings([])` (first
variant of the untagged enum)
2. The map index filter_impl returns `iter::empty()` for empty
cross-type variant, instead of matching everything
The test covers:
- except: [] without any index (baseline, currently works)
- except: [] with an integer index (currently broken)
- except: [] with a keyword index (currently broken)
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix `match: {except: []}` returning zero results with payload index
Fixes https://github.com/qdrant/qdrant/issues/9050
Root cause: `except: []` deserializes as `AnyVariants::Strings([])`
due to the untagged serde enum trying `Strings` first. On an integer
index, the filter_impl matched `Except + Strings(empty)` and returned
`iter::empty()` (zero results). The same issue existed symmetrically
on keyword indexes with `Integers(empty)` and on UUID indexes with
`Integers(empty)`.
The fix: when the cross-type variant reaches the Except branch (e.g.
Strings on an integer index, Integers on a keyword/UUID index), return
`None` unconditionally — regardless of whether the set is empty. This
delegates to the fallback condition checker, which already handles
`except: []` correctly by matching all values.
The `estimate_cardinality_impl` functions had the same bug (returning
`CardinalityEstimation::exact(0)` for the empty cross-type case) and
are fixed the same way.
Affected index types: integer, keyword (str), UUID.
Bool index is NOT affected — it already returns `None` for all
Any/Except conditions.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Expand match-except test to cover all index types and cross-type filters
Extends the regression test for #9050 to comprehensively cover:
- Integer, keyword, and UUID field indexes
- Empty except list (the original bug) for each index type
- Non-empty except list with matching types (normal filtering)
- Cross-type except values (e.g. strings on integer index, integers on
keyword/UUID index) — type mismatch should exclude nothing → all match
- Exclude-all: listing every value should return zero results
- All scenarios tested both with and without a payload index
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor Agent <agent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
||
|
|
d596261d9b |
Fix indexed integer range filter with float values (#9054)
* Add test, assert integer range bounds work properly on float fractionals * Use special float range conversion, respect integer bounds |
||
|
|
35ea107d02 |
Propagate S to ImmutableIdTracker (#9047)
|
||
|
|
1a7a5ac61a |
test: failing test for values_per_hash drift on duplicate geo point removal (#9043)
* test: add failing test for values_per_hash drift on duplicate geo point removal Reproduces the bug where `remove_point` only calls `decrement_hash_value_counts` once per unique geohash, while `add_many_geo_points` increments it once per value. When a point has duplicate geo coordinates (same geohash), the counters drift upward permanently after removal. Ref: https://github.com/qdrant/qdrant/pull/9033#discussion_r3241154045 Co-authored-by: Cursor <cursoragent@cursor.com> * fix: decrement values_per_hash once per value on geo point removal `add_many_geo_points` increments `values_per_hash` once per value, but `remove_point` deduplicated geohashes with a HashSet + `continue` that also skipped the per-value decrement. A point with duplicate geo coordinates (same geohash) therefore left the counters drifted upward permanently after removal. Move `decrement_hash_value_counts` above the dedup guard so it runs once per value, matching the increment side. `points_map` and `points_per_hash` track points, not values, so they stay deduplicated. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Cursor Agent <agent@cursor.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: generall <andrey@vasnetsov.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
3818ff593d |
refactor: read-only-fulltext (#9035)
* refactor: read only full text * fix: linter * refactor: follow new file structure * chore: remove storage enums * fix: linter * fix: FullTextReadIndex methods * chore: remove duplicated impls * fix: format * refactor: unify FullTextIndex telemetry via trait default method Add a `get_telemetry_data` default method to `FullTextIndexRead` built from the existing telemetry methods, so the `ReadOnlyFieldIndex` match arm is consistent with every other variant. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: move full-text lifecycle methods off FullTextIndexRead trait `populate`, `clear_cache`, `files` and `immutable_files` are lifecycle concerns, not part of the read surface. Move them from the `FullTextIndexRead` trait to inherent methods alongside the rest of each index's lifecycle code (open / wipe / flusher). Also split the `mutable_text_index` tests into their own `tests.rs`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: restructure full_text_index module layout Move the `FullTextIndex` enum into `full_text_index/mod.rs` and split its impls following the inner-module convention: lifecycle (constructors, builders, ValueIndexer, PayloadFieldIndex) into `lifecycle.rs`, read-path impls and shared free functions into `read_ops.rs`. Extract the `FullTextIndexRead` trait into its own `full_text_index_read.rs`. Drops the inherent `FullTextIndex::get_telemetry_data`, which duplicated the `FullTextIndexRead::get_telemetry_data` default method. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat: read-only counterpart of MutableFullTextIndex Extract the shared in-memory state of `MutableFullTextIndex` into `MutableFullTextIndexInner` (inverted index, config, tokenizer), which implements `FullTextIndexRead`. `MutableFullTextIndex` now wraps it alongside a writable `Gridstore`. Add `ReadOnlyAppendableFullTextIndex<S>`, wrapping the same inner state with a `GridstoreReader` over generic `UniversalRead`. Mirrors the `MutableMapIndex` / `ReadOnlyAppendableMapIndex` split; loading and lifecycle land in a follow-up. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: turn ReadOnlyFullTextIndex into a two-variant enum Rename the `read_only_text_index` module to `read_only` and reshape `ReadOnlyFullTextIndex` from a thin `MmapFullTextIndex` wrapper into an `Appendable` / `Immutable` enum, mirroring `ReadOnlyMapIndex`. Its `FullTextIndexRead` impl is now a dispatcher forwarding to the active variant. Drop the no-op inherent `populate` (mutable, immutable) and `immutable_files` (mutable) methods that tripped `clippy::unused_self` and `clippy::unnecessary_wraps` after the earlier trait-to-inherent move; `FullTextIndex::populate` inlines the in-RAM no-op directly. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: generall <andrey@vasnetsov.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
4ef1417443 |
[UIO] Introduce Populate enum (#8946)
* introduce `Populate` enum * fix: apply CodeRabbit auto-fixes Fixed 6 file(s) based on 6 unresolved review comments. Co-authored-by: CodeRabbit <noreply@coderabbit.ai> * fix coderabbit fix * Use from rather than into * add `Auto` variant * fix rebase * fix rebase again --------- Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> Co-authored-by: CodeRabbit <noreply@coderabbit.ai> Co-authored-by: timvisee <tim@visee.me> |
||
|
|
fd027ca4a2 |
refactor: read-only numeric index (#9038)
* refactor: split numeric index variants into dedicated modules Move MutableNumericIndex, ImmutableNumericIndex, and MmapNumericIndex into their own directories, each split into mod.rs (struct definitions), lifecycle.rs (open/build/wipe/mutations), and read_ops.rs (accessors), mirroring the map_index layout. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: drop single-variant Storage enums in mutable/immutable numeric index Replace `Storage<T>` wrappers around the only backing store with the store types directly: `Gridstore<Vec<T>>` for `MutableNumericIndex` and `Box<MmapNumericIndex<T>>` for `ImmutableNumericIndex`. Collapses the trivial single-arm matches into direct method calls. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: introduce NumericIndexRead trait Mirror `MapIndexRead` from the map_index refactor: define a `NumericIndexRead<T>` trait in `numeric_index/read_ops.rs` and implement it on each of the three storage variants (`MutableNumericIndex`, `ImmutableNumericIndex`, `MmapNumericIndex`). Trait signatures are unified across variants — in-memory variants accept and ignore the `hw_counter` argument that the mmap-backed variant uses for IO tracking, and `total_unique_values_count`, `values_range`, and `orderable_values_range` return `OperationResult` everywhere so the dispatcher in `NumericIndexInner` can call them generically. Variant-specific helpers that don't fit the shared shape stay as inherent methods: `MutableNumericIndex::map()`, `ImmutableNumericIndex::values_range_size()`, and `MmapNumericIndex::{values_range_size, is_on_disk}`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: add ReadOnlyAppendableNumericIndex Counterpart to `MutableNumericIndex`, mirroring `ReadOnlyAppendableMapIndex` from the map_index refactor. It reuses the shared `InMemoryNumericIndex` in-memory state but is backed by a `GridstoreReader` over generic `UniversalRead` instead of a writable `Gridstore`, and implements `NumericIndexRead` by forwarding to the in-memory index — no mutation surface. Loading / lifecycle (constructor, files, populate, clear_cache) will follow in a separate change; the storage field is held only to pin the on-disk layout for now. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: rename MmapNumericIndex to UniversalNumericIndex, expose UniversalRead param Mirror `UniversalMapIndex`: the type is now generic over `S: UniversalRead` with a `MmapFile` default, so the index can be served from any `UniversalRead` backend (io_uring, disk-cache wrappers, …) rather than the hard-coded `MmapFile`. The `NumericIndexRead` impl and read-side helpers are generic over `S`; `build` / `open` and the other lifecycle methods stay `MmapFile`-only since they construct mmap-backed storage from a path. The `NumericIndexInner::Mmap` enum variant keeps its name and uses the default `S = MmapFile`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: split numeric_index/storage/mod.rs into lifecycle and read_ops `storage/mod.rs` now holds only the `NumericIndexInner` enum and module wiring. The variant-dispatch impls are split into sibling modules matching the layout of the individual storage variants: - `lifecycle.rs`: construction, persistence, file listing, cache control, and `remove_point`. - `read_ops.rs`: read-path forwarding — value lookups, telemetry, RAM accounting, `is_on_disk`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: implement NumericIndexRead for NumericIndexInner The enum-level read dispatch was a set of inherent methods scattered across storage/read_ops.rs and storage/statistics.rs with signatures that drifted from the variant trait (`values_count` returned `usize`, `max_values_per_point` vs `get_max_values_per_point`, a hand-rolled `get_telemetry_data`). Make `NumericIndexInner` implement `NumericIndexRead` directly so it shares one interface with the three storage variants. - All 12 trait methods are forwarded via match dispatch in storage/read_ops.rs; `values_range` / `orderable_values_range` box the per-variant iterators. - `get_histogram`, `get_points_count`, `total_unique_values_count` move out of statistics.rs into the trait impl; `values_is_empty` and `get_telemetry_data` now come from the trait defaults. - `point_ids_by_value` and `is_on_disk` stay as enum-only inherent helpers (not part of the shared trait). - Callers updated: `NumericIndex::values_count` unwraps the now `Option`-returning trait method; `filter` boxes `point_ids_by_value`; `field_index.rs` and `numeric_field_index.rs` import the trait. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: add ReadOnlyNumericIndexInner Read-only counterpart to `NumericIndexInner`, mirroring `ReadOnlyMapIndex` from the map_index refactor. Lives under `numeric_index/storage/read_only` and selects across the two read-only storage backends: - `Appendable(ReadOnlyAppendableNumericIndex<T, S>)` — loaded into RAM from the appendable Gridstore format. - `Immutable(UniversalNumericIndex<T, S>)` — served directly from the immutable stored format. Implements `NumericIndexRead` by forwarding each method to the active variant; `values_is_empty` / `get_telemetry_data` come from the trait defaults. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: rename numeric_index read_ops to numeric_index_read, split mod.rs Two changes: - Rename `numeric_index/read_ops.rs` (the `NumericIndexRead` trait definition) to `numeric_index_read.rs`, freeing the `read_ops` name. - Split the leftover content of `numeric_index/mod.rs` into sibling modules, matching the per-variant layout: - `lifecycle.rs`: the `Encodable` key-format trait + impls and the `HISTOGRAM_*` construction constants. - `read_ops.rs`: the `StreamRange` trait and the `Range` → index-key-bounds conversion. `mod.rs` now only wires modules and re-exports. `Encodable` and `StreamRange` keep their public paths via re-export; `tests.rs` gains explicit imports for the symbols it previously picked up through the `mod.rs` glob. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: dissolve numeric_index/index.rs into mod, lifecycle, read_ops `index.rs` only held the `NumericIndex` wrapper and its two impl blocks; spread them to match the per-module layout used elsewhere: - `NumericIndex` struct + `NumericIndexIntoInnerValue` trait → `mod.rs` (type definitions live with the module wiring). - The inherent `impl NumericIndex` (open / build / cache control / storage introspection) → `lifecycle.rs`, alongside the `HISTOGRAM_*` seed constants. - The `PayloadFieldIndexRead` impl → `read_ops.rs`. Also move the `Encodable` key-format trait out of `lifecycle.rs` into its own `encodable.rs`. `mod.rs` keeps re-exporting `Encodable` so its public path is unchanged. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat: add ReadOnlyNumericIndex with NumericIndexRead + PayloadFieldIndexRead Read-only counterpart to `NumericIndex`, wrapping `ReadOnlyNumericIndexInner` plus the payload value type parameter `P`. Implements both `NumericIndexRead` and `PayloadFieldIndexRead` by forwarding to the inner storage-variant enum. To support `PayloadFieldIndexRead` without duplicating the query logic, the cardinality/filter/payload-block/condition-checker code is extracted into a new `query` module of generic free functions over `NumericIndexRead<T>`. `ReadOnlyNumericIndexInner` implements `PayloadFieldIndexRead` by plugging into those helpers; `ReadOnlyNumericIndex` delegates to its inner. The writable `NumericIndexInner` path is left untouched — its existing variant-specialized `estimate_points` heuristic stays in `storage`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: make query.rs the single source of truth for numeric index queries The generic `query` helpers and the per-`NumericIndexInner` impls in `storage/{trait_impls,statistics}.rs` had duplicated cardinality / filter / payload-block / condition-checker logic. Collapse them onto the shared `query` helpers: - `storage/trait_impls.rs`: `PayloadFieldIndexRead for NumericIndexInner` now forwards each method to `query::*` instead of carrying its own copy. - `storage/statistics.rs`: deleted — `range_cardinality` and `estimate_points` were duplicates of the `query` versions. - `estimate_points` needs a range size; add `values_range_size` to the `NumericIndexRead` trait with a default that counts `values_range`, overridden by the `Immutable` / `Mmap` variants with their `O(log n)` boundary search. The `MutableNumericIndex::map()` accessor (its only caller was the old `estimate_points`) is removed. - `values_range_size` takes `hw_counter` and threads it into `values_range` rather than fabricating a disposable counter. `tests.rs` calls `query::range_cardinality` directly now that the inherent method is gone. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat: integrate read-only numeric index into ReadOnlyFieldIndex Wire the four numeric variants (`IntIndex`, `DatetimeIndex`, `FloatIndex`, `UuidIndex`) into `ReadOnlyFieldIndex`, mirroring `FieldIndex`: - `PayloadFieldIndexRead` / `FieldIndexRead` dispatch covers the new variants — telemetry, value counts, value retrievers, `as_numeric`. - `ReadOnlyNumericFieldIndex` is the read-only counterpart of `NumericFieldIndex` (Int/Float order-by erasure over `ReadOnlyNumericIndexInner`); `as_numeric` returns it for the Int/Datetime/Float variants (UUIDs aren't numerically order-by-able, matching `FieldIndex`). - `ReadOnlyNumericIndex` gains per-`(T, P)` `value_retriever` methods (in `read_only/value_retriever.rs`) and an `inner()` accessor. `StreamRange` is now backed by a shared generic `query::stream_range` helper over `NumericIndexRead`, implemented for both `NumericIndexInner` and `ReadOnlyNumericIndexInner` — replacing the bespoke `EitherVariant` dispatch in `storage/trait_impls.rs`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor: collapse ReadOnlyNumericIndex value retrievers onto one generic method The four per-`(T, P)` `value_retriever` methods were identical except for the per-value `T -> Value` conversion. Extract that conversion into a `NumericValueToJson` trait (one tiny impl per `(T, P)`) and keep a single generic `value_retriever` that builds the retriever closure once. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fmt * refactor: dedup NumericFieldIndex / ReadOnlyNumericFieldIndex The two enums were structurally identical — same `StreamRange`, `get_ordering_values`, and `NumericFieldIndexRead` bodies — differing only in the backing storage type. Collapse them onto one generic `NumericFieldIndexView<'a, I, F>` with a single set of impls (over `I: NumericIndexRead<i64> + StreamRange<i64>` and the `f64` counterpart). `NumericFieldIndex` and `ReadOnlyNumericFieldIndex` are now type aliases of the generic view, so every existing call site is unchanged. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
e784896753 | feat: restructure null index (#9042) | ||
|
|
652af06f9a |
refactor: read only geoindex (#9033)
* refactor: read only geo index * fix: linter * fix: drop open temp * fix: read_ops * fix: address review fixes and use same file structure like map index * fix: review comments * fix: linter |
||
|
|
1acb4d8166 | Determinism for TQ Bits2 HNSW tests (#9036) | ||
|
|
045b1de45e |
Refactor quantized multi-vector scorers for io_uring support (#8988)
* Only read each vector once, when scoring quantized multi-vectors * Simplify `QuantizedMultiQueryScorer` Replace generic `TEncodedVectors` with concrete `QuantizedMultivectorStorage` * Simplify `QuantizedMultiCustomQueryScorer` Replace generic `TEncodedVectors` with concrete `QuantizedMultivectorStorage` * Further simplify `QuantizedMultiCustomQueryScorer` Remove `TElement` and `TMetric` type-parameters * Add `iter_offsets` method to quantized multi-vector storage types * Add `QuantizedMultivectorStorage::score_multi` method... ...which accepts `MultivectorOffset` instead of `PointOffsetType` * Implement `QuantizedMultiQueryScorer::score_stored_batch` * Implement `QuantizedMultiCustomQueryScorer::score_stored_batch` |
||
|
|
9c709e7891 | feat: add readonly bool (#9030) | ||
|
|
5f5e802395 | Fix flaky oversampling assertion in HNSW quantization tests (#9029) | ||
|
|
731293bcb8 |
refactor: read only boolindex (#9023)
* feat: add read only boolean index * fix: linter * chore: merge BoolIndex impl |
||
|
|
dc2a6df244 |
read only map index (#9025)
* rename mmap -> universal * [AI] Introduce `ReadOnlyMapIndex` and implement `MapIndexRead` for it * add missing file * [AI] Implement read-only PayloadFieldIndexRead / FieldIndexRead (#9026) Wires `ReadOnlyFieldIndex` into the read-only field-index surface that `ReadOnlyMapIndex` and `ReadOnlyNullIndex` already provide, sharing the non-trivial bodies with the existing `MapIndex` impls. - Extracted per-`N` `PayloadFieldIndexRead` bodies (filter, estimate_cardinality, for_each_payload_block, condition_checker) in `payload_index_impl/{str,int,uuid}.rs` into free functions over `T: MapIndexRead<N>`. Both `MapIndex<N>` and `ReadOnlyMapIndex<N, S>` now impl `PayloadFieldIndexRead` via thin delegation to those bodies. - Lifted `value_retriever` bodies in `value_indexer_impl.rs` the same way; `ReadOnlyMapIndex<N, S>` gets the per-K closure for free. - Added `telemetry_index_type` (required) + `get_telemetry_data` (default) to `MapIndexRead`, mirroring the `NullIndexRead` pattern; every map-index variant now reports telemetry without an inherent method on the enum. - Promoted `MapIndexRead` + its module to `pub` so `field_index_base/read_only/` can name them. - Implemented `FacetIndex for ReadOnlyMapIndex<N, S>` and extended `FacetIndexEnum` with `S: UniversalRead = MmapFile` plus the three read-only variants. `FieldIndex::as_facet_index` pins `S = MmapFile` via turbofish to keep inference happy. - `ReadOnlyFieldIndex::as_numeric` returns `None` for now; no concrete read-only numeric type exists yet (the doc on `NumericFieldIndexRead` already notes this is intended for a future `ReadOnlySegment`). |
||
|
|
aabc99fdd3 |
Replace Meta with U: UserData (#9009)
|
||
|
|
5e0f3a380c |
refactor: split geo_index/mod.rs into separate files (#9018)
The geo_index module was a single 2134-line file. Split it into focused files for better readability and maintainability: - mod.rs: GeoMapIndex enum definition and core methods (~290 lines) - builders.rs: GeoMapIndexMmapBuilder and GeoMapIndexGridstoreBuilder - payload_index.rs: ValueIndexer, PayloadFieldIndex, PayloadFieldIndexRead impls - tests.rs: all test code (~1340 lines) No logic changes; all 40 existing tests pass. Co-authored-by: Cursor Agent <agent@cursor.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
452473099e |
refactor(map_index): extract inner state; add ReadOnlyAppendableMapIndex skeleton (#9020)
* refactor(map_index): extract MutableMapIndexInner; add ReadOnlyAppendableMapIndex
Extract the four in-memory fields (`map`, `point_to_values`, `indexed_points`,
`values_count`) of `MutableMapIndex<N>` into a shared inner struct
`MutableMapIndexInner<N>` under `mutable_map_index/inner.rs`. The
`MapIndexRead<N>` impl moves wholesale onto the inner. `MutableMapIndex<N>`
becomes a thin wrapper `{ inner, storage: Gridstore<...> }`; its
`MapIndexRead<N>` impl forwards to the inner. The Gridstore load loop in
`open_gridstore` collapses to `MutableMapIndexInner::empty()` +
`inner.ingest(idx, value)` per row, sharing the per-value ingestion path with
the future read-only loader.
Add a new sibling `mutable_map_index/read_only/` module containing
`ReadOnlyAppendableMapIndex<N>` (inner + `GridstoreReader<Vec<...>, MmapFile>`),
with its `MapIndexRead<N>` impl forwarding to the same inner. No constructor
yet — loading and lifecycle for the read-only variant will land in a
follow-up. Field `storage` is held to pin the on-disk type and tagged
`#[allow(dead_code)]` until then.
No behavioural change. The 17 existing map_index tests pass unchanged.
* expose storage generic
|
||
|
|
eb4bac0782 |
refactor: read only nullindex (#9019)
* feat: add read only null index * fix: linter * chore: remove unused impl * chore: revert impl deletion * fix: linter errors and rename shared->read_ops * fix: linter * review fixes --------- Co-authored-by: generall <andrey@vasnetsov.com> |
||
|
|
67295ef87f |
refactor(map_index): split mod.rs into read_ops, lifecycle, tests (#9015)
* refactor(map_index): split mod.rs into read_ops, lifecycle, tests
mod.rs is now ~40 lines containing only the MapIndex enum, type
aliases, and module declarations.
- read_ops.rs: read-only inherent methods (get_values, get_iterator,
for_each_*, except_cardinality, except_set, telemetry, ram_usage,
mutability/storage type)
- lifecycle.rs: open/builder/flush/wipe/remove_point/files/populate/
clear_cache
- tests.rs: all #[cfg(test)] tests
Also folded payload_index_impl_{int,str,uuid}.rs into a dedicated
payload_index_impl/ submodule.
* refactor(map_index): split storage submodules into dedicated dirs (#9016)
* refactor(map_index): split storage submodules into dedicated module dirs
Turn each of the three storage implementation files into a directory
module split into read_ops + lifecycle, mirroring the parent module
layout introduced in #9015.
- mutable_map_index/{mod,lifecycle,read_ops}.rs
- immutable_map_index/{mod,lifecycle,read_ops}.rs
- mmap_map_index/{mod,lifecycle,read_ops}.rs
Each mod.rs holds only the struct definitions, internal Storage type,
and config/constants. read_ops.rs holds the read-only query methods;
lifecycle.rs holds open/build/flush/wipe/files/remove_point and
internal mutation helpers.
Pure refactor — no behavior change.
* refactor(map_index): introduce MapIndexRead trait
Define a unified read-only trait `MapIndexRead<N>` describing the
methods every storage variant exposes (check_values_any, get_values,
get_iterator, for_each_*, storage_type, ram_usage_bytes, etc.).
Each storage variant's read_ops.rs now contains a trait impl instead
of inherent methods. Signatures are unified across variants:
- `hw_counter` is accepted by every method that needs it for the mmap
variant; mutable / immutable accept and ignore it.
- `check_values_any` returns `bool` (mmap absorbs IO errors internally
with the existing FIXME, matching the parent's prior `.unwrap_or`).
- `for_each_count_per_value` takes `deferred_internal_id` uniformly;
the immutable variant `debug_assert!`s it is `None`.
`for_points_values` keeps its variant-specific callback signatures
and stays as an inherent method — it's only used by FacetIndex with
explicit pattern matching.
Pure refactor — no behavior change.
* refactor(mutable_map_index): drop single-variant Storage enum
The `Storage<T>` enum had only one variant (`Gridstore`), so every
`match &self.storage { Storage::Gridstore(s) => ... }` was just
unwrapping the same path. Replace the field with `Gridstore<Vec<...>>`
directly and inline every match.
|
||
|
|
d17514fb61 |
refactor(map_index): split monolithic mod.rs into focused submodules (#8999)
The ~2150-line map_index/mod.rs was difficult to navigate. Split it into focused files while preserving all logic and public API: - key.rs: MapIndexKey trait and impls for str, IntPayloadType, UuidIntType - builders.rs: MapIndexBuilder, MapIndexMmapBuilder, MapIndexGridstoreBuilder - payload_index_impl_str.rs: PayloadFieldIndex/Read for MapIndex<str> - payload_index_impl_int.rs: PayloadFieldIndex/Read for MapIndex<IntPayloadType> - payload_index_impl_uuid.rs: PayloadFieldIndex/Read for MapIndex<UuidIntType> - facet_index_impl.rs: FacetIndex impl for MapIndex<N> - value_indexer_impl.rs: ValueIndexer impls and value_retriever methods mod.rs retains the MapIndex enum, its core inherent methods, type aliases, constants, and tests. Co-authored-by: Cursor Agent <agent@cursor.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
89c488a497 |
refactor(field_index): make FieldIndexRead a supertrait of PayloadFieldIndexRead (#8996)
* refactor(field_index): make FieldIndexRead a supertrait of PayloadFieldIndexRead Removes the `get_payload_field_index_read() -> &dyn PayloadFieldIndexRead` bridge from `FieldIndexRead` and the five default impls that forwarded through it. The overlapping read methods now come from the supertrait directly, eliminating one layer of dynamic dispatch on the hot read path: `FieldIndex::filter` → variant match → concrete typed-index method. `FieldIndex` gains a direct `impl PayloadFieldIndexRead` block with per-method match arms, mirroring the existing dispatch shape used by `get_telemetry_data`, `values_count`, etc. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore: apply nightly rustfmt Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(numeric_index): split mod.rs into focused submodules (#8997) * refactor(numeric_index): split mod.rs into focused submodules `numeric_index/mod.rs` was 1308 lines mixing the storage-dispatch enum, the public `NumericIndex<T, P>` wrapper, three builders, per-(T, P) `ValueIndexer` impls, and the `Encodable`/`StreamRange` traits — hard to navigate. Split into: - `mod.rs` keeps the shared traits (`Encodable`, `StreamRange`), `Range<T>::as_index_key_bounds`, and re-exports. - `wrapper.rs` — `NumericIndex<T, P>` + inherent impl + `NumericIndexIntoInnerValue` trait. - `builders.rs` — `NumericIndexBuilder`, `NumericIndexMmapBuilder`, `NumericIndexGridstoreBuilder`. - `value_indexer.rs` — `ValueIndexer` and per-(T, P) `value_retriever` inherent impls. - `storage/` — the `NumericIndexInner` dispatch enum: - `storage/mod.rs` — enum + simple match-and-forward (constructors, lifecycle, telemetry, per-point access). - `storage/statistics.rs` — histogram-driven cardinality and point-count helpers. - `storage/trait_impls.rs` — `PayloadFieldIndex`, `PayloadFieldIndexRead`, `StreamRange` impls. Pure code reorganization — no behavior change. A few inherent methods on `NumericIndexInner` had to widen from private to `pub(super)` / `pub(in crate::index::field_index::numeric_index)` to remain reachable across the new module boundaries (and from `tests.rs`); the new constructors (`NumericIndexMmapBuilder::new`, `NumericIndexGridstoreBuilder::new`) replace direct field construction across files. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(numeric_index): rename wrapper.rs -> numeric_index.rs, move point_ids_by_value out of statistics - Renamed `wrapper.rs` to `numeric_index.rs`, matching the central `NumericIndex` type and the surrounding module name. - Moved `point_ids_by_value` from `storage/statistics.rs` to `storage/mod.rs` next to `get_values`. It is an exact value->points lookup primitive, not a cardinality estimate; the statistics module is left to the histogram-driven helpers. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(clippy): rename numeric_index submodule to index to fix module_inception lint Clippy's module_inception rule disallows a module with the same name as its containing module. Rename numeric_index.rs -> index.rs and update the three internal references. Co-authored-by: Cursor <cursoragent@cursor.com> * refactor(field_index): push PayloadFieldIndexRead through NumericIndex and move special_check_condition per-variant (#8998) Two related cleanups that share a theme — drop enum-level read dispatch in favor of per-variant trait impls: 1. `NumericIndex<T, P>` now implements `PayloadFieldIndexRead` directly (forwarding to its inner storage enum). The four numeric arms in `FieldIndex`'s `impl PayloadFieldIndexRead` drop their `.inner()` calls, so all eleven variants now use a uniform `idx.<method>(...)` form. 2. `special_check_condition` moves from `FieldIndexRead` to `PayloadFieldIndexRead` with a default `Ok(None)` body. `FullTextIndex` (the only variant with non-trivial logic) overrides it. `NumericIndex<T, P>` forwards through to the inner enum, and the `FieldIndex` enum's match dispatch moves out of `FieldIndexRead` into `PayloadFieldIndexRead` for consistency with the other five trait methods. Removed the now-redundant declaration from `FieldIndexRead` (inherited via the supertrait). Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Cursor Agent <agent@cursor.com> Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Cursor Agent <agent@cursor.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
47943f9b20 |
Integrate UniversalHashMap (#8954)
* Callback-based InvertedIndex::for_each_vocab_with_postings_len * Integrate UniversalHashMap into full-text index * Integrate UniversalHashMap into immutable_map_index |
||
|
|
a6934da495 |
refactor(field_index): extract FieldIndexRead impl into separate file (#8995)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
ca757cd13a |
refactor(read_view): make StructPayloadIndexReadView generic over F: FieldIndexRead (#8994)
The read view's `field_indexes` was concretely typed as `&IndexesMap = &HashMap<PayloadKeyType, Vec<FieldIndex>>`. Replace the FieldIndex with a generic parameter F: FieldIndexRead, so the view advertises a dependency on the trait surface rather than the concrete enum. `StructPayloadIndex::with_view` instantiates `F = FieldIndex`; any future consumer that implements FieldIndexRead (e.g. a read-only segment with a narrower field-index representation) can plug in directly. Changes: 1. StructPayloadIndexReadView gains a fourth generic parameter F bound by FieldIndexRead. `field_indexes` is now `&HashMap<PayloadKeyType, Vec<F>>`. All five impl blocks across read_view/ propagate the F generic. 2. `check_field_condition` / `select_nested_indexes` / `check_payload` in payload_storage/query_checker.rs become generic over `FI: FieldIndexRead` (and `R: AsRef<Vec<FI>>`). The bodies only call `special_check_condition` which is on FieldIndexRead, so the generalization is free. 3. `variable_retriever` in read_view/value_retriever/helpers.rs gains an F generic; the body already only uses FieldIndexRead methods after the previous commits in this PR. 4. `with_view` and `SegmentReadViewFor` instantiate F = FieldIndex. No external consumer breaks. 5. Tests construct the view with an inferred F. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |