* Migrate InvertedIndexCompressedMmap to UniversalRead
* Rehaul search_scratch.rs (was scores_memory_pool.rs)
- Use `typed_arena::Arena` instead of `bumpalo::Bump`.
Reason: bump don't drop.
Drawback: `Arena::new()` is not free, and `Arena::clear()` doesn't
exist; but this commit has workarounds.
- Use names that make more sense.
* Misc fixes
* InvertedIndexCompressedMmap: explicit S type parameter
* [ai] TQDT in the API
* [ai] unify new sparse error for TQDT
* Rename to `turbo4`
* Add `Turbo4` to comments and doc strings.
* Also validate named sparse vector creation
* ci(windows): skip IO-heavy tests that aren't OS-specific
On the Windows CI runner, several tests are 3-25x slower than on Ubuntu
purely due to slow filesystem IO. These tests exercise platform-agnostic
logic (optimizer, snapshot, WAL recovery, dedup, deferred points) and
are fully covered by the Linux and macOS jobs.
Mark them with `#[cfg_attr(target_os = "windows", ignore = "...")]` so:
- Windows CI skips them and finishes faster.
- They're still listed and runnable via `cargo test -- --ignored`
on Windows for local debugging.
Based on JUnit timings from CI run 26462785436 (PR #8827), this should
save ~5 minutes wall-clock on the Windows job, taking it closer to the
~13min Ubuntu and ~9min macOS jobs (currently 20m24s).
Tests affected:
- lib/wal: check_wal, check_last_index, check_clear, check_reopen,
check_truncate, check_prefix_truncate, test_prefix_truncate_parametric
- lib/edge/optimize: full tests module
- lib/segment deferred-point tests: read_operations,
dense_segment_combinations, sparse, facets
- lib/collection: snapshot_test, points_dedup, wal_recovery,
collection_test::test_ordered_read_api, snapshot_recovery_test
Co-authored-by: Cursor <cursoragent@cursor.com>
* revert(ci/windows): keep WAL and WAL-recovery tests on Windows
Reviewer correctly pointed out that WAL is mmap-backed and has
substantial Windows-specific code paths:
- Different segment allocation (fs4 vs rustix::ftruncate)
- Windows-specific delete_windows() with mmap-drop + retry loop
- Windows-specific sync_all() because directory fsync is unavailable
- Windows-specific lock proxy file (directories aren't lockable)
So those tests genuinely need Windows coverage. Reverted skips for:
- lib/wal/src/lib.rs: all check_* tests and test_prefix_truncate_parametric
- lib/collection/src/tests/wal_recovery_test.rs: all three tests
Still skipped on Windows (no OS-specific code in their production paths):
- lib/edge/optimize.rs (no cfg(windows) in source)
- lib/segment deferred-point tests (segment/ has no cfg(windows))
- lib/collection snapshot/dedup tests (collection/ has no cfg(windows))
- lib/collection integration snapshot_recovery + ordered_read_api
Co-authored-by: Cursor <cursoragent@cursor.com>
* revert(ci/windows): keep collection integration and snapshot_test
Per reviewer request, keep running these on Windows:
- lib/collection/tests/integration/* (snapshot_recovery_test,
collection_test::test_ordered_read_api)
- lib/collection/src/tests/snapshot_test.rs
These exercise higher-level collection/snapshot behavior that benefits
from cross-platform validation.
Remaining Windows skips (production code has no cfg(windows) branches):
- lib/edge/src/optimize.rs: 14 optimizer tests
- lib/segment/src/segment/tests/mod.rs: 4 deferred-point tests
- lib/collection/src/tests/points_dedup.rs: 2 dedup tests
Co-authored-by: Cursor <cursoragent@cursor.com>
* ci(windows): also skip HNSW/quantization integration tests
Per reviewer, also skip these segment integration test modules on Windows:
- hnsw_quantized_search_test::* (25 tests)
- multivector_filtrable_hnsw_test::* (rstest cases)
- multivector_quantization_test::* (rstest cases)
- byte_storage_quantization_test::* (rstest cases)
- payload_index_test::test_struct_payload_index_nested_fields
These exercise pure HNSW/quantization correctness on top of standard
segment IO that is already covered by tests we keep running on Windows.
Adds ~930s of sequential time to the Windows skip list, bringing the
expected wall-clock saving from ~3 min to ~8-10 min.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: serialize optimization proxy install against shard updates
execute_optimization captures `target_config` from the optimizer's frozen
config and then wraps source segments in proxies. Between those two points,
`CollectionUpdater::update` can apply a `CreateVectorName(V)` to the source
segments via `apply_segments`, leaving the optimizer with sources that have
V but a target_config that does not. The optimization then produces a
merged segment without V, and a follow-up optimization (running with the
refreshed config that includes V) fails to use that segment as a source:
"Cannot update from other segment because it is missing vector name X".
Close the race by extending the scope of the existing
`LockedSegmentHolder::acquire_updates_lock` to cover the proxy install
window. `CollectionUpdater::update` already takes this lock before
processing any shard update, so concurrent writers wait until proxies are
in place — at which point further mutations hit the proxies (recorded as
intent and propagated to the merged segment in `finish_optimization`)
instead of the originals. The guard is dropped right after proxy install so
the slow build phase does not extend it.
Tests:
- Three `SegmentBuilder::update` tests document the precondition the lock
now guarantees: with a target schema that adds a named vector the source
lacks, update errors with "missing vector name X". Quantized and
mixed-source variants exercise the same error path.
- `test_optimize_blocks_proxy_install_on_updates_lock` asserts the
invariant directly: while the updates lock is held, proxies are not yet
installed. Verified to fail when the new guard is removed (otherwise it
passes because `finish_optimization` also takes the same lock, so a
naive "did optimize finish?" check would not catch a missing proxy-install
guard).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(optimize): finish_optimization lock order; drop redundant tests
Address review of #9110:
1. `finish_optimization` was acquiring `upgradable_read` before
`acquire_updates_lock`, while the new guard at the start of
`execute_optimization` acquires them in the reverse order. With two
optimizer threads in flight, thread A in `finish_optimization` could
hold `upgradable_read` and wait on `updates_lock` while thread B at the
top of `execute_optimization` held `updates_lock` and waited on
`upgradable_read` (parking_lot allows only one upgradable reader),
deadlocking. Swap `finish_optimization` to take `updates_lock` first so
both halves agree.
2. Drop the quantized and mixed-source variants of the inverted
`SegmentBuilder` unit test — all three asserted the same error path
(the mismatch check fires before quantization training or per-source
branching), so only one is useful as documentation of the precondition.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: reject source-superset schema mismatch at SegmentBuilder
Drop the lock approach (deadlocked test_continuous_snapshot) and fix the bug
at the merge layer instead.
Snapshot's proxy_all_segments_and_apply acquires the segment_holder
upgradable_read first and then takes acquire_updates_lock tactically inside
the snapshot operation. The previous commits' lock-extension acquired
updates_lock before upgradable_read, so a snapshot in flight and an
optimization just entering execute_optimization could deadlock holding each
other's required next lock. Snapshot cannot easily reverse its order — that
would hold updates_lock for the entire snapshot duration, blocking all
writes.
Move the fix to where the actual harm happens: SegmentBuilder::update
iterates the target's vector_data and silently drops source vectors that
aren't in target. That silent drop is what produces the broken merged
segment in the CreateVectorName-vs-optimizer race. Add a check that every
source vector name is in the target schema; the optimization aborts cleanly
on mismatch and the next round (with refreshed config) merges correctly.
This is strictly stronger than the lock: the lock only closed the window
where V arrived *during* the proxy-install region. The schema check catches
both that window and the window where V's apply_segments completed before
the optimizer's lock acquisition.
Diff is contained to lib/segment; no locking changes, no cross-crate
plumbing. test_continuous_snapshot passes again.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(segment_builder): use Cancelled instead of ServiceError for schema mismatch
ServiceError flips the shard to RED status (via `report_optimizer_error` →
`segments.optimizer_errors`) and stays sticky until the next
`recreate_optimizers_blocking` clears it. That's the right shape for
hardware/IO failures but wrong for the schema-mismatch case here, which is
an expected, recoverable race outcome — the next optimizer round with a
refreshed target_config merges the same originals cleanly.
`Cancelled` is the variant the optimization worker treats as a recoverable
cancellation: logged at debug, tracker marked Cancelled, no
`report_optimizer_error` call, no RED status.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(segment_builder): also use Cancelled for the existing target-superset error
The existing "missing vector name" check at the start of the merge loop
also fires during a race — specifically the optimizer-vs-DeleteVectorName
shape, where V is removed from originals before J wraps proxies but J's
frozen target_config still has V. Like the new source-superset check, this
is an expected, recoverable race outcome, so use Cancelled instead of
ServiceError to avoid flipping the shard to RED.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor: move deferred-point ownership into the ID tracker
Re-implements the idea from #8512 against current `dev`.
Deferred-point state (`deferred_internal_id` + `deferred_deleted_count`)
moves out of `Segment.deferred_point_status` and the cached
`SparseVectorIndex.deferred_internal_id` field into `PointMappings`,
exposed through `IdTrackerRead`. The threshold is set once at
`MutableIdTracker::open` time; `PointMappings::drop` now maintains the
deleted counter inline (with double-delete protection), removing the
manual increment in `delete_point_internal` and the
`calculate_deleted_deferred_point_count` rescan.
Read paths consume the threshold through the id tracker:
- The segment read view drops the `deferred_point_status` field and
`with_view` no longer threads it in; `read_view/{deferred,info}.rs`
call `self.id_tracker.deferred_*()` directly.
- `SparseVectorIndex` no longer stores its own copy and its
`update_vector` / search debug-assert read from
`self.id_tracker.borrow().deferred_internal_id()`.
- `VectorQueryContext.deferred_internal_id` and the
`SegmentQueryContext::get_vector_context` parameter are gone; the
three downstream readers (`plain_vector_index`, sparse search, sparse
`update_vector`) consult their own id tracker.
`PointMappingsRefEnum` centralises the dispatch:
- `iter_internal_with_behavior(DeferredBehavior)` replaces ad-hoc
branches in `iter_filtered_points` impls.
- `external_iter_cutoff(DeferredBehavior)` covers iterators sourced
outside the mapping (field-index outputs in
`struct_payload_index::iter_filtered_points`).
- The internal `deferred_internal_id()` accessor is private; the raw
threshold no longer leaks to consumers.
- `iter_from_visible` / `iter_random_visible` read the mapping's own
threshold; callers that previously passed
`DeferredBehavior::apply(...)` now branch on
`deferred_behavior.include_all_points()` (scroll / order_by) or simply
drop the argument (sampling / facet).
`PayloadIndexRead::query_points` drops the now-redundant
`deferred_internal_id` parameter; `iter_filtered_points` takes
`DeferredBehavior` directly so HNSW build/search can request
`IncludeAll` while normal reads request `Exclude`.
RocksDB-related parts of the original PR are skipped — that tracker is
already gone from `dev`.
Tests adapted: sites that mutated `segment.deferred_point_status`
directly now construct a parallel non-deferred segment via
`create_deferred_segment(..., 0)` for comparison;
`test_deleted_deferred_point_count` reads counters through the id
tracker.
See `docs/plans/deferred-points-owned-by-id-tracker.md` for the design
write-up.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(benches): drop stale deferred_internal_id arg from query_points calls
The boolean / range / conditional bench files weren't built by
`cargo test -p segment`, so they slipped through. `cargo clippy
--workspace --all-targets` catches them.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor: drop id_tracker / point_mappings args from iter_filtered_points
Both impls already hold an id tracker on `self`:
- `StructPayloadIndexReadView` carries `id_tracker: &'a I`, so
`self.id_tracker.point_mappings()` borrows from `'a` and the lazy
iterator chain keeps working unchanged.
- `PlainPayloadIndex` carries `id_tracker: Arc<AtomicRefCell<...>>`,
where the mapping borrow is local; collect into a `Vec` and return
`into_iter()`. PlainPayloadIndex::iter_filtered_points has no direct
callers — only `query_points` was using it — so eager collection is a
non-issue.
While here, take `self` by value on `iter_internal_visible`,
`iter_from_visible`, `iter_random_visible`,
`iter_internal_with_behavior`, and `external_iter_cutoff`.
`PointMappingsRefEnum` is `Copy`; this matches the existing
`iter_internal` / `iter_from` / `iter_random` shape and lets the
iterator outlive a local `let point_mappings = ...;` binding.
The HNSW `condition_points` helper drops its now-unused `id_tracker`
parameter.
All callers (sampling, scroll, order_by, facet ×2, hnsw build/search)
just drop the two arguments.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore: ignore /docs/plans/ and untrack the previously-committed plan
`docs/plans/` is a scratch directory for per-feature planning notes —
not something we want under source control. Add it to `.gitignore` and
drop the deferred-points plan that slipped into history; the design is
captured in the PR description.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor: replace external_iter_cutoff with filter_deferred iterator wrapper
Instead of exposing a raw `Option<PointOffsetType>` cutoff that every
caller has to apply with their own `.filter(...)`, give
`PointMappingsRefEnum` an iterator wrapper:
fn filter_deferred<I: Iterator<Item = PointOffsetType>>(
self,
iter: I,
deferred_behavior: DeferredBehavior,
) -> impl Iterator<Item = PointOffsetType>
It returns the iterator unchanged for `IncludeAll` (or when the mapping
has no threshold) and otherwise wraps it in a cutoff `.filter`,
dispatched via `itertools::Either` so the no-cutoff path stays
allocation-free.
The struct payload index's `iter_filtered_points` swaps its open-coded
filter for a single `point_mappings.filter_deferred(...)` call. The
deferred threshold no longer leaks out of `PointMappingsRefEnum`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor: move deferred wrapping out of peek_top_all, gate it as test-only
`BatchFilteredSearcher::peek_top_all` baked the deferred cutoff into
its iterator construction, which was the last place outside
`PointMappingsRefEnum` that knew about the threshold. Split the
deleted-iteration concern out into a new accessor:
fn iter_not_deleted(&self) -> impl Iterator<Item = PointOffsetType> + 'a
It borrows `&'a BitSlice` directly (not via `&self`), so callers can
chain `filter_deferred` and then move `self` into `peek_top_iter`
without lifetime conflicts.
Sparse + plain vector index call sites now do:
let iter = id_tracker
.point_mappings()
.filter_deferred(searcher.iter_not_deleted(), DeferredBehavior::Exclude);
searcher.peek_top_iter(iter, &is_stopped)
leaving `BatchFilteredSearcher` completely ignorant of deferred state.
With deferred handling lifted out, `peek_top_all` itself is now used
only by tests (3 inline `#[cfg(test)] mod tests`, 1 integration test,
1 bench) — gate it under `#[cfg(feature = "testing")]` to match
`new_for_test`. Production code goes through the
`iter_not_deleted` + `filter_deferred` + `peek_top_iter` composition.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor: optimize Segment::retrieve and thread user data through read_vectors
Two interlocking changes that together collapse the per-point lookups
and intermediate allocations in `Segment::retrieve` down to one
external-to-internal pass.
## `IdTrackerRead::resolve_external_ids` (new default trait method)
Single-pass translation of a `&[PointIdType]` slice into two parallel
vectors `(Vec<PointIdType>, Vec<PointOffsetType>)`. Folds deferred
filtering (compare offset against the threshold inline — no separate
`point_is_deferred` lookup) and missing-id errors (eager
`PointIdError`) into resolution.
Lives on the trait so the deferred threshold never leaks out of the id
tracker; the parallel-vector shape lets a future batched payload /
vector fetcher consume `&offsets` straight without unzipping. The
`appendable_flag` guard previously in `point_is_deferred` is gone:
non-appendable trackers always carry `deferred_internal_id() == None`
(set only via `MutableIdTracker::open`, guarded by the segment
constructor), so the check was load-bearing nowhere.
## User-data threading through `read_vectors`
`VectorStorageRead::read_vectors` now takes
`IntoIterator<Item = (U, PointOffsetType)>` and yields
`(U, PointOffsetType, CowVector)`. The user-data tag rides alongside
each offset all the way through, so callers can map results back into a
parallel array without keeping a separate `offset → ...` lookup table.
- Default trait impl: one-line per-key loop.
- Dense impl: `unzip()` into parallel `(Vec<U>, Vec<PointOffsetType>)`
in a single pass — same allocation count as before, just U riding
alongside.
- Enum delegations (`VectorStorageEnum`, `VectorStorageReadEnum`)
forward unchanged.
- `for_each_in_batch` and below stay untouched.
`SegmentReadView::vectors_by_offsets<U: Copy>` becomes a lazy filter
chain — no parallel `Vec<(orig_idx, offset)>` allocation. The dead
`SegmentReadView::read_vectors` helper is removed.
## `Segment::retrieve` end-to-end
Per N points / V vectors / payload:
| Operation | Before | After |
|----------------------------|---------------------|-------|
| `id_tracker.internal_id` | N × (1 + V + 1) | N |
| `id_tracker.external_id` | N × V | 0 |
| `point_is_deferred` | N (when applicable) | 0 |
| `offset_to_id` HashMap | N entries | none |
| `Vec` in `vectors_by_offsets` | 1 | 0 |
The vectors stage passes the external id as `read_vectors`'s user
data — the callback gets `id` directly without any index lookup.
The payload stage uses `payload_by_offset` against the already-resolved
offsets. The shape is also batch-friendly: swapping in a future
`IdTrackerRead::batch_internal_id` or `payload_index.batch_get_payload`
needs no changes outside the two call sites.
## Behavioural notes
- Missing-id now errors eagerly inside resolution, instead of in the
vectors stage (`WithVector::Bool(true)` / `Selector`) or payload
stage (`with_payload.enable`). The previous `WithVector::Bool(false)`
+ no-payload path silently inserted an empty record; that is now also
an error. None of the existing callers (search post-processing,
external retrieve API, the deferred-points test on tests/mod.rs:1179)
pass non-existent ids.
- Added a per-payload `check_stopped`; the vectors stage already had
`stop_if` on its iterator chain.
- `vector_by_offset` (the single-element helper) passes `()` as the
no-op user data.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* style: apply rustfmt to optimised retrieve / read_vectors paths
Pre-push hook failure on the previous commit was rustfmt. Same content,
formatted.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* do not error out on missing points in retrieve
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(field_index): make FieldIndexRead a supertrait of PayloadFieldIndexRead
Removes the `get_payload_field_index_read() -> &dyn PayloadFieldIndexRead`
bridge from `FieldIndexRead` and the five default impls that forwarded
through it. The overlapping read methods now come from the supertrait
directly, eliminating one layer of dynamic dispatch on the hot read path:
`FieldIndex::filter` → variant match → concrete typed-index method.
`FieldIndex` gains a direct `impl PayloadFieldIndexRead` block with
per-method match arms, mirroring the existing dispatch shape used by
`get_telemetry_data`, `values_count`, etc.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore: apply nightly rustfmt
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(numeric_index): split mod.rs into focused submodules (#8997)
* refactor(numeric_index): split mod.rs into focused submodules
`numeric_index/mod.rs` was 1308 lines mixing the storage-dispatch enum,
the public `NumericIndex<T, P>` wrapper, three builders, per-(T, P)
`ValueIndexer` impls, and the `Encodable`/`StreamRange` traits — hard
to navigate.
Split into:
- `mod.rs` keeps the shared traits (`Encodable`, `StreamRange`),
`Range<T>::as_index_key_bounds`, and re-exports.
- `wrapper.rs` — `NumericIndex<T, P>` + inherent impl + `NumericIndexIntoInnerValue` trait.
- `builders.rs` — `NumericIndexBuilder`, `NumericIndexMmapBuilder`, `NumericIndexGridstoreBuilder`.
- `value_indexer.rs` — `ValueIndexer` and per-(T, P) `value_retriever` inherent impls.
- `storage/` — the `NumericIndexInner` dispatch enum:
- `storage/mod.rs` — enum + simple match-and-forward (constructors, lifecycle, telemetry, per-point access).
- `storage/statistics.rs` — histogram-driven cardinality and point-count helpers.
- `storage/trait_impls.rs` — `PayloadFieldIndex`, `PayloadFieldIndexRead`, `StreamRange` impls.
Pure code reorganization — no behavior change. A few inherent methods
on `NumericIndexInner` had to widen from private to `pub(super)` /
`pub(in crate::index::field_index::numeric_index)` to remain reachable
across the new module boundaries (and from `tests.rs`); the new
constructors (`NumericIndexMmapBuilder::new`,
`NumericIndexGridstoreBuilder::new`) replace direct field
construction across files.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(numeric_index): rename wrapper.rs -> numeric_index.rs, move point_ids_by_value out of statistics
- Renamed `wrapper.rs` to `numeric_index.rs`, matching the central
`NumericIndex` type and the surrounding module name.
- Moved `point_ids_by_value` from `storage/statistics.rs` to
`storage/mod.rs` next to `get_values`. It is an exact value->points
lookup primitive, not a cardinality estimate; the statistics module
is left to the histogram-driven helpers.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(clippy): rename numeric_index submodule to index to fix module_inception lint
Clippy's module_inception rule disallows a module with the same name as
its containing module. Rename numeric_index.rs -> index.rs and update
the three internal references.
Co-authored-by: Cursor <cursoragent@cursor.com>
* refactor(field_index): push PayloadFieldIndexRead through NumericIndex and move special_check_condition per-variant (#8998)
Two related cleanups that share a theme — drop enum-level read dispatch
in favor of per-variant trait impls:
1. `NumericIndex<T, P>` now implements `PayloadFieldIndexRead` directly
(forwarding to its inner storage enum). The four numeric arms in
`FieldIndex`'s `impl PayloadFieldIndexRead` drop their `.inner()`
calls, so all eleven variants now use a uniform `idx.<method>(...)`
form.
2. `special_check_condition` moves from `FieldIndexRead` to
`PayloadFieldIndexRead` with a default `Ok(None)` body.
`FullTextIndex` (the only variant with non-trivial logic) overrides
it. `NumericIndex<T, P>` forwards through to the inner enum, and
the `FieldIndex` enum's match dispatch moves out of `FieldIndexRead`
into `PayloadFieldIndexRead` for consistency with the other five
trait methods. Removed the now-redundant declaration from
`FieldIndexRead` (inherited via the supertrait).
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Cursor Agent <agent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Cursor Agent <agent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* refactor: implement FieldIndexRead for FieldIndex
Move the 10 read-only inherent methods on FieldIndex (special_check_condition,
count_indexed_points, filter, estimate_cardinality, for_each_payload_block,
get_telemetry_data, values_count, values_is_empty, as_numeric, as_facet_index)
into impl FieldIndexRead for FieldIndex. Bodies are unchanged.
Write/lifecycle methods (add_point, remove_point, wipe, flusher, files,
immutable_files, ram_usage_bytes, is_on_disk, populate, clear_cache,
get_full_index_type) stay inherent on FieldIndex.
Call sites that hold &FieldIndex now resolve through the trait — added
`use ... FieldIndexRead;` imports where the compiler asked
(read_view/{payload_index_read,filtering}.rs, query_optimization/
condition_converter.rs and match_converter.rs, payload_storage/
query_checker.rs, two integration tests, full_text_index tests module).
StructPayloadIndex::get_facet_index keeps its concrete
OperationResult<FacetIndexEnum<'_>> return: the trait method as_facet_index
returns Option<impl FacetIndex + '_> (RPITIT) which can't be downcast back
to FacetIndexEnum, so the function inlines the variant match. Same shape as
before, just expressed locally.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor: remove unused StructPayloadIndex::get_facet_index
Confirmed dead code — the function had no callers anywhere in the workspace.
It was preserved across recent refactors (most recently #8967) but was never
referenced. Removing it also drops the now-unused JsonPath and FacetIndexEnum
imports in struct_payload_index/mod.rs.
This obsoletes the inline variant match introduced in the previous commit:
that match existed solely to keep get_facet_index returning the concrete
FacetIndexEnum<'_> when as_facet_index moved to the FieldIndexRead trait
with an opaque return. With get_facet_index gone, no workaround is needed.
The OperationError::MissingMapIndexForFacet variant stays — it is still used
by segment::read_view::facet and lib/collection/src/operations/types.rs.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(index): drop PayloadIndex: PayloadIndexRead super-trait
`PayloadIndex` now only declares the mutating surface; reads live on
the sibling `PayloadIndexRead` trait. The previous super-trait
relationship forced any type that implemented `PayloadIndex` to also
implement `PayloadIndexRead`, blocking a future `PayloadIndexRead`-only
view that doesn't (and shouldn't) own the writable index machinery.
No behavioural change. Audit before committing showed no generic
bound site on `PayloadIndex` exists in the workspace, and every
caller that uses read methods already imports `PayloadIndexRead`
explicitly (the trait was already used as a generic bound on
`SegmentReadView`'s `TPayloadIndex` parameter and on
`iter_filtered_points`). The full workspace builds clean and all
segment / storage / collection tests pass without any consumer
update.
Doc comment on `PayloadIndex` updated to point readers at
`PayloadIndexRead` for the read surface.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(index): introduce StructPayloadIndexReadView<P, I, V> (#8970)
Move the read surface of `StructPayloadIndex` onto a new borrowed view
struct generic over `<P: PayloadStorageRead, I: IdTrackerRead, V: VectorStorageRead>`.
The view holds exactly the fields that `PayloadIndexRead` requires --
no more, no less:
pub struct StructPayloadIndexReadView<'a, P, I, V> {
payload: &'a Arc<AtomicRefCell<P>>,
id_tracker: &'a I,
vector_storages: &'a HashMap<VectorNameBuf, Arc<AtomicRefCell<V>>>,
field_indexes: &'a IndexesMap,
config: &'a PayloadConfig,
visited_pool: &'a VisitedPool,
}
`StructPayloadIndex` now exposes a `with_view(|v| ...)` accessor that
borrows `id_tracker` once at the top and constructs the view for the
closure scope. All read-method bodies move onto the view, which is the
sole `PayloadIndexRead` implementor for this index.
Why three generics
==================
- `P: PayloadStorageRead` -- already generic via PR #8968.
- `I: IdTrackerRead` -- direct method calls; held as `&I` (not
`&Arc<AtomicRefCell<I>>`) because the cell is collapsed at the
`with_view` boundary, saving a per-method `borrow()` atomic op.
`dyn IdTrackerRead` does not satisfy `I: IdTrackerRead` bounds in
Rust without an explicit blanket impl, so generic is the only
consistent option here.
- `V: VectorStorageRead` -- the only access site is
`available_vector_count()` for the `HasVector` cardinality branch
(`condition_cardinality` in `read_view/filtering.rs`).
Why `payload` keeps the `Arc`
=============================
`PayloadProvider<P>::new(...)` (introduced in PR #8968) takes
`Arc<AtomicRefCell<P>>` so that the returned `FormulaScorer<'q>` /
`Box<dyn FilterContext + 'a>` can outlive the caller frame. The view
therefore holds `&'a Arc<AtomicRefCell<P>>` (asymmetric vs the bare
`&I` for `id_tracker`). Switching to a borrow-based provider would
require reworking `formula_scorer` / `filter_context` to callback
style; deferred to a follow-up if needed.
What does NOT move
==================
- `build_field_indexes` and `clear_index_for_point` stay on
`StructPayloadIndex`. `build_field_indexes` is read-shaped but only
has write-side callers, and pulls in the `selector` machinery which
uses `path` + `storage_type`. Keeping it on the writable struct
means `path` and `is_appendable` do not need to leak into the view.
- The `selector` / `selector_with_type` helpers stay on the writable
struct for the same reason.
- The free helpers in `query_optimization/condition_converter.rs`
(range / geo / null / is-empty checkers) stay where they are; their
visibility is bumped from `fn` to `pub(in crate::index)` so the view
can still call them.
Module layout
=============
lib/segment/src/index/struct_payload_index/
mod.rs # owning struct + with_view
build.rs # write-side build coordination
payload_index.rs # impl PayloadIndex (mutating only)
tests.rs
read_view/
mod.rs # view struct + module wiring
payload_index_read.rs # impl PayloadIndexRead for view
filtering.rs # struct_filtered_context, condition_cardinality, query_field, estimate_field_condition
condition_converter.rs # impl block from query_optimization/
optimizer.rs # impl block from query_optimization/
value_retriever.rs # impl block from query_optimization/
tests.rs # smoke test that builds the view directly
Consumer migration
==================
- `Segment::with_view` nests the new `StructPayloadIndex::with_view`
inside it; `SegmentReadViewFor<'s>` uses the view as its
`TPayloadIndex` parameter.
- HNSW (`hnsw.rs`), sparse (`sparse_vector_index.rs`), plain
(`plain_vector_index.rs`) call sites wrap their read-method calls in
`payload_index.borrow().with_view(|v| ...)`.
- `Segment::get_indexed_fields`, `update_all_field_indices`, and
`SegmentBuilder::build` switch to `with_view` for `indexed_fields()`
/ `get_payload_sequential()`.
- Integration tests and benches similarly migrate.
- `set_payload` (still on `PayloadIndex` write impl) inlines its
former `self.get_payload(...)` call as `self.payload.borrow().get(...)`
to avoid going through `with_view` from a `&mut self` write path.
Smoke test (`read_view/tests.rs`) constructs the view directly over
`InMemoryPayloadStorage` + `InMemoryIdTracker` + an empty vector-storage
map, and exercises `indexed_fields()`, `query_points()`, and
`available_point_count()` -- proving the view is genuinely decoupled
from `StructPayloadIndex`. This is the abstraction PR 4 will use to
wire a read-only segment.
Verified
========
- `cargo build --workspace --tests --benches` -- green
- `cargo test -p segment --lib` -- 666 passed (665 + 1 new smoke test), 0 failed
- `cargo test -p segment --tests` -- 120 integration tests pass
- `cargo test -p storage --lib` -- 44 passed
- `cargo test -p collection --lib` -- 197 passed
- `cargo clippy -p segment --tests --benches` -- clean
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Step 5 leftover: with PayloadIndexRead.iter_filtered_points now on the
trait, filtered_read_by_index can move to the view alongside the other
three scroll helpers.
* `read_view/scroll.rs` gains `filtered_read_by_index` and the
`read_filtered` orchestrator.
* `segment/scroll.rs` is deleted entirely; `mod scroll;` removed from
`segment/mod.rs`.
* `Segment::read_filtered` collapses to a single
`with_view(|v| v.read_filtered(...))` delegator.
* `scroll_filtering_test.rs` integration test calls
`filtered_read_by_index` via `with_view`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Step 5 of the SegmentReadView migration.
New `read_view/scroll.rs` module hosts:
* `should_pre_filter` — payload-index cardinality estimation,
used by all three scroll-shaped trait methods (read_filtered,
read_ordered_filtered, read_random_filtered).
* `read_by_id_stream` — streamed enumeration of visible points.
* `filtered_read_by_id_stream` — streamed enumeration with a
payload-filter context applied.
`Segment`-side `read_filtered` collapses to a `with_view` orchestrator
(except for the `filtered_read_by_index` branch, see below).
`read_ordered_filtered` and `read_random_filtered` now route their
`should_pre_filter` calls through `with_view` (their own bodies migrate
in steps 6 and 7).
`filtered_read_by_index` stays on `Segment` for now: it depends on
`StructPayloadIndex::iter_filtered_points`, which is an inherent method
that takes the concrete `&IdTrackerEnum` and returns `impl Iterator`.
Migrating it cleanly requires extending `PayloadIndexRead` with a
trait-object-friendly version of that method, which is the same
prerequisite Steps 6/7/8 will need (sampling, order_by, facet all use
`iter_filtered_points`). I will do that as a focused pre-step before
Step 6.
`deferred_internal_id` / `deferred_deleted_count` view helpers bumped
from private to `pub(super)` so the new `scroll` module can call them.
`scroll_filtering_test.rs` integration test updated to call
`filtered_read_by_id_stream` via `with_view`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Read-only methods now live on a separate IdTrackerRead trait, with the
mutating IdTracker trait extending it. This lets read-only call sites
depend only on the read API.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* [ai] Replace manual into mappings with Into::into
* Reformat
* [ai] Use implicit .iter
* Don't iterate over keys too
* [ai] Replace unwrap_or
* Reformat
* [ai] Use as_deref and then_some
* [ai] Use more to_string
* [ai] Use explicitly typed into conversions
* Reformat
* [ai] More explicit into conversions
* Reformat
* measure filterable hnsw indexed vs unindexs + cpu vs gpu
* measure filterable hnsw indexed vs unindexs + cpu vs gpu
* limit GPU parallelism
* address review
* Propagate HardwareCounterCell through MmapMapIndex::get_values
Previously, `MmapMapIndex::get_values` used `ConditionedCounter::never()`,
which silently skipped all hardware IO counter tracking for mmap map index
value reads. This propagates a real `HardwareCounterCell` through the full
call chain so that disk IO from `values_iter` is properly measured.
Changes:
- `MmapMapIndex::get_values` now accepts `&HardwareCounterCell`, creates a
`ConditionedCounter` via `make_conditioned_counter`, and measures the
`deleted` bitvec access (matching `check_values_any` behavior).
- `MapIndex::get_values` forwards the counter to the `Mmap` variant.
- `FacetIndex::get_point_values` trait method now accepts `&HardwareCounterCell`,
propagated through `FacetIndexEnum` and both impls (MapIndex, BoolIndex).
- `indexed_variable_retriever` accepts and forwards the counter to
`IntMapIndex`, `KeywordIndex`, and `UuidMapIndex` calls.
- `SegmentBuilder::update` and `_get_ordering_value` accept and propagate
the counter instead of creating disposable instances.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat: add test
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Daniel Boros <dancixx@gmail.com>
* Skip broken tests
* Add rocksdb dropper
This is a temporary tool. It will be removed later in this PR.
* [automated] Drop rocksdb
This commit is made by running tools/rocksdb/drop.sh
* Touch-up after ast-grep
The previous automated commit removed items, but not their comments.
Also, some blocks are left with only one item.
Also, ast-grep can't handle macros like `vec![]`.
This commit completes the job.
* Remove RocksDB dropper
* Remove mentions of rocksdb feature in CI
* Fix clippy warnings
These `FIXME` comments added previously in this PR by ast-grep by
"peeling" cfg_attr like this:
-#[cfg_attr(not(feature = "rocksdb"), expect(...))]
+#[expect(...)] // FIXME(rocksdb): ...
This commit removes these allow/expect attributes and fixes clippy
lints.
* Remove leftover rocksdb-related code
Removed:
- Cargo.toml: rocksdb cargo feature flag and dependencies.
- code: rocksdb-related items that was not under feature flag.
- flags.rs: rocksdb-related qdrant feature flags.
Disabled:
- Some benchmarks because these do not compile now.
* Print warning if rocksdb leftovers found in snapshots
* Remove Clone from tokenizer
* Regenerate openapi.json
* Don't insert deferred points into sparse index
# Conflicts:
# lib/segment/src/segment/tests.rs
# lib/segment/src/segment_constructor/segment_constructor_base.rs
# Conflicts:
# lib/segment/src/segment_constructor/segment_constructor_base.rs
* Clippy
* Assert consistency of deferred_internal_id var
* Return before SparseVector conversion in case of deferred point
* Use debug_assert instead
* Refactor: `SearchSegmentEntry` for search ops
* Move `has_deferrend_points_method`
to `SearchSegmentEntry`
* Reorg imports
* Renamings per review
* fixup! Renamings per review
* `StorageSegmentEntry` trait
It contains methods dealing with storage and syncronization
* Move index methods from `ReadSegmentEntry`
to `SegmentEntry`
* Add `_concurrent` suffix to some methods
* move field index methods into NonAppendableSegmentEntry
---------
Co-authored-by: generall <andrey@vasnetsov.com>
* UIO traits: require Sized
I don't think we ever going to add unsized implementation. So, lets,
require it on the whole trait rather than on individual methods.
* Move `AccessPattern` to the `common` package
Was: segment::vector_storage::vector_storage_base::AccessPattern
Now: common::generic_consts::AccessPattern
* UIO traits: replace `bool` with `AccessPattern`
Instead of `IdTracker` to have `iter_*` methods, it now has a `point_mappings()` method, returning a guard to `PointMapping` value. It prepares for moving the point_mappings under a lock to allow parallel searches and deletions.
Also, an incorrect unsafe in `Segment::iter_points` is replaced by a safe version with `self_cell`.
Reduce test parameters across 5 areas to cut test runtime without
sacrificing coverage:
1. HNSW PQ tests: lower vector dimensionality from 131 to 64 for product
quantization variants, cutting PQ codebook training time (~31s -> ~8s)
2. HNSW index build: use 2 threads instead of 1 for index construction;
the tests only check accuracy above a threshold so determinism is not
required
3. Search attempts: reduce from 10 to 5 query vectors per test
4. Rescoring: skip the expensive re-upsert-all-as-zeros rescoring check
for PQ tests (already covered by scalar quantization variants)
5. Near-miss speedups:
- WAL: reduce QuickCheck iterations from 100 to 50
- Gridstore: halve operation counts, drop 64-byte block size case,
reduce proptest cases for gap search
- Continuous snapshot: reduce timeout from 20s to 10s
Made-with: Cursor
Co-authored-by: Cursor Agent <agent@cursor.com>
* Add ignore_deferred flag for internal scroll API
* Ignore deferred points in ProxyShard::update
* Use enum instead of bool as type in function parameters
* Use consistent parameter name
* fmt
* Apply suggestions from code review
Co-authored-by: Tim Visée <tim+github@visee.me>
* More explicit naming
* Add DeferredBehavior to `count()` and disable filtering in stream_records
* Don't filter `retrieve()` everywhere
* Use consistent filter behavior in read operations
* Add TODO for ignored parameter in remote_shard
* Rebase fixes
---------
Co-authored-by: timvisee <tim@visee.me>
Co-authored-by: Tim Visée <tim+github@visee.me>
* Implement filtering: Search
* Add tests for search and read_filtered
* Implement filtering: read_ordered_filtered + Tests
* Codespell + Clippy
* Reduce code duplication in tests
* Use helper methods of Id tracker
* Implement filtering: Random Scroll
* fixes after rebase
* [Unittest] Only set deferred ID if deferred points exist
* [Unit test] Reduce combinations in segment creation test
---------
Co-authored-by: Ivan Pleshkov <pleshkov.ivan@gmail.com>
* Introduce EdgeShardConfig for edge shard
- Add EdgeShardConfig and EdgeOptimizersConfig in lib/edge/src/config.rs
- Segment config (vector_data, sparse_vector_data, payload_storage_type)
- Global hnsw_config and per-vector HNSW in segment config
- Optimizer params: deleted_threshold, vacuum_min_vector_number,
default_segment_number, max_segment_size, indexing_threshold,
prevent_unoptimized (excludes memmap_threshold, flush_interval_sec,
max_optimization_threads)
- Persist/load as edge_config.json in shard path
- EdgeShard uses RwLock<EdgeShardConfig>; load() accepts Option<EdgeShardConfig>,
falls back to file or infer from segments; compatibility checked on load
- load_with_segment_config() for backward compatibility (SegmentConfig -> EdgeShardConfig)
- optimize() uses EdgeShardConfig for hnsw and optimizer thresholds
- Public methods: set_hnsw_config(), set_vector_hnsw_config(), set_optimizers_config()
(update and persist)
- Python and examples use load_with_segment_config with existing config API
Made-with: Cursor
* Refactor EdgeShardConfig: user-facing params only, config module
- Replace SegmentConfig inside EdgeShardConfig with user-facing fields:
- on_disk_payload (bool) instead of payload_storage_type
- vectors: HashMap<VectorNameBuf, EdgeVectorParams> with on_disk per vector,
no per-vector quantization; global quantization_config only
- sparse_vectors: HashMap<VectorNameBuf, EdgeSparseVectorParams> with on_disk
- EdgeVectorParams / EdgeSparseVectorParams use on_disk (bool) instead of
storage_type; conversion to VectorDataConfig/SparseVectorDataConfig in
to_segment_config()
- Add config module: mod.rs, optimizers.rs, vectors.rs, shard.rs
- from_segment_config(&SegmentConfig) fills all inferrable params
- to_segment_config() builds SegmentConfig for segments and optimize()
- load_with_segment_config takes Option<SegmentConfig>, uses from_segment_config
Made-with: Cursor
* Move optimizer threshold helpers to shard crate
- Add get_number_segments, get_indexing_threshold_kb, get_max_segment_size_kb,
get_deferred_points_threshold_bytes in shard::optimizers::config
- Collection OptimizersConfig and edge EdgeOptimizersConfig delegate to these
- Single place for threshold logic; collection and edge use shard helpers
Made-with: Cursor
* Use destructuring in config conversions to avoid missing new fields
- EdgeVectorParams: destructure VectorDataConfig in from_*, destructure self in to_vector_data_config
- EdgeSparseVectorParams: destructure SparseVectorDataConfig and SparseIndexConfig in from_*, destructure self in to_sparse_vector_data_config
- EdgeShardConfig: destructure SegmentConfig in from_segment_config, destructure self in to_segment_config
Adding new fields to source structs will now cause compile errors until conversions are updated.
Made-with: Cursor
* refactor: centralize on_disk_payload→payload_storage_type, on_disk→storage_type, and appendable quantization logic
- PayloadStorageType::from_on_disk_payload(bool) in segment (Mmap/InRamMmap)
- VectorStorageType::from_on_disk(bool) in segment (ChunkedMmap/InRamChunkedMmap)
- QuantizationConfig::for_appendable_segment(Option<&Self>) in segment (feature flag + supports_appendable)
- collection: use from_on_disk_payload in non-rocksdb branch
- edge shard/vectors: use new helpers; remove duplicated conditionals
- shard optimizers: use from_on_disk and for_appendable_segment
Made-with: Cursor
* refactor(edge): use EdgeShardConfig directly, drop segment_config
- Add plain_segment_config() for create_appendable_segment (no HNSW)
- Add segment_optimizer_config() built from EdgeShardConfig for blocking optimizers
- Add vector_data_config(name) for query/MMR
- build_blocking_optimizers: use segment_optimizer_config() instead of SegmentConfig
- create_appendable_segment: use plain_segment_config()
- search/query: use config().vectors and vector_data_config() instead of segment_config()
- Remove segment_config() from EdgeShardConfig and EdgeShard
- Add to_plain_vector_data_config on EdgeVectorParams
Made-with: Cursor
* [manual] review changes
* refactor(edge-py): wrap EdgeShardConfig, add EdgeVectorParams/EdgeSparseVectorParams
- PyEdgeConfig now wraps EdgeShardConfig (vectors, sparse_vectors, on_disk_payload, etc.)
- PyEdgeVectorParams / PyEdgeSparseVectorParams wrap edge config types
- PyEdgeOptimizersConfig for optional optimizer settings
- EdgeShard.load() uses EdgeShardConfig; edge::config made pub for Python crate
- cargo fmt + clippy (remove map_identity)
Made-with: Cursor
* refactor(edge-py): simplify config API, remove unused Py* types, add EdgeConfig
- Remove unused PyPayloadStorageType, PyVectorDataConfig, PyVectorStorageType,
PySparseVectorDataConfig, PySparseVectorStorageType from Python bindings
- Move PyEdgeOptimizersConfig to lib/edge/python/src/config/optimizers.rs
- Update qdrant_edge.pyi: EdgeConfig with vectors/sparse_vectors,
EdgeVectorParams, EdgeSparseVectorParams, EdgeOptimizersConfig
- Update examples (common.py, repr.py) to use new config API
- Run cargo fmt
Made-with: Cursor
* [manual] review changes
* [manual] review changes
* [manual] fix test
* Address CodeRabbit review comments for PR 8322 (#8324)
* Address CodeRabbit review comments for PR 8322
- Python examples: explicit imports (repr.py, common.py) and new EdgeConfig API
- HnswIndexConfig: add max_indexing_threads param and property in .pyi and Rust bindings
- EdgeConfig: make vectors optional for sparse-only configs; validate at least one of vectors/sparse_vectors
- EdgeShardConfig::load: use try_exists(), propagate I/O errors
- from_segment_config: infer hnsw_config from per-vector HNSW when all agree
- EdgeShard setters: atomic clone-mutate-save-then-replace; persist config save errors
- Segment compat: prefix vector name in error messages; resolve None datatype to Float32
- max_indexing_threads: preserve 0 (auto) sentinel in trait default; remove per-optimizer overrides
- SegmentOptimizerConfig:🆕 build plain and optimizer maps in single pass
- config_mismatch_optimizer tests: use VectorNameBuf::from() instead of .into()
- vectors.rs: doc updates for per-vector quantization
Made-with: Cursor
* Address @generall review: SaveOnDisk for config, resolve num_rayon_threads in optimizer
- Use SaveOnDisk<EdgeShardConfig> for EdgeShard config (generall: 'We have SaveOnDisk struct for this')
- Create via SaveOnDisk::new() after resolving config; setters use .write() for atomic persist
- set_vector_hnsw_config: clone then mutate then write (fallible setter)
- max_indexing_threads: resolve 0 (auto) via num_rayon_threads inside impl (generall: 'proper solution would be to resolve num_rayon_threads inside the optimizer impl')
- max_indexing_threads_sentinel_aware() now returns Some(num_rayon_threads(raw)) so callers get actual thread count
Made-with: Cursor
* [manual] reorganize num_rayon_threads -> get_num_indexing_threads to better account per-vector configuration
---------
Co-authored-by: Cursor Agent <agent@cursor.com>
Co-authored-by: generall <andrey@vasnetsov.com>
* update docstring and pyi
* fmt
* fmt
* clipy
---------
Co-authored-by: Cursor Agent <agent@cursor.com>
Co-authored-by: generall <andrey@vasnetsov.com>
* Use `IdTrackerEnum` type instead of `dyn IdTracker`
It would allow to be more flexible on the IdTracker trait, making it
dyn-incompatible eventually.
Coauthored with Claude Code.
* Review fixes
* Unify parking_lot/arc_lock feature
* Move lib/common/{io,memory}/* -> lib/common/common/*
- Mmap-related items are grouped into `common::mmap` sub-module:
- `memory/src/chunked_utils.rs` -> `common/src/mmap/chunked.rs`
- `memory/src/madvise.rs` -> `common/src/mmap/advice.rs`
- `memory/src/mmap_ops.rs` -> `common/src/mmap/ops.rs`
- `memory/src/mmap_type_readonly.rs` -> `common/src/mmap/mmap_readonly.rs`
- `memory/src/mmap_type.rs` -> `common/src/mmap/mmap_rw.rs`
- Filesystem-related items are grouped into `common::fs` sub-module:
- `common/src/fs.rs` -> `common/src/fs/sync.rs`
- `io/src/file_operations.rs` -> `common/src/fs/ops.rs`
- `io/src/move_files.rs` -> `common/src/fs/move.rs`
- `io/src/safe_delete.rs` -> `common/src/fs/safe_delete.rs`
- `memory/src/checkfs.rs` -> `common/src/fs/check.rs`
- `memory/src/fadvise.rs` -> `common/src/fs/fadvise.rs`
- Rest is moved straight into `common`:
- `io/src/storage_version.rs` -> `common/src/storage_version.rs`
The old `io` and `memory` are now hollow crates that re-export items
from `common`. These hollow crates will be removed in next commits.
* Replace uses of `io` and `memory` with new paths in `common`
Since `io` and `memory` are just re-exports of `common`, these
replacements are no-op.
* Remove `io` and `memory` crates
* Split `SegmentEntry`
Add `ImmutableSegmentEntry` for operations that can be applied to
immutable segments, and make the `SegmentEntry` as it subtrait.
* Rename to `NonAppendableSegmentEntry`
It differs semantically from `ImmutableSegmentEntry` by allowing point
deletion. Move point deletion to the trait too.
* Fix docstring
* use NonAppendableSegmentEntry where possible
---------
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
* download tar
* compute sha256 for stream download
* wip: propagate unpacking into down to the logic, todo: validation
* unpacked snapshot validation
* Minor tweaks
* Fix typo
* validation during unpack
* cancellation token
* update docstring
* remove redundant dep
* Rearrange unpack functions
- Rename `safe_unpack.rs` into `tar_unpack.rs` so it would be listed
near `tar_ext.rs` in IDEs.
- Replace calls like `ar = open_snapshot_archive(…); safe_unpack(ar, …);`
with a single call to `tar_unpack_file(…)`.
- Put calls to `Archive::new(); Archive::set_overwrite(false);` inside
`tar_unpack_reader` (was `safe_unpack`). So, now it is the only place
that does `set_overwrite`.
* Let clippy complain if tar::Archive::unpack used
* Mock snapshot download URL
Instead of downloading from storage.googleapis.com every time the test
runs, put small snapshot file to the repo.
The snapshot file is created using this command:
curl -s \
https://storage.googleapis.com/qdrant-benchmark-snapshots/test-shard.snapshot \
| tar \
--delete segments/4ea958d8-0b64-4312-9a53-0cd857e93535.tar \
--delete segments/65ac6276-8cca-4f5c-b767-9722190cee8b.tar \
> lib/storage/src/content_manager/snapshots/test-shard.snapshot
File contents:
$ tar tf lib/storage/src/content_manager/snapshots/test-shard.snapshot
wal/
wal/closed-255
newest_clocks.json
replica_state.json
shard_config.json
$ du -sh lib/storage/src/content_manager/snapshots/test-shard.snapshot
12K lib/storage/src/content_manager/snapshots/test-shard.snapshot
$ sha256sum < lib/storage/src/content_manager/snapshots/test-shard.snapshot
5d94eac5c1ede3994a28bc406120046c37370d5d45b489a0d2252531b4e3e1f2 -
---------
Co-authored-by: timvisee <tim@visee.me>
Co-authored-by: xzfc <xzfcpw@gmail.com>