* genericize ChunkedVectors.status, remove `Sized` bound
* check exists with `UniversalReadFileOps`
* Inline UioChunkedVectors bound, drop the alias (#8952)
The empty trait + blanket impl was a stable-Rust trait-alias workaround
that hid a fairly short bound (UniversalWrite<T> + UniversalWrite<Status>
+ Send + 'static) at the cost of an indirection readers had to mentally
unwind. Spelling it out at the three sites that need it is shorter overall
and immediately tells the reader what is required.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The test killed the last peer immediately after start_cluster, but
start_cluster only waits for cluster size and a known leader — not for
all peers to be promoted from learner to voter. If the last peer caught
up first, it became the only other voter alongside the leader; killing
it left a 2-of-2 quorum with one voter dead, and the subsequent
CreateCollection commit timed out after 10s.
Wait for all peers to be voters before killing one, so the survivors
form a 2-of-3 voter quorum.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The rapid drop/create loops in test_rejoin_cluster intentionally use
short 3s timeouts to accumulate Raft log entries quickly. Under CI
load the consensus apply for CreateCollection can exceed 3s (segment
setup competes with background flushes/optimizations), and the API
returns 500 even though the operation reaches consensus right after.
The matching upserts already pass `fail_on_error=False`; do the same
for `create_collection` to make the test resilient to that race.
* Initial commit
* Edits
* Change non-descriptive “here” link texts
* Review feedback
* Add Edge code snippet, Web UI section with screenshot, mention autit logging, mention NVIDIA and AMD GPU support
* Resize Web UI image; Move to end of Features section
* Restore Web UI section in README
Reintroduce Web UI section to README with visual aid.
---------
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
* Add `EncodedStorage::for_each_in_batch` method
* Implement `for_each_in_batch` method for `QuantizedChunkedMmapStorage`
* Add `EncodedVectors::for_each_in_batch` method
* Add `EncodedVectors::score` method
* Implement `score_stored_batch` for `QuantizedQueryScorer` and `Quanti…
* Remove `TElement` and `TMetric` type parameters from `QuantizedMultiQ…
* Use `QuantizedMultiQueryScorer` when building `raw_internal_scorer`...
`special_check_condition` in `FieldIndex` did not handle
`Match::TextAny`, so nested-condition evaluation fell through to the
`ValueChecker` fallback which uses naive `String::contains` instead of
proper full-text tokenization. This caused "good" to match "goodness"
inside nested filters.
Add `check_payload_match_any` to `FullTextIndex` and wire it into
`special_check_condition`. Also replace the catch-all `_ => None` with
explicit match arms for all `Match` variants.
Co-authored-by: Cursor Agent <agent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Match the REST API by adding an optional `WriteOrdering` field to
`CreateVectorNameRequest` and `DeleteVectorNameRequest`, and propagate
it through the tonic handlers and remote-shard forwarding paths.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
These tests rely on the collection size stats cache being refreshed to
detect that a size limit has been exceeded. Without wait=true, upsert
operations are written to WAL and acknowledged immediately without being
applied to segments. When the cache refreshes, it reads segment data
which may not yet reflect the pending WAL operations, causing the size
check to see stale values and not reject the request.
Adding wait=true ensures operations are applied to segments before the
response returns, so the cache refresh sees the correct sizes.
Co-authored-by: Cursor Agent <agent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* propagate to BufferedDynamicFlags
* use duplicate for tests
* propagate to Bitvec/Roaring flags
* propagate to RoaringFlags
* fixup! propagate to BufferedDynamicFlags
* propagate to BitvecFlags
Extract the read-side API into ChunkedVectorsRead<T, S: UniversalRead<T>>
and rebuild the existing ChunkedVectors<T, S: UniversalWrite<T>> on top of
it via composition + Deref. Lets read-only consumers use the storage
without pulling in UniversalWrite, and avoids method duplication.
Reorganizes the file into a chunked_vectors/ module: config (constants,
ChunkedVectorsConfig, Status), chunks (read_chunks/create_chunk helpers),
read (ChunkedVectorsRead), write (ChunkedVectors wrapper). read_chunks
now takes a writeable flag so both open paths share it.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The deadline loop only waited for the download phase to become
"streaming", but the assertion also required bytes > 0. On CI the
assertion could fire before iter_content yielded the first chunk,
causing a flaky failure. Wait for bytes > 0 in the deadline loop
and improve the error message to show bytes/error state.
Co-authored-by: Cursor Agent <agent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* test: keep optimizers disabled during snapshot transfer in deferred test
The snapshot variant of test_shard_transfer_includes_deferred_points was
flaky because optimizers were enabled before the transfer, letting the
optimizer race ahead and fully index the segment before the snapshot was
captured (~1s of HNSW build for 500 small vectors fits comfortably before
the snapshot is taken). The deferred-state assertion then fails since all
points are already visible.
Only enable optimizers before the transfer for stream_records (which needs
them for its internal wait=true). For snapshot, leave optimizers disabled
through the transfer so deferred state is preserved on the wire, then
enable them afterwards for trigger_upsert_wait_true. The hung server-side
wait=true from the timeout-and-retry block does not block the snapshot —
wait_for_deferred_points_ready runs in a detached tokio::spawn.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test: skip wait=true probe for snapshot variant
CI showed that with optimizers kept disabled through the snapshot transfer
(needed to preserve deferred state on the wire), the wait=true probe at
the start of the test leaves a hung server-side request: update_local
holds local.read() until the deferred wait resolves, and there is no
optimizer to resolve it. The subsequent shard transfer's apply path
deadlocks against that held read lock when queue_proxify_local tries to
take local.write().
For stream_records the config update later cancels the hung worker, so
the probe is fine there. Move the probe (and config update) under the
stream_records branch so the snapshot variant doesn't leave a hung
update around. The probe was auxiliary behaviour verification, not
central to the snapshot-of-deferred-points assertion.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Revert "test: skip wait=true probe for snapshot variant"
This reverts commit d07aac78a263d7a91691e43444b9dae44e3d179f.
* test: add reproducer for deferred-wait shard-transfer deadlock
Adds test_shard_transfer_with_hung_deferred_wait_does_not_deadlock as a
focused reproducer for the engine bug surfaced by the snapshot variant
of test_shard_transfer_includes_deferred_points.
Lock-ordering chain:
1. With prevent_unoptimized=true and max_optimization_threads=0, a
wait=true upsert on deferred points enters
wait_for_deferred_points_ready (update_worker.rs:241), which loops
on tokio::select over cancel and optimization_finished. The
optimization_worker hits limit==0 and `continue`s without firing
optimization_finished_sender (optimization_worker.rs:172-174), so
neither branch of the select ever fires.
2. update_local (replica_set/update.rs:49) holds self.local.read()
across the entire update await. actix-web does not cancel the
response future on client disconnect, so the read guard stays alive
even after the client's 5s timeout.
3. A subsequent snapshot transfer eventually calls queue_proxify_local
(replica_set/shard_transfer.rs:122), which needs self.local.write().
tokio::sync::RwLock is write-preferring: the queued writer blocks
new readers, including is_local() calls on the consensus apply
path itself (shard_transfer.rs:129-130). The apply never returns,
the consensus broadcast never fires, POST /cluster times out with
"Waiting for consensus operation commit failed".
The new test asserts the symptom (POST /cluster must return promptly)
without papering over the bug, so it stays red until the engine is
fixed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(replica_set): release local read guard around deferred-points wait (#8862)
* fix(replica_set): drop remotes read guard early in update_impl
`update_impl` was holding `self.remotes.read()` and `self.local.read()`
across the entire update await, including the deferred-points wait that
can park indefinitely under prevent_unoptimized + max_optimization_threads=0.
When a shard transfer is started concurrently with a parked wait=true
update, the consensus apply runs `add_remote`, which calls
`self.remotes.write().await`. tokio::sync::RwLock is write-preferring:
the queued writer is blocked behind the held read, the apply never
returns, and `POST /cluster` times out with "Waiting for consensus
operation commit failed".
Fix: snapshot updatable remote shards into owned `Vec<RemoteShard>` and
drop the read guard before the await. The remote_update futures now own
the cloned RemoteShards, so they no longer borrow from the guard.
The `local` guard is still held across the await (futures borrow
`&Shard` from it). Releasing it would unblock `queue_proxify_local`'s
`local.write()` too, but that requires wrapping `Shard` in `Arc` —
deferred to a follow-up. For the consensus-commit-timeout deadlock
exposed by `test_shard_transfer_with_hung_deferred_wait_does_not_deadlock`,
dropping `remotes` is sufficient.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(updater): wake deferred wait on caller-receiver drop
`wait_for_deferred_points_ready` parked on a `tokio::select` over
`cancel.cancelled()` and `optimization_finished_receiver.changed()`.
Under prevent_unoptimized + max_optimization_threads=0, neither fires:
optimization_worker.rs:171-174 hits `limit == 0` and `continue`s
without notifying, and the cancel token is the worker's lifecycle
token (only fired by stop_update_worker on config update / shutdown).
The top-of-loop `is_closed()` poll didn't help — the loop never
re-runs once the select parks.
Take `feedback_sender` by `&mut` and add `feedback_sender.closed()`
as a third select branch. When the matching `Receiver` is dropped
(by upstream cancellation, client-supplied timeout, or any future
cancellation), the detached task wakes immediately and exits with
WaitTimeout instead of staying parked until the next worker restart.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* [AI] split update operarion into submit and independent wait function
* [AI] refactor `update_local` to drop local shard lock after submitting update operation
* [AI] refactor `update_impl` for early release of the lock in case of local shard update
* fmt
* Apply suggestion from @generall
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Cleanup lifetimes and generic type parameters
- Rename read pipeline lifetime from `'a` into `'file`
- Use explicit `where` clauses everywhere
* Cleanup
* fixup! Cleanup lifetimes and generic type parameters
* fix(segment): immutable map index skips values with no live points on load
When MmapMapIndex::open ORs the id-tracker's runtime deletion bitvec
into the on-disk one at open time, ImmutableMapIndex::open_mmap could
insert a zero-count entry into value_to_points for any value whose live
points were all deleted, then immediately trip its own post-build sort
assert. Skip such values, mirroring the runtime invariant maintained by
remove_idx_from_value_list which already removes entries when their
count drops to zero.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(segment): file-system immutability test for payload indices
Builds an immutable segment with all 8 PayloadSchemaType variants
indexed, snapshots every byte under payload_index/, then asserts
byte-for-byte equality plus per-field query correctness after
delete_point, flush, drop+reload, and a second deletes+flush on the
reloaded segment. Each query exercises a different read path of an
immutable index variant: map exact-match (keyword/uuid/integer),
numeric range (float), datetime range, geo bounding box, full-text
token match, bool match. Reproduces the regression fixed in the
previous commit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Mirrors the layout of mutable_id_tracker: storage helpers for the
mappings, versions, and deleted bitslice files live in their own
submodules, leaving mod.rs focused on the ImmutableIdTracker type and
its trait impls.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Introduce a read-only payload storage that wraps `GridstoreReader<Payload>`
and implements `PayloadStorageRead`. Also rename the parameter on
`PayloadStorageRead::get`/`get_sequential` from `point_id` to `point_offset`
for consistency with the underlying storage API.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(index): extend VectorIndexRead and PayloadIndexRead for telemetry/info
Two trait extensions (no defaults — every implementor must opt in):
* \`VectorIndexRead::is_index\` — distinguishes a real index from a plain
full-scan one. Used by reporting code. Moved out of inherent
\`VectorIndexEnum::is_index\` into the trait. Explicit impls on Plain
(false), HNSW (true), Sparse* (true).
* \`PayloadIndexRead::get_telemetry_data\` — per-field-index telemetry.
Moved out of inherent \`StructPayloadIndex::get_telemetry_data\` into
the trait impl. \`PlainPayloadIndex\` returns an empty Vec.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(segment): migrate info/size_info/telemetry to SegmentReadView
Step 11 of the SegmentReadView migration — the final logical step.
New \`read_view/info.rs\` exposes builder methods:
* \`build_size_info(uuid, segment_type, is_appendable)\`
* \`build_info(uuid, segment_type, is_appendable)\` — same plus
\`index_schema\`
* \`build_telemetry(uuid, segment_type, is_appendable, config, detail)\`
The trivial segment-level fields (\`uuid\`, \`segment_type\`,
\`is_appendable\`, \`config\`) are passed in by the caller — they stay
direct on each segment-type rather than going through the view.
Everything else (vector data breakdown, payloads size, deferred
counts, vector-index telemetry, payload-field telemetry, …) is
computed once inside the view through the read traits.
\`Segment::size_info\`, \`info\`, and \`get_telemetry_data\` collapse to
one-line \`with_view\` delegators that pass in the trivial fields.
Cleanup: \`Segment::deferred_deleted_count\` is now unused (the view
has its own equivalent helper); deleted.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Leftovers from moving \`StructPayloadIndex::formula_scorer\` out of this
file. Caught by clippy.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
\`FormulaScorer<'a>\` is already self-contained — it holds the parsed
formula, prefetch scores, retrievers and condition checkers, all owned
or borrowed independently of any payload index. Wrapping it in a trait
adds nothing (a future ReadOnlySegment can construct one too).
* Remove the \`FormulaScorerRead\` trait. \`score(point_id)\` goes back
to being an inherent method on \`FormulaScorer\`.
* \`PayloadIndexRead::formula_scorer\` returns
\`OperationResult<FormulaScorer<'q>>\` directly.
* \`StructPayloadIndex\` and \`PlainPayloadIndex\` updated to match.
* \`PlainPayloadIndex\` no longer needs the turbofish placeholder —
\`Err(...)\` is enough.
* View's \`formula_rescore.rs\` drops the \`FormulaScorerRead\` import.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Step 10 of the SegmentReadView migration.
Move \`segment/formula_rescore.rs\` → \`read_view/formula_rescore.rs\`:
* \`do_rescore_with_formula\` (private helper).
* \`rescore_with_formula\` (\`ReadSegmentEntry\` orchestrator).
Both now use the trait-method \`PayloadIndexRead::formula_scorer\`
(added in the prior commit) and \`IdTrackerRead::internal_id\` instead
of inherent calls.
\`Segment::rescore_with_formula\` collapses to a single
\`with_view(|v| v.rescore_with_formula(...))\` delegator. The legacy
\`segment/formula_rescore.rs\` is deleted; \`mod formula_rescore;\`
removed from \`segment/mod.rs\`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
A read-only payload index implementation will have its own concrete
formula scorer type, so PayloadIndexRead can't return the appendable
\`FormulaScorer<'a>\` directly.
* New \`FormulaScorerRead\` trait next to \`FormulaScorer\` exposing only
what the rescore code path consumes (\`score(point_id)\`). Implemented
for \`FormulaScorer<'_>\` by moving its inherent \`score\` into the
trait impl.
* \`PayloadIndexRead::formula_scorer\` returns
\`OperationResult<impl FormulaScorerRead + 'q>\` (RPITIT).
* The inherent \`StructPayloadIndex::formula_scorer\` (which lived in
\`formula_scorer.rs\`) is moved into the trait impl block in
\`struct_payload_index.rs\`, with the body delegating to a new
\`FormulaScorer::new\` constructor (fields stay private).
* \`PlainPayloadIndex\` always returns
\`Err::<FormulaScorer<'q>, _>(...)\` — formula scoring is not
supported there. The turbofish supplies the placeholder type tag.
* Re-export \`FormulaScorer\` and \`FormulaScorerRead\` from
\`rescore_formula::mod\`. \`retrievers_map\` bumped from \`pub(super)\`
to \`pub(crate)\` so the trait impl can call it.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>