* Add a BM25-over-sparse baseline benchmark
The text payload index is meant to score about as fast as BM25 over sparse
vectors, so the number it has to match needs to exist before the scorer
does. Measures a local shard end to end, no HTTP.
- embeds the corpus through `lib/bm25` with its defaults, so the baseline
is the route a user migrates from rather than a reimplementation
- Zipf-like vocabulary. On a uniform one every term is equally selective,
IDF is flat and pruning has nothing to prune, which would flatter any
scorer measured against it
- two shards rather than one shard before and after optimization: a shard
that will optimize starts as soon as the upsert lands, so the first cut
timed a half-converted index and called it fresh. Both states assert
what they hold before anything is timed
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Measure the sparse BM25 baseline where the routes separate
- 200k documents by default (BM25_SPARSE_DOCS overrides): at 20k every shape
of both routes measures the same and half of a shard-level query is the
shard; reachable through the shard since #10682
- a third shard with the sparse index on disk, which the optimized state never
exercised
- recall at 10 against BM25 by definition, printed per state: the default
avg_len of 256 on a corpus averaging 110 tokens misses a quarter of the true
top 10, so the optimized state is also timed with the corpus average
- corpus, queries, reference and recall move to segment::fixtures::bm25_corpus,
to be shared with the text-index bench and the comparison harness
- module doc: the sparse shapes measure within a few percent of each other; the
states exist for the text comparison
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Refuse an empty corpus in the sparse BM25 baseline
BM25_SPARSE_DOCS=0 built empty shards, made the average length NaN and
scored every empty truth as recall 1.0.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Measure what the sparse BM25 baseline states claim
The fresh shard kept the default 10 MB indexing threshold and optimized
itself in the background, so it was timed as a second optimized state.
Disable indexing and re-check it after timing. Keep the shard storage
under CARGO_TARGET_TMPDIR so the on-disk index is not read from a tmpfs.
Fix doc comments that named missing files, a no-op IDF clamp and the
wrong reason for the empty-corpus guard.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Configurable key prefix for object storage snapshots
Snapshot object keys were always derived from the local snapshots path,
so every deployment sharing a bucket wrote under the same `snapshots/`
root. Each cloud config block gains an optional `prefix`, and objects
become `<prefix>/snapshots/...`.
The prefix is applied by wrapping the client in `PrefixStore`, so the
snapshot operations and the names returned by the API are unchanged.
Leading, trailing and repeated slashes are dropped, and an empty prefix
is a no-op.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Test that snapshot operations cannot escape the storage prefix
Hostile targets are handed to the cloud manager directly, past
`validate_snapshot_name`: parent references, absolute paths, encoded
slashes, backslashes and empty paths. Writes must stay under the prefix,
and objects planted outside the prefix must be invisible to list,
download, stream and delete.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Normalize empty prefix components
Refactor prefix handling to remove empty components.
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tim Visée <tim+github@visee.me>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* Add `match: { substring }` filter condition
Unindexed `text` and `text_any` matching became token-aware in #10341 and
#10593. Users who relied on the old raw substring behaviour get it back as
an explicit condition: `match: { "substring": "..." }` selects points with a
string value containing the given string, byte-wise and case-sensitive,
consistent with exact keyword and prefix matching.
Execution: a keyword index (with or without the `prefix` option) serves the
condition by scanning its value dictionary and uniting the postings of the
matching keys; cardinality reuses the prefix estimator, generalised into
`keys_union_cardinality`. The per-point checker goes through the forward
index. Without a keyword index the condition falls back to reading the
payload. Text, bool, integer and uuid indexes decline it.
Strict mode: the condition requires the `KeywordMatch` capability, so with
`unindexed_filtering_retrieve: false` it is rejected on unindexed and on
text-indexed fields and allowed on any keyword index.
API: `MatchSubstring` in the REST `Match` union with regenerated OpenAPI,
gRPC `Match.substring = 12`, edge python `MatchSubstring`, edge ffi
`Match::Substring`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Test substring fallback on a text-indexed field, document estimator params
A text index cannot serve `substring`, so on a field that has only a text
index the condition runs through the payload fallback; only strict mode may
reject it. Pin that in the OpenAPI suite and reword the strict-mode unit
test comment, which read as if the text index itself blocked the query.
Also spell out what `keys` and `postings` mean in `keys_union_cardinality`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Serve `match: { substring }` from the keyword key dictionary
The condition used to enumerate keys through `MapIndexRead::for_each_value`,
which on the on-disk variant drags the whole `value_to_points` file through
`for_each_entry`, plus one random read per matching key for its postings
count. Query planning paid that scan in full, before deciding whether to use
the clause at all.
Route it through the `prefix_index.bin` key dictionary instead: front-coded
keys with their postings counts inline, no postings. Estimation now reads
keys only and never touches `value_to_points`; filtering takes the matched
key list and resolves postings in one batched read, as prefix matching
already does.
This makes the `prefix` option a requirement: a keyword index without it has
no key dictionary, so it declines the condition and falls back to the payload
scan, the same as a text index. Strict mode follows — substring now infers
`KeywordPrefix`, so `unindexed_filtering_retrieve: false` names
`keyword (with prefix: true)` as the index to create.
`PrefixIndex::for_each_key` reads blocks in ~1 MiB chunks rather than the
whole key section at once: a substring cannot be pruned by the block index,
so the one-shot read would grow with the dictionary.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Sync MatchSubstring OpenAPI description with Rust docs
After routing substring matching through the prefix key dictionary,
the schema docstring required regenerating so OpenAPI stays consistent.
* Scan the map index keys for substring match without a dictionary
Without the keyword dictionary a substring condition was declined by the
field index and left to the per-point condition checker, which reads the
forward index for every candidate point. Enumerate the distinct keys of
`values_to_points` instead: the same one-pass-over-distinct-values shape as
the dictionary scan, only over a structure that interleaves keys with their
postings. Filtering and cardinality estimation are then always served, so
the condition can act as a primary clause on a plain keyword index.
Prefix matching keeps its per-point fallback: an ordered dictionary is what
makes a prefix a bounded range, and enumerating every key to answer one is
not a trade worth making implicitly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Reject substring matching in strict mode
A substring condition is answered by looking at every distinct value of the
field: no index gives it a bounded access path, so there is no index a user
could create to make it affordable. Reject it under strict mode instead,
wherever a filter reaches verification — read and write filters, nested
sub-filters, and prefetch filters.
Filter limits are now checked before the unindexed-field check, so the
rejection is not reported as "create an index for this key", advice that
would lead to the same rejection afterwards.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Update the strict mode substring test to the new rejection
The test asserted that substring filtering under strict mode asks for a
keyword index with the `prefix` option. It is now rejected whatever index
the field carries, so every case in the test gets the same answer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Estimate a substring condition without scanning
Counting the keys a substring matches costs the same scan as answering the
condition, and `filter` then repeats it to collect those keys. Report the
uninformed estimate instead — the one an unindexed condition has always
reported — and keep the primary clause, so the scan happens once, in
`filter`, and only when the planner picks the condition to drive iteration.
With no counts to collect, `substring_scan` collapses into `substring_keys`:
the in-RAM variants no longer look up a posting count per matched key.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Don't to parse everything as UTF-8
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com>
Co-authored-by: timvisee <tim@visee.me>
Snapshot storage was limited to S3 although the object_store dependency
already ships the GCS and Azure backends. `snapshots_storage` now accepts
`gcs` (alias `gcp`) and `azure`, configured through `gcs_config` and
`azure_config` blocks next to the existing `s3_config`. The legacy S3
shape is unchanged.
Client construction is split into one builder per backend, all sharing
the Qdrant user agent and the plain-HTTP rule for `http://` endpoints.
Startup warns when a config block for an unselected backend is present.
The e2e snapshot recovery test is parameterized over the cloud backends.
The GCS case is skipped because fake-gcs-server does not implement the
XML multipart upload API that object_store uses for GCS uploads.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* test: pin persisted proxy segments follow-up findings
* fix: tests
* Respect `up_to` in flushing pending proxy changes
* Fix unproxy divergence, centralize logic in single shared function
* Don't pass locked segments we don't use
* Close the post-swap window losing acknowledged proxy changes
`finish_optimization` left a window between `swap_new` and the end of the
function where the propagated proxy changes were durable in no reachable
place. The proxies had left the holder, so a flush pass no longer saw them
as unsaved work and acknowledged the WAL past what their pending changes
logs persisted, while their source files were still on disk contradicting
the newer state. Their `ack_pin` was only registered at the very end, and
the optimized segment, the only durable home of those changes, had no
version file yet, so `normalize_segment_dir` deleted it on the next load.
A failure or crash in that window lost every change past the proxies' logs.
Close it from both sides, reordering only:
- Save the optimized segment's version file right after the flush that made
the propagated changes durable, before the swap. That flush is what
`SegmentBuilder::build` postponed the version file for, so the segment is
loadable from the swap onwards.
- Register the proxies' deferred destruction (and with it their `ack_pin`)
under the same write lock that evicted them, so no flush pass can observe
the holder without either the proxies or their pin.
Collecting the deferred point ids moves up with the registration, as it
borrows the swapped-out proxies. The deferred destruction still cannot run
before the manifest is synced: `locked_proxies` holds the segments alive
until the end of the function, so `try_drop_data` retries until then.
* Rename the optimization test hook after the window it guards
The hook no longer sits before the version save, it marks a failure
anywhere in the window after the optimized segment was swapped in. Rename
it and the test accordingly, and restate the test doc as the invariant that
window must uphold rather than the bug it used to describe.
* Pin WAL ack while creating snapshot
* Patch test that was stuck
* Initialize necessary feature flags in tests
* Correctly propagate changes in two stages, lock updates on second stage
* Use existing WAL ack pinning infrastructure
* test: pin unproxy phase 2 propagation failure losing acknowledged changes
* test: pin WAL ack pin at zero suppressing clock persistence
* Fix propagate and unproxy data consistency error on failure
* Store clocks before checking WAL ack pin
* fix: wait for the flush worker when stopping it in tests
* fix: linter
---------
Co-authored-by: timvisee <tim@visee.me>
* Rename `wal_keep_from` to `wal_ack_pin`
The parameter pins the WAL acknowledge, the new name says so.
* Support pinning the WAL acknowledge from multiple places at once
The WAL acknowledge had a single pin slot, a shared `AtomicU64` any second
user would have clobbered. Replace it with `WalAckPins`, holding any number
of live pins. The flush worker never acknowledges at or past the lowest of
them.
Taking a pin hands out a `WalAckPinGuard` that releases when dropped, so the
queue proxy shard no longer has to release it by hand. The registry holds its
pins weakly, which means dropping the guard is all it takes and a queue proxy
lost to an unwind can no longer stall the WAL acknowledge forever.
* Tests for the WAL acknowledge pins
* Debug assert that set does not move version backwards
* Update tests
* fix: expose ReshardingStage in internal telemetry
Useful for cluster-manager operations.
Non-breaking change for existing API.
The current field values look stable enough,
so it's not worrisome to keep them stable.
Signed-off-by: Anton Antonov <anton.synd.antonov@gmail.com>
* fix: address comment, also document uuid for symmetry
Already returned, just document it in OpenAPI schema.
Signed-off-by: Anton Antonov <anton.synd.antonov@gmail.com>
---------
Signed-off-by: Anton Antonov <anton.synd.antonov@gmail.com>
* segment: move proxy pending change types into segment crate
Move the types describing the changes a proxy segment buffers — point
deletes (`ProxyDeletedPoint`), payload index changes (`ProxyIndexChange`,
`ProxyIndexChanges`) and vector name changes (`IntendedVector`,
`ProxyVectorNameChanges`) — from `shard::proxy_segment` into a new
`segment::pending_changes` module.
Pure move, no behavior change: the proxy segment re-exports them from
their old location. Having them in the segment crate lets both the proxy
segment and the segment load path share them, in preparation for
persisting pending proxy changes to disk and replaying them on restart.
* segment: add PendingChange describing a persisted proxy operation
Add the `PendingChange` enum with one variant per operation type a proxy
segment buffers — point delete, payload index change, vector name
change — each carrying the operation version it was issued with. This
is the shape in which pending proxy changes are persisted to disk.
Derive serde on it and on the buffered change types it embeds, so
entries can be serialized into a log file and read back. `PartialEq` on
those types lets a persisted batch be matched against the in-memory
pending buffer after a flush.
* segment: add PendingChanges component persisting proxy changes to a log
Add `PendingChanges`, the component that manages the operations a proxy
segment buffers for one proxy layer, and persists them to disk so they
no longer only live in memory.
It keeps the same per-type buffers the proxy segment served its reads
from (point deletes, payload index changes, vector name changes), plus a
single registration-ordered buffer of everything not yet persisted.
`flusher()` writes that buffer into an append-only log file inside the
wrapped segment's directory: `pending_changes.log` for the inner most
proxy layer, with the layer number as a suffix for each layer above it.
Appends follow the mutable ID tracker: all new entries are serialized
into one buffer and written with a single call on an append-mode file,
then fsynced, so a crash can only leave a torn entry at the very end.
Loading truncates such an entry — its operations were never durable and
thus never acknowledged in the WAL — but fails hard on a malformed entry
in the middle, which cannot be explained by a torn append.
The component tracks the highest operation version the log covers.
Every registered operation at or below it is either durable in the log
or was a no-op that does not need recovery; a flusher advances it to the
proxy's version even when there is nothing to write. The pending buffer
is deliberately not cleared when the proxy propagates its changes to
the wrapped segment, as that only makes them durable once the wrapped
segment flushes. Replaying an entry twice is a version-gated no-op.
A log file left behind by a previous proxy on the same segment is
adopted by `open()`: new entries are appended after it and its highest
version is taken over, while its entries are not loaded into the
buffers as they are already applied to the segment. `load()` also
reconstructs the buffers, for callers that do want the buffered state.
* segment: replay persisted pending proxy changes onto a segment on load
Add `recover_pending_changes`, to be called when a segment is loaded on
restart, before regular WAL replay. If the segment directory holds
pending changes log files, the proxies that wrote them did not
propagate their buffered state into the segment before the process
stopped. Instead of reconstructing the proxies, replay all logged
operations directly onto the segment: inner most proxy layer first,
each file in append order, through the regular version-gated segment
operations (`apply_change`). Entries the segment already applied are
silently skipped, so a stale file is harmless.
The segment is force-flushed before the files are removed; a crash in
between merely replays the files once more.
* segment: test PendingChanges component
Cover the pending changes component: registering and flushing each
operation type and reconstructing the buffers from the log, log file
naming per proxy layer and gap-tolerant listing, covering the proxy
version without entries, operations registered while a flusher is
captured, flushers of a dropped component, torn-tail truncation versus
mid-file corruption, adoption of an existing log, and replaying logs
onto a real segment: fresh, stale (already applied), multi-layer, and
vector name changes.
* segment: include pending changes logs in segment snapshots
Register the pending changes log files of a segment in its snapshot:
add them to `snapshot_files` next to the segment state and version
files, existence-guarded, and to the segment manifest as unversioned
files. Full, partial and streamed snapshots therefore all carry them.
The recovery side needs no changes: a restored segment is loaded like
any other, which replays and removes the logs.
* shard: back proxy segment pending changes by PendingChanges component
Replace the proxy segment's separate `deleted_points`, `changed_indexes`
and `changed_vector_names` fields with a single `PendingChanges`
component. Reads keep going through the same per-type buffers, now
behind accessors; writes go through the component's `register_*`
methods, which additionally queue every operation for persistence.
Opening the component is fallible, as it adopts a pending changes log
a previous proxy may have left in the wrapped segment's directory, so
`UnsyncedProxySegment::new` now returns a result. Wrapping another proxy
opens the next proxy layer up, writing to its own dedicated log file.
No behavior change yet: the proxy still flushes and reports persistence
exactly as before, nothing is written to the log.
* shard: persist proxy pending changes on flush, stop holding back WAL ack
Hook the pending changes component into the proxy segment's flush: the
proxy flusher first persists the buffered operations into the pending
changes log, then passes the flush along to the wrapped segment. The
proxy's `persistent_version` now covers what the log durably holds on
top of what the wrapped segment persisted itself.
That is what lifts the WAL cap proxies imposed so far. `flush_all`
compares each segment's version against its persistent version; a
proxy used to report only the wrapped segment's persisted version while
its own version climbed with every buffered operation, so the WAL could
never be acknowledged past the point the proxy was created at, and a
restart replayed all of it — potentially very expensive operations,
such as an update by filter, all over again. With the buffered state
durable on disk the generic rule acknowledges the full version, and a
restart recovers it from the log instead.
Dropping a proxy's data drops the component first, which waits for any
in-flight pending changes flusher so it cannot append to the segment
directory while that is being deleted.
Update the proxy flush test to the new semantics, add a segment holder
test asserting the acknowledged version advances past a proxied delete,
and update the ack pin rationale in `finish_optimization`: the pin is
still needed after the proxies leave the holder, it just snapshots a
persistent version that now includes the log.
* shard: propagate proxy changes when unwrapping on optimizer cancel
When an optimization is cancelled or fails, `unwrap_proxy` puts the
wrapped segments back into the segment holder. Propagate the changes
buffered in each proxy into its wrapped segment first, as the snapshot
unproxy path already does, instead of dropping them with the proxy.
The pending changes log is deliberately left in place when unwrapping:
deleting it before the wrapped segment has flushed the propagated
changes would not be crash safe. It is cleaned up on restart and when
the segment directory is dropped, and a new proxy on the same segment
adopts and appends to it; replaying a stale file is safe because all
operations are version gated.
* shard: test persisted proxy pending changes
Test the proxy segment against its persisted pending changes: buffered
changes survive dropping the proxy without propagation and are replayed
onto the segment when it is loaded again; unwrapping leaves the log in
place and a new proxy on the same segment adopts and appends to it;
layered proxies each persist into their own log file and a restart
replays both; and a persisted log is part of the segment manifest and
snapshot.
* collection, edge: recover persisted proxy changes on segment load
Replay the pending changes logs left behind by proxy segments onto each
segment when a shard loads its segments, right after consistency
repair and before the payload index rebuild, vector name reconciliation
and WAL replay. Proxy state that made it to disk no longer holds back
the WAL acknowledge, so this is where it must be recovered from.
Proxies are not reconstructed: the segment holder starts with plain
segments carrying the replayed operations, and the logs are removed
once the segment flushed them.
* collection: test crash recovery through persisted proxy changes
End-to-end test of the persisted pending changes: wrap every segment of
a local shard in a proxy, delete points so the deletes are only
buffered, flush, and assert the acknowledgeable version covers them.
Then acknowledge the WAL up to that version, drop the shard without
ever propagating the proxies, and load it again: the deletes are gone
from the WAL and must come back through the pending changes logs.
The delete under test is deliberately not the last WAL entry, as the
acknowledge never passes the last entry and that one is always
replayed.
* segment: make replaying persisted proxy changes on load an explicit mode
Add `PersistedProxyChanges` to state whether persisted pending proxy
changes are replayed onto a segment when it is loaded. `Replay`, the
default, recovers them and removes the logs as before. `Ignore` leaves
both the segment and the log files untouched and logs at debug level
that replaying was skipped; it is for segment files that mirror those
of another writer, where replaying would make the local copy diverge
from what the writer's manifest describes.
All callers pass `Replay` for now, no behavior change.
* collection: do not replay persisted proxy changes on partial snapshot recovery
Partial snapshots are recovered by read replicas in a read/write
segregation setup. A read replica must not mutate its segments, so it
cannot replay the persisted proxy segment changes on load and must
ignore them instead: its segment files are a local copy of the writer's
that must stay a faithful mirror of them, as later partial snapshots are
diffed against what the writer's manifest describes. Replaying would
mutate the segment files and remove the logs, making the copy diverge.
Thread the replay mode through `LocalShard::load` as a dedicated
`PersistedProxyChanges` argument, derived from the recovery type:
`RecoveryType::Full` replays as before, `RecoveryType::Partial` ignores
the persisted changes and leaves the logs in place. Regular shard loads
replay.
Extend the crash recovery test with an ignoring load first: the delete
under test must not come back and the logs must survive, before a
replaying load recovers it.
* Persist wrapped segment before pending changes
Prevents raising version of proxy segment too early
* Fix comment
* Fix crash window, only ready optimized segment after propagating changes
The optimizer renamed a newly built segment into segments_path and wrote
its version file before finish_optimization propagated the proxies'
buffered changes into it. A crash in that window left the segment
restart-loadable but stale, permanently losing or resurrecting points.
Defer the version file save until finish_optimization has fully
reconciled proxy changes into the segment, including the post-swap dedup
pass, so it stays invisible to restart and snapshot recovery until then.
SegmentBuilder::build() gains a `ready` flag; load_segment gains
`ignore_missing_version` for the one caller reloading before that point.
Incidentally also closes the crash-unsafe cancellation-orphan cleanup
gap noted in #9217, since a cancelled build is discarded on restart the
same way.
* Force flush optimized segment, otherwise we may lose proxy changes
* Don't force flush after replay, defer deleting log files until flush
* Include persisted proxy changes log file in segment manifest
* Add random ID to proxy log files, prevent instance conflicts
* Rename proxy log file, always include level
* Delete proxy log file on unproxy, defer until next flush cycle
* Fix truncation
* Reformat
* Lock persisted segments behind runtime feature flag
* Enable necessary feature flags in tests
* Fix linters
A hung harness_no_restarts run blocked ubuntu CI for ~50m with only SLOW
markers and no failure dump. Kill tests after five slow-timeout periods,
and print stage/op breadcrumbs to stdout so the timeout failure includes
seed, storage path, and the last op/stage that never returned.
* Send the queue proxy batch as a pre-encoded gRPC body
The parent commits moved the WAL read, the operation clone and the request build
off the async runtime. Two passes over the batch were left on it.
Measured cost of each synchronous pass over a 26.2 MiB send batch (800 ops x 30
points x 256 dims), release build:
pass 1 clone WAL operations 72.4 ms moved by the parent commits
pass 2 build gRPC request 22.2 ms moved by the parent commits
pass 3 clone request to send 65.7 ms on the async runtime
pass 4 protobuf encode (tonic) 41.4 ms on the async runtime
Pass 3 is there because `with_points_client` takes `impl Fn` and the channel pool
calls that closure once per attempt, so each attempt needs its own owned message.
Pass 4 runs inside `poll_next`: for a unary call tonic encodes the whole message
in a single synchronous `encode_item`, so a worker is blocked for the full 41 ms,
once per attempt.
Encoding the batch up front removes both. The generated client cannot take a
pre-encoded body, `update_batch` is typed `impl IntoRequest<UpdateBatchInternal>`,
but all it does is pick a codec and a path and call `Grpc::unary`, and we already
build the client ourselves from a pooled channel. `update_batch_pre_encoded` does
the same three things with a codec that writes the encoded bytes through and
decodes the response with prost.
What stays on the runtime is the copy of the encoded body into tonic's send
buffer: 14.9 ms for 26.2 MiB under jemalloc, nearly all of it faulting in freshly
mapped pages rather than the copy itself (0.9 ms when the allocator hands back
warm pages). Retries share the same refcounted bytes instead of cloning and
re-encoding, so they drop with it.
on the runtime before 107.2 ms per attempt
on the runtime after 14.9 ms per attempt
Bypassing the generated client means the RPC path and message types no longer
follow the proto automatically, so a test checks them against the compiled
descriptor set.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WZjWYpKeYdGpKiLeEZU9oc
* Hand the pre-encoded update batch a configured Grpc, take the service name from the generated code
`update_batch_pre_encoded` took a bare channel plus a `max_decoding_message_size`
argument, and its only caller passed `usize::MAX`. Every other internal client
applies that limit inside its `with_*_client` helper, so do the same: `with_grpc`
hands out the `tonic::client::Grpc` the generated clients wrap, already
configured, and the argument goes away.
The service half of the RPC identity now comes from the generated
`points_internal_server::SERVICE_NAME` instead of a second literal. Only the
method name and the path literal remain hand-written, still pinned to the
descriptor set by the test.
`PreEncodedMessage::encode` uses `encode_to_vec`: one pass instead of a separate
`encoded_len` call, and no `expect`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Adapt pre-encoding to the build/forward split from #10599
`forward_update_batch` takes `Arc<UpdateBatchInternal>` and encodes it
once on the blocking pool, so the channel pool's attempts share the
bytes. The queue proxy keeps the built request in that `Arc` across
`BATCH_RETRIES` and for the per-operation isolation path.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: generall <andrey@qdrant.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* Reject snapshot upload without collection config before loading it
The raw IO error from `CollectionConfigInternal::load` embedded the
server-side temporary path in the API response. Check for the file first
and return a fixed bad-input error instead.
Part of #10553
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A3o9eWSZNMa6WAs5F2HMZC
* Fix missing-config check to run before restore_snapshot loads config
The path-leak guard lived after Collection::restore_snapshot, but that
function already calls CollectionConfigInternal::load and surfaced the
temp path as a 500. Require a regular config.json file before loading,
and cover a directory-shaped config entry in the openapi test.
* Use a valid empty TAR in the missing-config snapshot upload test
Avoid depending on malformed-archive handling; exercise the missing
collection-config path with a real TAR that has no entries.
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com>
* Move queue proxy WAL read off the async runtime
`read_wal_batch()` read up to MAX_BATCH_BYTES (32 MiB) from the WAL and
deserialized it synchronously on the async runtime. That runtime also serves
all internal gRPC, including the health check the transport channel pool uses
to decide whether a peer is alive, so during the queue-replay phase of a
snapshot shard transfer the sender periodically stopped answering internal
requests for the duration of a 32 MiB disk read plus decode.
Measured on a 3-node cluster with no CPU/memory/IO limits and ~30% free RAM:
the sender's health-check p99 went from 3-6 ms during the download phase to
81-93 ms during replay, across three separate transfers, while a peer probed
at the same instant stayed flat at 3-5 ms.
Take the lock with `lock_owned().await` before `spawn_blocking` rather than
`blocking_lock()` inside it, so a blocking-pool thread is only ever occupied
by the read and never by waiting for a concurrent writer. Lock scope is
unchanged - the mutex was already held across the whole synchronous read.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WZjWYpKeYdGpKiLeEZU9oc
* remove unwanted comment
* Build the queue proxy send batch off the async runtime too
The previous commit moved the WAL read to the blocking pool, but measuring it
showed no improvement: the sender's health-check p99 during queue replay stayed
at ~90-100 ms. perf on the patched build explained why - optimizer/HNSW load is
ambient (~90% of CPU in both the download and replay phases, so not what makes
replay special), while one general-runtime worker burns 7.7% of all CPU during
replay against ~0.5% during download. That worker is doing the *send* half of
the loop, which is still synchronous:
- transfer_operations_batch() clones every operation in the batch
- forward_update_batch() converts each one into its gRPC representation
Both are full passes over up to MAX_BATCH_BYTES (32 MiB) of point data, on the
runtime that also answers internal health checks.
Move both to the blocking pool. WalBatch now holds its operations behind an Arc
so the batch can be shared into a blocking task without being copied first, and
the gRPC request construction is split out of forward_update_batch into
RemoteShard::build_update_batch_request so it can be spawned. The extracted
function keeps the original body and indentation, so the diff is the move plus
the plumbing rather than a reindent.
forward_update_batch has exactly one caller (the queue proxy), so this adds no
blocking-pool hop to the normal replication path.
Still on the runtime and not addressed here: the per-attempt clone of the
request inside with_points_client, and tonic's own protobuf encoding.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WZjWYpKeYdGpKiLeEZU9oc
* Use single clone, don't iterate manually
* Methods are droppable, don't hang on spawn_blocking with full runtime
* Build the WAL transfer batch once, reuse it across retries
The batch was deep cloned on every send so the original operations stayed
available for retries and for the one-by-one isolation path. Instead, move the
operations into the gRPC request once and retry on the prebuilt request, which
`with_points_client` already clones per attempt.
Stripping WAL indices and setting the force flag now happens in the WAL read
task, so no per-operation work runs on the async runtime.
Drop the pre-1.14.1 fallback that transferred operations individually, all
peers support batched updates by now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: generall <andrey@qdrant.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: timvisee <tim@visee.me>
Thread a `Populate` through the disk-resident id tracker's open paths
(`DiskMappingReader`, `DiskIdTracker::open`, `ReadOnlyDiskIdTracker`,
`ReadOnlyIdTrackerEnum`) so a `cached` placement primes the page cache with the
mapping files on load instead of leaving them to page in on demand. The
populate is derived from the segment config's placement at load time, clamped
by low-memory mode, in both the writable segment open and the read-only one.
The update-only lookup path keeps its transfer-nothing policy, and the
build-time open stays cold: the built segment is reloaded anyway.
`cached` is no longer rejected by validation.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Add `id_tracker: { memory: cold | pinned }` to CollectionParams,
CollectionParamsDiff and CreateCollection (REST + gRPC `IdTrackerParams`),
mirroring `payload: { memory }`. `cold` builds the disk-resident id tracker,
`pinned` the in-RAM immutable one. Unset keeps the current behavior: the
`serverless_compatible` feature flag decides.
The requested placement is persisted as an optional `id_tracker_memory` on
SegmentConfig (skipped when unset, so existing configs are unchanged); the
segment builder resolves it through `SegmentConfig::id_tracker_memory_placement`
instead of reading the feature flag directly.
The config mismatch optimizer rebuilds non-appendable segments whose effective
placement differs from the requested one. Appendable segments are skipped: they
always use the mutable tracker and get the current config when indexed.
`cached` is rejected by validation: the disk mapping reader has no
populate-on-open path.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* fix: do not claim an unfinished operation when flushing
A flush pass can capture a segment between the separately locked steps of
one update operation. Persisting it under that operation's version marks
the segment clean while the rest is still in memory, so every later pass
skips it and the WAL acknowledge moves past the operation.
Clamp what a flush claims to the last fully applied operation.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Fix rustfmt in alias_mapping test after merging dev
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com>
* Fix MMR pagination with offsets
(cherry picked from commit 6a778bbde8)
* Clamp MMR selection capacity at the candidate count
Move the guard into `maximal_marginal_relevance`, which every MMR caller
routes through, so the collection, edge and local-shard paths are covered
instead of just the collection one. The collection-level limit is then
plainly `limit + offset`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit a63d744131)
---------
Co-authored-by: mikemikimike <13286568797@163.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Remove MessagePack (rmp-serde)
The WAL switched from msgpack to CBOR in v0.3.5 (2021-07-11), so v0.3.4
is the last version that wrote msgpack entries. Drop the read fallback
kept for those entries, plus the remaining test and bench usages.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Drop unused fs4 dependency from collection
Not referenced anywhere in the crate. Still used by wal and common, so
the workspace entry stays.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
- 'mising' -> 'missing' in .github/review-rules.md (flagged by the
repo's own codespell config)
- 'Do not create segments larger this size' -> 'larger than this size'
in config.yaml, the optimizer builder/diff sources and the grpc
proto + generated rust comment
- 'bigger then' -> 'bigger than' in the query scorer rustdoc
- three duplicated-word rustdoc fixes ('override in in', 'any of of',
'and and')
* Direct call for shard transfer method and keys
* Reuse cardinality estimate in sparse plain search
* Avoid recounting available points in segment size info
* Avoid cloning segment config when updating quantization
* Avoid cloning search request for load profile
* Direct call for counting read-only segments
* Avoid re-reading point range for values count
* Direct call to check replica states when initializing collection
* Direct call to look up transfer on restart
* Direct call for shard replicas after snapshot recovery
* Direct call for local replica states in health check
* Direct call for payload index schema keys when applying state
* Direct calls for sharding method and key mapping when creating shard key
* Direct call to check if peer has shards
* Direct call for sharding method and keys when dropping shard key
* Avoid cloning collection params for group by ordering
* Avoid cloning collection params in local shard search
* Direct call for peer address when sending Raft messages
* Direct call for peer address in who_is
* Avoid cloning remote query batch request
* Avoid cloning operation in queue proxy update
* Avoid cloning gRPC search groups request
* Fetch cluster status once in cluster telemetry
* Direct call to validate transfer exists on finish
* Direct call for sharding method when dropping shard key
* Avoid cloning peer address map when listing peers
* Avoid cloning peer address map when adding peer to known
* Avoid cloning shard key mapping when routing writes with fallback
* Avoid cloning shard key mapping when checking resharding start
* Avoid cloning gRPC recommend groups request
* Avoid cloning operation when retaining forwarded point IDs
* Direct call for counting collections in telemetry
* Direct call to validate transfer exists on recovery
* Direct call for shard IDs by shard key
* Direct call for shard keys
* Direct call to check if peer has shards in consensus
* Direct call for replica state on transfer recovery
* Direct call to check for active replicas when routing writes with fallback
* Direct call to validate transfer exists on abort
The test asserts live WAL contents after applied updates, but the default
fixture uses flush_interval_sec = 0 so the flush worker can truncate
earlier rewritten records before the assertion runs.
Fixes#10423
* perf: skip retrieval in scroll when no payload or vectors are requested
Every scroll variant went through SegmentsSearcher::retrieve to build its
records, even when neither payload nor vectors were asked for. That is a
has_point lookup per id per segment plus a version and id resolution per
hit, only to yield records holding nothing but the id. The universal query
API always scrolls this way and fetches payload separately afterwards.
Build the bare records from the ids directly in that case. The retrieve
could only have dropped ids deleted in between, which the update lock held
across the scroll rules out.
* perf: fetch payload and vectors in the leaf of plain query requests
A query without prefetches and without rescoring is served by a single
leaf search or scroll whose result is returned as is. The planner still
built that leaf without payload or vectors and filled them in afterwards
through SegmentsSearcher::retrieve, which resolves every result id in
every segment again: the same cost #10312 removed from the search API,
paid once more at the end of each query.
Let the leaf carry the requested payload and vectors instead, so the
segment attaches them to the results it already holds by offset, and
clear the root plan so the fill step is skipped. Prefetch leaves and
rescored roots (MMR) are unchanged. As with the search API, this fetches
payload for each segment's candidates rather than for the merged top
`limit` alone.
* Use new_empty function
* fix: fetch payload and vectors in scroll leaves only (#10384)
A search leaf hydrates every segment's local top-k before merging, so
`with_payload` there multiplies payload I/O by the segment count — the
regression #6279 fixed and `test_payload_io_read_is_within_limit[query]`
guards. Scroll leaves retrieve once for the merged page, so they keep
fetching directly; search leaves stay bare and the root plan retrieves
for the final result.
Claude-Session: https://claude.ai/code/session_01SUWh5PqUeSefrUqUwXxU3E
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* Read vector runs that straddle a chunk boundary
Resolve a run into per-chunk parts instead of a single range, borrowing
when it lands in one chunk and copying when it spans two. The read
pipeline schedules one range per read, so a straddling run is read
outside it.
No writer produces such a run yet, so this changes nothing on its own.
It is what a reader needs before one does — including edge and
live-reload readers, which read files a different version wrote.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Place multivector runs without regard to chunk boundaries
Writers appended a multivector's inner vectors at the end of the row
space unless the run would cross a chunk boundary, in which case they
skipped the chunk tail — the batch writers padding the skipped rows with
explicit zero rows. That made chunk geometry part of the interface every
multivector storage had to reuse.
Runs now go at the end unconditionally and the chunked storage splits
the write across chunks, as it already did for a batch of single
vectors.
What is left of the geometry is a size cap: a multivector may not exceed
one chunk. It is fill-independent, so it constrains nothing about
placement, and it is what the volatile storage needs anyway — that one
returns a plain slice and so cannot serve a straddling run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Split a run at chunk boundaries in one place
Reading, writing in place and appending each derived the split from
`remaining_chunk_capacity`, so every one of them had to know that a run
does not necessarily fit where it starts.
`split_run` hands out the parts instead: one per chunk the run covers,
each carrying where it goes and how much of the run it takes. Nothing
asks how much room is left any more, and `get_chunk_offset` goes with it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Keep straddling runs on the read pipeline
Reading a straddling run outside the pipeline blocked the scheduling
loop on one read, which costs a round trip on a backend that fetches
remotely and drops the batch back to sequential.
A run is now scheduled as one read per chunk it covers. Parts complete
in any order, so each run holds what has landed until the last part
does, then hands the callback the stitched vectors. Runs taking a single
read carry the caller's data in the tag and never touch that table.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Stop capping a multivector at one chunk
The cap outlived its reason on disk, but the volatile storage still
needed it: its `get_many` handed out a slice of one chunk, so a run that
crossed a boundary had nowhere to come from. And since a volatile
storage is a target of the batched copy that builds a segment, dropping
the cap only on disk would have turned a rejected write into a failed
merge.
So the volatile storage splits and stitches too. Both are a few lines
each, and placing a run no longer skips a chunk tail, so `extend` is now
`insert_many` at the end of the storage.
Nothing user-facing moves: `MAX_MULTIVECTOR_FLATTENED_LEN` caps a
multivector at 1M elements, far inside a 32 MiB chunk, so the storages
only ever rejected what reached them unvalidated.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Schedule a single-read run without the queue
The scheduling loop resolved every run into the queue and then took it
straight back out, so the overwhelmingly common run — one that fits a
chunk — paid a push and a pop for nothing. It now goes to the pipeline
directly, and the queue holds only what a straddling run leaves behind.
Worth ~10% on the multivector read benchmark, and it collapses the
"top up, then take" pair into one decision. Extracting that bookkeeping
into helpers instead was measured and is much worse: the mmap pipeline
alternates one schedule with one wait, so the loop body is a few dozen
nanoseconds, and a helper carrying the cold map and stitching paths is
too big for the compiler to inline back into it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Repoint the multivector WAL-replay test at a live rejection
The test upserted a multivector too large for a storage chunk, which no
longer fails: the storages stopped capping one at a chunk. Nothing else
covered a multivector operation that only the apply path rejects.
A raw blob that is not a whole number of quantized records still does,
so the test now uses that, alongside its dense and sparse siblings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Test reading multivectors with legacy chunk-tail padding
Locks the compatibility contract that pre-straddle files — runs that
skip a chunk's leftover slots — still reopen as single-chunk borrows.
* chore: retrigger CI after flaky test-consensus-compose
* Move ReadTag into for_each_vector, its only user
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AN4Hgbd65gDhesthJk5bUY
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com>
* Add SetFlushInterval op to the model tester
Changes the collection's flush_interval_sec mid-run through the same path
update_collection takes (persist the optimizer-config diff, then recreate
the optimizers in the background). The model is untouched: what it perturbs
is the flush cadence, so how much of the workload is still WAL-only when a
restart hits, plus the worker stop/start race in on_optimizer_config_update.
Kept in FORCE_OFF for now: with the optimizer on it makes stale point state
visible within a few ops of the config change. Narrowed to
recreate_optimizers_background, see the comment on Swarm::BASE.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011AjmS5GFeGztfP3JqutnXj
* Keep SetFlushInterval enabled in the swarm
Drops it from FORCE_OFF so the divergence it surfaces is reachable without
--enable-force-off (which would also enable the broken vector-name ops).
The evidence moves from the FORCE_OFF comment onto the op's own doc.
The two optimizer-on harness gates now fail whenever the swarm draws the op.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011AjmS5GFeGztfP3JqutnXj
* Propagate proxied changes when unwrapping proxies on optimization failure
unwrap_proxy puts the wrapped segments back into the segment holder, so the
changes recorded on the proxy while the optimization ran (deleted points,
index and vector-name changes) have to reach the wrapped segment first. They
did not, so every point deleted or overwritten during the optimization kept
its pre-optimization copy live next to the new copy in the write segment, and
reads saw both: counts too high, scroll and search returning the stale copy.
The snapshot unproxy path already does this; the optimizer failure path was
the only place putting a wrapped segment back without it. It is reachable
whenever the shard outlives the cancellation, in particular an update_collection
that recreates the optimizers while an optimization is in flight.
Lock order is holder-then-updates, matching try_unproxy_segment: updates-then-
holder-write deadlocks against the snapshot path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011AjmS5GFeGztfP3JqutnXj
* Drop the stale failure note from the SetFlushInterval doc
The divergence it described is fixed in this branch.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011AjmS5GFeGztfP3JqutnXj
* test as well with 0s as flushing interval
* Fail optimization unwrapping when proxy propagation fails
Losing proxied deletes and index changes is data corruption, so return the
error instead of logging it: no proxy is unwrapped and the changes stay
served by the proxies. The cancelled-segment cleanup moves ahead of
unwrap_proxy so the orphan is still removed when that error fires.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CpoaAtGbAQAEuxEi8ScHc5
* Drop the model tester --flush-interval-sec flag
SetFlushInterval covers the interval now, so the run starts at the shipped
5s default (fixture::INITIAL_FLUSH_INTERVAL_SEC, still traced in the header)
and the ops move it from there. Also documents what 0 does now that it is a
generated value.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CpoaAtGbAQAEuxEi8ScHc5
* Fail snapshot unproxying when proxy propagation fails
Both paths logged the error and unwrapped anyway, dropping the deletes and
index changes that never reached the wrapped segment. Same reasoning as
unwrap_proxy in the optimizer.
try_unproxy_segment hands the lock back and leaves the proxy installed, the
failure mode its doc already describes: the caller keeps it in `proxies` and
unproxy_all_segments retries the propagation right after. unproxy_all_segments
returns before touching the holder, so the temp segment the surviving proxies
write into stays in place (remove_segment_if_not_needed only checks whether it
is empty and appendable, not whether a proxy still references it).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CpoaAtGbAQAEuxEi8ScHc5
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(strict-mode): enforce max_query_limit on scroll requests when limit is omitted (fixes#10373)
* style: rustfmt scroll query_limit for CI lint
---------
Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com>
* Add `AliasMapping::remove` and `AliasMapping::rename` methods
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Implement `ChangeAliases` operation
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `ChangeAliases` to replay-safety tests
Multi-action operations that rename an alias are skipped in the convergence
property: the current implementation does not replay them convergently, and
the machine reproduces that.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `ChangeAliases` tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add test-only `PeerMetadata::new` constructor
Test state needs peers at a version other than this build. The `version` field
is crate-private and `current()` is the only constructor, so gate the new one
on the `testing` feature and enable it for the `storage` test build.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Implement `UpdatePeerMetadata` operation
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `UpdatePeerMetadata` to replay-safety tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `UpdatePeerMetadata` tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Implement `UpdateClusterMetadata` operation
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `UpdateClusterMetadata` to replay-safety tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `UpdateClusterMetadata` tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Implement `SetQuotaConfig` operation
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `SetQuotaConfig` to replay-safety tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `SetQuotaConfig` tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Implement `TestSlowDown` and `TestTransientError` operations
Both are node-local: one sleeps, the other fails at random. They plan no
actions, like `Nop`, so the replay-safety properties have nothing to add.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `TestSlowDown` and `TestTransientError` tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fixup! Add `UpdatePeerMetadata` to replay-safety tests
* fixup! Add `UpdateClusterMetadata` tests
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(schema): declare enforced 1..=65536 bound on VectorParams.size
The REST layer enforces an upper bound of 65536 on VectorParams.size
via a custom validator (validate_nonzerou64_range_min_1_max_65536),
but custom validators contribute no bounds to the generated JSON
schema - so the published OpenAPI document only declared minimum: 1,
while DenseVectorConfig.size already documented both bounds.
Add an explicit #[schemars(range(min = 1, max = 65536))] attribute so
clients validating requests against the schema see the same contract
the server enforces, update docs/redoc/master/openapi.json
accordingly, and pin the bound with a unit test asserting the
generated schema.
Fixes#9942
* test(openapi): accept documented size bound as a rejection path
test_vector_dimension_limit asserted that an oversized VectorParams.size
reaches the server and returns the exact runtime 422 message. Now that the
enforced 1..=65536 bound is documented in the served OpenAPI schema (#9942),
request_with_validation rejects such payloads client-side before sending.
Accept either layer: a client-side jsonschema.ValidationError or the
server-side validation error.
* test(openapi): handle both rejection layers in dimension limit
pytest.raises only covered the client-side jsonschema rejection; if the
request reached the server instead, the test would fail on an unhandled
response. Use try/except around request_with_validation and assert the
server-side status and exact error message in the else branch.
* test(openapi): assert exact HTTP 422 on server-side rejection
A broad not-ok check would pass on any error status carrying the same
error text; pin the documented contract to 422.
* refactor(tests): address review feedback
Remove the unit test asserting the generated VectorParams schema shape -
it only restates the schemars attribute and adds maintenance cost.
Reduce test_vector_dimension_limit to its actual contract: an oversized
dimension is rejected by the documented OpenAPI schema before the
request is sent.
* Drop obsolete clippy large-error-threshold override
The 256 threshold was pinned for clippy 1.87 while tonic's `Status` was a
large error type. Upstream boxed its contents in `5de7bad` (hyperium/tonic#2253),
which is in the pinned 0.14.6 fork, so `Status` is now a single `Box` and the
default threshold of 128 passes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Remove stale clippy allows
These 11 allows no longer suppress anything under any of the three CI clippy
configurations (default, --all-targets, --all-targets --all-features).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Steer writes away from appendable segments at max_segment_size
* pick a write target that stays under the configured size cap, instead of
growing an appendable segment past it
* clamp the deferred points threshold to max_segment_size, treating a zero cap
as uncapped
* apply the same cap when replaying the WAL, so recovery matches live updates
* plumb the cap through the update worker and cover the silent-failure gaps
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Extract the per-segment capacity check into a helper
Lets has_appendable_segment_with_capacity short-circuit on the first segment
below the cap instead of collecting every eligible ID.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-up to #10287, which added `max`. Expressing a minimum still
required spelling out `(a + b - |a - b|) / 2`, the sign flip of the max
identity — drop the `neg` and you silently get a maximum instead. It also
only works for two operands and mentions each one twice, so the scorer
walks every sub-tree twice per candidate point.
The pair is what makes clamping expressible:
{"max": [0.0, {"min": [1.0, "$score"]}]}
`min` mirrors `max` throughout, and both guard helpers introduced in
#10287 already took an `operator: &str`, so they are reused unchanged: an
empty operand list is rejected at parse time rather than folding to
+infinity, and the Edge FFI rejects it at construction time. The result
needs no `is_finite` check, since `min` cannot produce a non-finite value
from finite inputs.
The unindexed-field walker shares one arm for `Max | Min` as the bodies
are identical, with a test pinning `min` separately so a later split
cannot silently drop it.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
11 names across 7 files. Renames that did not reach the comment above them
(further_searches for further_results, query_context for segment_query_context,
block_ranges for local_block_ranges, op for operation twice, request for
requests, max_threads for max_kmeans_threads), and 3 arguments that were
removed from a signature and left documented (is_on_disk, collection_params,
search_runtime_handle with timeout).
Documentation only, no behaviour change.
* feat: add a dedicated max operator to score formulas
Expressing a maximum in a score formula required spelling out the
arithmetic identity `(a + b + |a - b|) / 2`. That is easy to get wrong
(the `/ 2` is load-bearing), only works for two operands, and mentions
each operand twice, so the scorer evaluates every sub-tree twice per
candidate point.
`max` is variadic, mirroring `sum` and `mult`:
{"max": ["$score", {"mult": [0.5, "popularity"]}]}
Unlike `sum` and `mult`, `max` has no identity element for the empty
case, so an empty operand list is rejected at parse time rather than
folding to -infinity and scoring every point with a non-finite value.
The check lives in `ExpressionInternal::parse_and_convert`, which every
entry point passes through, and the Edge FFI additionally rejects it at
construction time to match how that crate validates elsewhere.
The result needs no `is_finite` check: unlike `log10`, `exp`, `div`,
`sqrt` and `pow`, `max` cannot produce a non-finite value from finite
inputs, so it follows the existing `sum` convention.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: cover max error propagation and datetime operands
An operand that fails must fail the whole expression rather than being
passed over in favour of a finite sibling. Covered with the failure both
before and after the finite operand: `mult` short-circuits on zero and
so can skip evaluating later operands, and this pins down that `max`
must not grow a similar shortcut that would swallow an error.
Also covers `max` over datetime operands, which reach the scorer through
a separate conversion to seconds, so that "score by whichever timestamp
is newer" is verified rather than assumed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Use a warmup baseline and adaptive timeout so the search is cancelled
mid-flight on both fast macos ARM runners and overloaded hosts, instead
of relying on a fixed 350ms cutoff.