With `enforce_internal_auth` enabled, the p2p server reused the public
auth layer. A read-only key or any JWT passed that layer, and the Raft
service does not check per-request access, so those credentials could
add a peer to consensus.
The internal auth layer now runs in a dedicated scope that accepts only
`api_key` and `alt_api_key`. The read-only key, JWTs, and missing or
wrong keys are rejected with Unauthenticated. Warn at startup when
internal auth is enforced with only a read-only key configured, since
peers would have nothing to authenticate with.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* Configurable key prefix for object storage snapshots
Snapshot object keys were always derived from the local snapshots path,
so every deployment sharing a bucket wrote under the same `snapshots/`
root. Each cloud config block gains an optional `prefix`, and objects
become `<prefix>/snapshots/...`.
The prefix is applied by wrapping the client in `PrefixStore`, so the
snapshot operations and the names returned by the API are unchanged.
Leading, trailing and repeated slashes are dropped, and an empty prefix
is a no-op.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Test that snapshot operations cannot escape the storage prefix
Hostile targets are handed to the cloud manager directly, past
`validate_snapshot_name`: parent references, absolute paths, encoded
slashes, backslashes and empty paths. Writes must stay under the prefix,
and objects planted outside the prefix must be invisible to list,
download, stream and delete.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Normalize empty prefix components
Refactor prefix handling to remove empty components.
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tim Visée <tim+github@visee.me>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* Add `match: { substring }` filter condition
Unindexed `text` and `text_any` matching became token-aware in #10341 and
#10593. Users who relied on the old raw substring behaviour get it back as
an explicit condition: `match: { "substring": "..." }` selects points with a
string value containing the given string, byte-wise and case-sensitive,
consistent with exact keyword and prefix matching.
Execution: a keyword index (with or without the `prefix` option) serves the
condition by scanning its value dictionary and uniting the postings of the
matching keys; cardinality reuses the prefix estimator, generalised into
`keys_union_cardinality`. The per-point checker goes through the forward
index. Without a keyword index the condition falls back to reading the
payload. Text, bool, integer and uuid indexes decline it.
Strict mode: the condition requires the `KeywordMatch` capability, so with
`unindexed_filtering_retrieve: false` it is rejected on unindexed and on
text-indexed fields and allowed on any keyword index.
API: `MatchSubstring` in the REST `Match` union with regenerated OpenAPI,
gRPC `Match.substring = 12`, edge python `MatchSubstring`, edge ffi
`Match::Substring`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Test substring fallback on a text-indexed field, document estimator params
A text index cannot serve `substring`, so on a field that has only a text
index the condition runs through the payload fallback; only strict mode may
reject it. Pin that in the OpenAPI suite and reword the strict-mode unit
test comment, which read as if the text index itself blocked the query.
Also spell out what `keys` and `postings` mean in `keys_union_cardinality`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Serve `match: { substring }` from the keyword key dictionary
The condition used to enumerate keys through `MapIndexRead::for_each_value`,
which on the on-disk variant drags the whole `value_to_points` file through
`for_each_entry`, plus one random read per matching key for its postings
count. Query planning paid that scan in full, before deciding whether to use
the clause at all.
Route it through the `prefix_index.bin` key dictionary instead: front-coded
keys with their postings counts inline, no postings. Estimation now reads
keys only and never touches `value_to_points`; filtering takes the matched
key list and resolves postings in one batched read, as prefix matching
already does.
This makes the `prefix` option a requirement: a keyword index without it has
no key dictionary, so it declines the condition and falls back to the payload
scan, the same as a text index. Strict mode follows — substring now infers
`KeywordPrefix`, so `unindexed_filtering_retrieve: false` names
`keyword (with prefix: true)` as the index to create.
`PrefixIndex::for_each_key` reads blocks in ~1 MiB chunks rather than the
whole key section at once: a substring cannot be pruned by the block index,
so the one-shot read would grow with the dictionary.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Sync MatchSubstring OpenAPI description with Rust docs
After routing substring matching through the prefix key dictionary,
the schema docstring required regenerating so OpenAPI stays consistent.
* Scan the map index keys for substring match without a dictionary
Without the keyword dictionary a substring condition was declined by the
field index and left to the per-point condition checker, which reads the
forward index for every candidate point. Enumerate the distinct keys of
`values_to_points` instead: the same one-pass-over-distinct-values shape as
the dictionary scan, only over a structure that interleaves keys with their
postings. Filtering and cardinality estimation are then always served, so
the condition can act as a primary clause on a plain keyword index.
Prefix matching keeps its per-point fallback: an ordered dictionary is what
makes a prefix a bounded range, and enumerating every key to answer one is
not a trade worth making implicitly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Reject substring matching in strict mode
A substring condition is answered by looking at every distinct value of the
field: no index gives it a bounded access path, so there is no index a user
could create to make it affordable. Reject it under strict mode instead,
wherever a filter reaches verification — read and write filters, nested
sub-filters, and prefetch filters.
Filter limits are now checked before the unindexed-field check, so the
rejection is not reported as "create an index for this key", advice that
would lead to the same rejection afterwards.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Update the strict mode substring test to the new rejection
The test asserted that substring filtering under strict mode asks for a
keyword index with the `prefix` option. It is now rejected whatever index
the field carries, so every case in the test gets the same answer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Estimate a substring condition without scanning
Counting the keys a substring matches costs the same scan as answering the
condition, and `filter` then repeats it to collect those keys. Report the
uninformed estimate instead — the one an unindexed condition has always
reported — and keep the primary clause, so the scan happens once, in
`filter`, and only when the planner picks the condition to drive iteration.
With no counts to collect, `substring_scan` collapses into `substring_keys`:
the in-RAM variants no longer look up a posting count per matched key.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Don't to parse everything as UTF-8
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com>
Co-authored-by: timvisee <tim@visee.me>
Snapshot storage was limited to S3 although the object_store dependency
already ships the GCS and Azure backends. `snapshots_storage` now accepts
`gcs` (alias `gcp`) and `azure`, configured through `gcs_config` and
`azure_config` blocks next to the existing `s3_config`. The legacy S3
shape is unchanged.
Client construction is split into one builder per backend, all sharing
the Qdrant user agent and the plain-HTTP rule for `http://` endpoints.
Startup warns when a config block for an unselected backend is present.
The e2e snapshot recovery test is parameterized over the cloud backends.
The GCS case is skipped because fake-gcs-server does not implement the
XML multipart upload API that object_store uses for GCS uploads.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* test(consensus): prove leader removal can stall
A departing leader can clear peer addresses before its queued commit
notification reaches the surviving voter. Control the removal append
and its acknowledgement to reproduce that ordering without fixed sleeps
or stopping the leader process.
Check the actual notification failure and the survivor's commit index.
Keep follower-removal and three-voter controls, and cover both the default
removal wait and timeout=60.
* test: timeout=60s with follower or leader
Signed-off-by: Anton Antonov <anton.synd.antonov@gmail.com>
* test: fail when leader removal loses its commit
Require the surviving voter to receive the committed removal and remain
operational. Fail explicitly when the commit notification never reaches
its gate instead of treating the stuck cluster as success.
Both two-node leader-removal cases now fail against the existing bug.
The follower-removal and three-node controls pass.
---------
Signed-off-by: Anton Antonov <anton.synd.antonov@gmail.com>
* test: add a gated proxy for peer RPCs
Pause one selected internal request while other peer traffic continues.
Preserve payloads, metadata, deadlines, and cancellation so consensus
tests can control transfer timing without blocking unrelated requests.
Cover forwarding, independent gates, and cleanup with socket tests.
* test: connect peer proxies to consensus clusters
Let consensus tests route internal RPCs through request gates. Keep each
proxy alive across peer restarts so advertised addresses remain stable,
and close all proxies during test cleanup.
Wait for the upstream gRPC connection before returning from proxied
startup. Verify consensus progress during a held WAL-delta request,
recovery data, and restart behavior with both URI configuration modes.
* test: fix potentially misleading peer proxy method names
Explicitly state the guarantees, or lack of.
Signed-off-by: Anton Antonov <anton.synd.antonov@gmail.com>
* test: add support for hold_snapshot_download
Removes flakiness from snapshot-related consensus tests too
Signed-off-by: Anton Antonov <anton.synd.antonov@gmail.com>
* test: improve asserts when force deleting peer
Actually verify survivors recover and retain the expected data.
Making sure no data loss happens.
Signed-off-by: Anton Antonov <anton.synd.antonov@gmail.com>
* test: add OsError socket handling + explicit wal_delta tests
Signed-off-by: Anton Antonov <anton.synd.antonov@gmail.com>
* test: reject zero as a defined consensus leader
* test: recheck leader agreement on each poll
After a restart, the leader can change during election. Let the cluster
wait resample the leader on each poll and require agreement on a nonzero
leader before the snapshot test starts its transfer.
Keep explicit leader checks for existing callers, membership-size checks,
and the existing timeout. Cover election changes and offline peers.
* test: verify independent snapshot download gates
* test: use a positive peer connection deadline
* test: cover recovery after the removed source exits
* chore: add clarifying comment on timeout=0 usage
It's not obvious at first why it's like so.
Signed-off-by: Anton Antonov <anton.synd.antonov@gmail.com>
* test: share consensus response gates
Move response gates, their tests, and Raft decoding from the leader
removal proof into the base test infrastructure. Both removal scenarios
can then use the same successful-response check.
* test: support selective RPC blocking
Keep a removed source unaware of membership changes while its transfer
continues. Block its Raft traffic in both directions so election
attempts cannot disrupt survivor recovery.
* test: make source removal scenarios deterministic
Separate recovery after source exit from late data sent by a removed
source. Require a successful receiver response in the late scenario,
and retain complete data and replica-state checks in both cases.
* test: refactor timeouts and deadlines
* Cancellation happens after observing the intended phase, without an RPC deadline.
* Separate deadline tests cover held requests, upstream work, and held responses.
Signed-off-by: Anton Antonov <anton.synd.antonov@gmail.com>
* test: bound peer probes and removal requests
Give cluster probes and peer removal finite client timeouts so a stalled
HTTP request cannot leave the test waiting indefinitely.
* test: separate RPC release from termination
Keep the upstream handler blocked until the test releases it or the RPC
terminates. Use a separate termination event for cancellation assertions,
and release the handler during teardown instead of racing a fixture timer.
* test: use monotonic polling deadlines
Measure elapsed polling time with a monotonic clock so system clock
adjustments cannot shorten or extend the wait.
* test: allow more time to observe proxy events
Allow ten seconds for proxy observations and ordinary test requests.
Event and future waits still return as soon as they complete. Keep the
one-second expiry tests and document the HTTP deadline setup race.
* test: bound leader and replication requests
Limit how long leader lookup and transfer submission wait for an HTTP
response. A stalled submission must fail so the test can release its
transfer gates and clean up the peers.
* test: preserve readiness failures in diagnostics
Catch request failures while collecting cluster diagnostics, including
read timeouts. Report the original readiness failure instead of replacing
it with a diagnostic error.
* test: assert points calls for the correct collection
Signed-off-by: Anton Antonov <anton.synd.antonov@gmail.com>
* test: make sure check_cluster_size and check_leader cannot stall
Have an explicit timeout.
Signed-off-by: Anton Antonov <anton.synd.antonov@gmail.com>
* test: retry timeouts during initial leader lookup
Treat request timeouts as retryable while discovering the expected leader,
matching the subsequent leader and membership checks. Keep polling after
a transient timeout instead of aborting the cluster-status wait.
* test: ensure batch data is different
Signed-off-by: Anton Antonov <anton.synd.antonov@gmail.com>
---------
Signed-off-by: Anton Antonov <anton.synd.antonov@gmail.com>
The test compared raft commit indices across peers before issuing shard
cleanup. A follower advances its commit index before it applies the entry,
so cleanup could still race the follower applying CommitRead, which
invalidates running clean tasks and makes the endpoint return 500.
Wait for every peer to report the applied `read_hash_ring_committed`
resharding stage via telemetry instead. Reading it takes the same shard
holder lock as the consensus handler, so an observed stage is fully applied.
Repurpose the unused stage check helper to read the stage from telemetry,
since the `comment` field it read no longer exists.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* tests: use local cache when building image for e2e tests
* tests: make e2e local-cache image build safe under pytest-xdist
- Serialize builder creation and the build with a `filelock` lock (the
approach recommended by pytest-xdist for session fixtures), since
every xdist worker runs the session fixture; workers that wait on the
lock reuse the image built by the first one instead of rebuilding.
- Keep the buildx builder between sessions instead of removing it in a
finalizer, which also avoided one worker deleting a shared builder.
- Fix README: CI pre-builds the image in the workflow (there is no
`build-e2e-image` artifact); note the cache dir is never pruned.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
test_recover_from_snapshot_2 and test_upload_snapshot_2 start snapshot
recovery on a freshly joined peer as soon as it lists the collection. The
collection appears once the creation entry is applied, while the peer is
still replaying the rest of the raft log, including the removal of the
killed peer. Recovery then decides which other replicas to remove or mark
dead from that stale local view and drops a healthy replica, leaving a
shard with a single replica.
Add a helper that waits until all peers share the same commit index and
have no pending operations, and use it in both tests before recovering.
Also fix a misleading comment in the recovery replica cleanup branch.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* Reject snapshot upload without collection config before loading it
The raw IO error from `CollectionConfigInternal::load` embedded the
server-side temporary path in the API response. Check for the file first
and return a fixed bad-input error instead.
Part of #10553
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A3o9eWSZNMa6WAs5F2HMZC
* Fix missing-config check to run before restore_snapshot loads config
The path-leak guard lived after Collection::restore_snapshot, but that
function already calls CollectionConfigInternal::load and surfaced the
temp path as a 500. Require a regular config.json file before loading,
and cover a directory-shaped config entry in the openapi test.
* Use a valid empty TAR in the missing-config snapshot upload test
Avoid depending on malformed-archive handling; exercise the missing
collection-config path with a real TAR that has no entries.
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com>
Thread a `Populate` through the disk-resident id tracker's open paths
(`DiskMappingReader`, `DiskIdTracker::open`, `ReadOnlyDiskIdTracker`,
`ReadOnlyIdTrackerEnum`) so a `cached` placement primes the page cache with the
mapping files on load instead of leaving them to page in on demand. The
populate is derived from the segment config's placement at load time, clamped
by low-memory mode, in both the writable segment open and the read-only one.
The update-only lookup path keeps its transfer-nothing policy, and the
build-time open stays cold: the built segment is reloaded anyway.
`cached` is no longer rejected by validation.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Add `id_tracker: { memory: cold | pinned }` to CollectionParams,
CollectionParamsDiff and CreateCollection (REST + gRPC `IdTrackerParams`),
mirroring `payload: { memory }`. `cold` builds the disk-resident id tracker,
`pinned` the in-RAM immutable one. Unset keeps the current behavior: the
`serverless_compatible` feature flag decides.
The requested placement is persisted as an optional `id_tracker_memory` on
SegmentConfig (skipped when unset, so existing configs are unchanged); the
segment builder resolves it through `SegmentConfig::id_tracker_memory_placement`
instead of reading the feature flag directly.
The config mismatch optimizer rebuilds non-appendable segments whose effective
placement differs from the requested one. Appendable segments are skipped: they
always use the mutable tracker and get the current config when indexed.
`cached` is rejected by validation: the disk mapping reader has no
populate-on-open path.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* Fix flaky shard snapshot API CI readiness race
Run the prebuilt binary and poll /readyz instead of cargo run + fixed sleep, which can miss startup when cargo recompiles.
* Move shard snapshot API CI runner into a dedicated script
Keep workflow YAML thin by starting Qdrant, waiting for /readyz, and invoking shard-snapshot-api.sh from tests/shard-snapshot-api-tests.sh.
* docs(schema): declare enforced 1..=65536 bound on VectorParams.size
The REST layer enforces an upper bound of 65536 on VectorParams.size
via a custom validator (validate_nonzerou64_range_min_1_max_65536),
but custom validators contribute no bounds to the generated JSON
schema - so the published OpenAPI document only declared minimum: 1,
while DenseVectorConfig.size already documented both bounds.
Add an explicit #[schemars(range(min = 1, max = 65536))] attribute so
clients validating requests against the schema see the same contract
the server enforces, update docs/redoc/master/openapi.json
accordingly, and pin the bound with a unit test asserting the
generated schema.
Fixes#9942
* test(openapi): accept documented size bound as a rejection path
test_vector_dimension_limit asserted that an oversized VectorParams.size
reaches the server and returns the exact runtime 422 message. Now that the
enforced 1..=65536 bound is documented in the served OpenAPI schema (#9942),
request_with_validation rejects such payloads client-side before sending.
Accept either layer: a client-side jsonschema.ValidationError or the
server-side validation error.
* test(openapi): handle both rejection layers in dimension limit
pytest.raises only covered the client-side jsonschema rejection; if the
request reached the server instead, the test would fail on an unhandled
response. Use try/except around request_with_validation and assert the
server-side status and exact error message in the else branch.
* test(openapi): assert exact HTTP 422 on server-side rejection
A broad not-ok check would pass on any error status carrying the same
error text; pin the documented contract to 422.
* refactor(tests): address review feedback
Remove the unit test asserting the generated VectorParams schema shape -
it only restates the schemars attribute and adds maintenance cost.
Reduce test_vector_dimension_limit to its actual contract: an oversized
dimension is rejected by the documented OpenAPI schema before the
request is sent.
Mirror the fix from #9938 for test_upload_snapshot: assert every shard
has n_replicas active replicas across local+remote shards instead of
assuming peer 0 always sees exactly 2*n_replicas remote shards.
Follow-up to #10287, which added `max`. Expressing a minimum still
required spelling out `(a + b - |a - b|) / 2`, the sign flip of the max
identity — drop the `neg` and you silently get a maximum instead. It also
only works for two operands and mentions each one twice, so the scorer
walks every sub-tree twice per candidate point.
The pair is what makes clamping expressible:
{"max": [0.0, {"min": [1.0, "$score"]}]}
`min` mirrors `max` throughout, and both guard helpers introduced in
#10287 already took an `operator: &str`, so they are reused unchanged: an
empty operand list is rejected at parse time rather than folding to
+infinity, and the Edge FFI rejects it at construction time. The result
needs no `is_finite` check, since `min` cannot produce a non-finite value
from finite inputs.
The unindexed-field walker shares one arm for `Max | Min` as the bodies
are identical, with a test pinning `min` separately so a later split
cannot silently drop it.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: add a dedicated max operator to score formulas
Expressing a maximum in a score formula required spelling out the
arithmetic identity `(a + b + |a - b|) / 2`. That is easy to get wrong
(the `/ 2` is load-bearing), only works for two operands, and mentions
each operand twice, so the scorer evaluates every sub-tree twice per
candidate point.
`max` is variadic, mirroring `sum` and `mult`:
{"max": ["$score", {"mult": [0.5, "popularity"]}]}
Unlike `sum` and `mult`, `max` has no identity element for the empty
case, so an empty operand list is rejected at parse time rather than
folding to -infinity and scoring every point with a non-finite value.
The check lives in `ExpressionInternal::parse_and_convert`, which every
entry point passes through, and the Edge FFI additionally rejects it at
construction time to match how that crate validates elsewhere.
The result needs no `is_finite` check: unlike `log10`, `exp`, `div`,
`sqrt` and `pow`, `max` cannot produce a non-finite value from finite
inputs, so it follows the existing `sum` convention.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: cover max error propagation and datetime operands
An operand that fails must fail the whole expression rather than being
passed over in favour of a finite sibling. Covered with the failure both
before and after the finite operand: `mult` short-circuits on zero and
so can skip evaluating later operands, and this pins down that `max`
must not grow a similar shortcut that would swallow an error.
Also covers `max` over datetime operands, which reach the scorer through
a separate conversion to seconds, so that "score by whichever timestamp
is newer" is verified rather than assumed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
`test_reinit_removed_peer_readyz_ignores_old_cluster` restarted the
reinitialized peer on a fresh port. A changed `--uri` makes the peer
announce its new address to every address-book entry, including the
injected old first peer, which re-adds it to the *old* cluster as a
learner and starts replicating its log to it. Normally the restarted
peer is a term ahead and ignores those messages, but when the reinit
run's hard state save did not finish before the kill, it restarted at
the old term, accepted the old leader, took its log (commit 13 > 12)
and failed the guard.
The scenario is a plain restart, so keep the URI; then nothing is
announced and only the `/readyz` membership filter is exercised.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Unary inverse hyperbolic cosine, parallel to sqrt/ln/exp/log10, in REST,
gRPC, and edge (FFI + Python) interfaces. Inputs below 1 produce the same
NonFiniteNumber error as an invalid sqrt or ln.
Closes#10186
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* test: cover TurboQuant, turbo4 datatype and keyword prefix in compat data
Extend the storage compatibility fixture with vector and payload index
features that landed since the generator was last updated:
* TurboQuant quantization, one collection per persisted blob layout
(bits1_5 and the default bits4)
* the turbo4 storage datatype, on both dense and multivector storage
* the keyword index `prefix` option, plus a matching prefix scroll in the
query battery
Sparse vector configs reject the turbo4 datatype, so create_collection omits
the sparse datatype for that collection rather than forwarding it. This is a
no-op for every other collection.
Archives are generated once per release and keep the collection set of their
own generation, so expected collections are now resolved per version and the
new ones are only required from v1.19.0 onward. The prefix scroll stays
ungated: archives without the prefix index answer it by scanning.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: fail instead of skip when a compatibility archive is missing
A 404 from the compatibility bucket means the archive was never published,
which no amount of retrying fixes. Skipping it reported the version as
covered while nothing ran, so a pull request adding a version could stay
green with its new coverage never executing.
Fail on 404 and keep skipping connection resets and timeouts, so a bucket
outage still does not turn unrelated pull requests red.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Tolerate not-yet-ready collection upserts in test_rejoin_cluster and give
JWT snapshot uploads more headroom while still bounding auth-rejection hangs.
The crashing peer was picked positionally, so it could be the raft leader.
The staging crash exits inside the apply of the `Dead` entry while raft
messages leave through an async send queue, so a crashing leader takes the
append carrying the new commit index down with it. The other live peer is
then left holding that entry appended but uncommitted and, as one voter out
of three, can never commit it: it never aborts resharding, and the test
times out waiting for its resharding state to clear — the restart that
would restore quorum only comes after that wait.
Pick the victim after the receiver is killed instead: wait for the two live
peers to agree on a live leader, keep that leader as the survivor and crash
the follower. The survivor then commits and applies the abort locally, with
no commit index left to escape a dying process.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add /profiler/consensus_lag to measure apply lag between peers
Raft commit index advances on a peer whose apply loop is stalled, so the
existing signals - `raft_info.commit` and the `all_nodes_have_same_commit`
test helper - report a stuck peer as healthy. Nothing exposes how long a
peer has been behind at *applying* entries, which is what shard transfer's
`await_consensus_sync` barrier actually waits on.
Each peer now keeps a ring of the last 32 entries it applied, stamped with
its own wall clock and the time that entry took to apply. The ring is in
memory on ConsensusManager, not in Persistent, so the on-disk format is
untouched.
`/profiler/consensus_lag` collects those rings from every peer over a new
internal RPC and lines them up on the entry indices they share. Each entry
is measured from whichever peer applied it first, so a lag is never
negative; the peer that is first can differ per entry, so the baseline is
per entry rather than a single chosen peer. Entries only one peer still
remembers are excluded, otherwise a peer would be measured against itself.
A peer stalled part-way through an entry keeps healthy lag statistics -
everything it did apply, it applied on time - so the report carries
`behind_entries` and `newest_applied_age_ms` alongside, which is what
actually exposes the stall.
Peers that fail or time out are listed rather than failing the request: a
partial answer is more useful than none when the point is to find a peer
that stopped answering.
The endpoint follows `/profiler/slow_requests`: manage access, and outside
OpenAPI, so no endpoint-count or ACTION_ACCESS guard applies.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Move applied-entry log into its own module
Keeps the new code out of files that are already large. The ring, its entry
type and the snapshot served over RPC move to
`content_manager/consensus/applied_log.rs`, alongside the other consensus
internals; `ConsensusManager` is left with a field, an accessor and the one
`record` call in the apply loop.
The grpc encoding moves next to the decoding it mirrors, in
`common/consensus_lag.rs`, leaving the internal service handler three lines
instead of thirty.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Test that a consensus stall is still in the report after the peer catches up
* Take each peer's applied index from consensus state, not its ring
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: tellet-q <elena.dubrovina@qdrant.com>
* feat: global quota API
Memory and disk are node-wide resources, so configuring their thresholds
per collection through strict mode makes little sense. Move them behind a
single cluster-wide `QuotaManager`.
The quota config is seeded from `storage.quotas` in the settings (and so
from env vars), overridden by `quota.json` in the storage directory, and
updated cluster-wide through a new `SetQuotaConfig` consensus operation
which rewrites that file on every peer. Raft snapshots carry it too, so a
peer that joins by snapshot picks it up.
Quotas are enforced wherever the strict mode memory and disk checks used
to run, but no longer gated behind `strict_mode.enabled`: a value set in
an enabled strict mode config still wins per resource, the quota is the
default. Rejections name both the condition that tripped and the config
that governs it.
`GET /quotas` reports the config plus current utilization to global read
users; `PUT /quotas` replaces it for global manage users.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: cover the quota endpoints in the API consistency checks
`test_all_rest_endpoints_are_covered` and the OpenAPI endpoint count both
break on any new REST endpoint. Add `GET`/`PUT /quotas` to `ACTION_ACCESS`
with their JWT access tests, and bump the expected API count. The quota
endpoints stay out of `REST_ENDPOINT_WHITELIST`: that list is for
data-plane endpoints reported per-endpoint in metrics.
Also add a Raft snapshot CBOR compatibility test — snapshots are exchanged
between peers of different versions during a rolling upgrade, so
`quota_config` must be absent-tolerant in both directions.
Review feedback: persist through `SaveOnDisk`, which already implements the
write-before-swap protocol this was doing by hand; validate the config at
both persistence boundaries, since a hand-edited quota file or a config
arriving through consensus does not pass the REST handler's validation, and
a `0%` limit would reject every update forever.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: assert seeding a quota from invalid settings persists nothing
Follow-up to review feedback claiming `SaveOnDisk::load_or_init` writes the
init value before it is validated. It does not — only `SaveOnDisk::new`
persists — but the property matters: were seeding to persist first, invalid
settings would leave a `quota.json` that fails validation on every
subsequent start, and the node could only be recovered by deleting it by
hand. Pin it down with a test.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor: make QuotaManager the single reader of memory and disk
The quota checks measured memory and disk themselves, while the optimizer
and the WAL disk watcher each called `fs4::available_space` behind their
own ad-hoc caches. Fold all of it into QuotaManager: it owns the readings,
the freshness policy, and the limits they are compared against.
Moves the module to `lib/shard`, since the optimizer sits below `storage`
and has to reach it; `storage::quota` re-exports it, so consensus, the
`/quotas` API and StorageConfig are unchanged. The manager is installed as
a process singleton by TableOfContent, ahead of loading any collection.
- Callers hand in QuotaLimits overrides instead of a StrictModeConfig, and
an override can now only tighten. A collection-level admin could raise
`max_disk_usage_percent` past a cluster-wide limit that needed global
manage rights to set; ties resolve to the quota so the rejection names
the knob that actually has to change.
- Measurements are cached for 5s, but a reading at or above its limit is
never reused: a rejected client retries, and freeing the resource has to
take effect on the next request rather than a TTL later.
- `fits_on_disk` sizes an optimization against physical free space only,
never the configured limits. Optimizations are what free a full disk, so
the quota must not be what stops one.
- `percent_of` widens to u128 instead of saturating the multiply, which
under-reported utilization (failing open) above ~184 PB.
- StorageConfig::quotas is optional; absent means no quota is enforced.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: don't recover dead replicas onto a node at a resource limit
Recovering a dead replica pulls a whole copy of its shard onto this node.
If it is already at its memory or disk quota that transfer cannot finish,
and starting it only pushes the node further past the limit. Skip it and
reconsider on a later sync, once the resource frees up.
Adds QuotaManager::check_capacity for work that lands bytes here without
being an update. Unlike fits_on_disk the configured limits do apply:
taking on a replica is not what frees a full node, so there is no deadlock
to avoid by letting it through.
The check is hoisted out of the per-shard loop because a node over its
limit re-measures on every call, so checking per dead shard would cost a
statvfs each. It is free when no quota is configured.
Also trims the comments across the quota module, which had grown well past
what the code needs.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: drop trivial and duplicated quota tests
Six tests removed, ~140 lines, with no loss of coverage:
- a_rejection_names_the_knob_that_has_to_change asserted that a format!
contains its own literals; the message is covered end-to-end by the
override test and by test_global_quota.py.
- a_node_over_its_quota_has_no_capacity_to_take_on_a_replica was 30 lines
for check_capacity, a one-line delegation to check_update the test above
it already calls.
- a_rejecting_measurement_is_never_served_from_the_cache duplicated the
meter test, which proves the same rule with an injected reader instead
of inferring it from the real filesystem.
- free_space_is_reported_without_enforcing_anything covered a one-line
accessor, and its point is what the fits_on_disk test is for.
- The two resolve tests and the three meter tests each collapse into one.
DiskFit::Unknown keeps its coverage as two lines inside the fits_on_disk
test rather than its own fixture. The snapshot compat pair becomes one
test: the second only asserted cluster_metadata.is_empty(), which says
nothing about quotas — the real check was the deserialize, now an expect
that states it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: re-measure free space as the disk fills, and drop a Windows-only assert
Two CI failures, both from this branch.
e2e test_low_disk: the DiskUsageWatcher I replaced escalated to checking
on every call once free space fell below 512 MB. Folding it into the quota
manager lost that — available_bytes passed no limit, so a reading was
reused for the full 5s however little space was left. On a disk filling as
fast as that test fills it, 5s blind is enough to actually run out and the
WAL write dies instead of returning "No space left on device".
available_bytes now takes a watch_below level and never reuses a reading
under it, which is what the old ladder was expressing. The watcher passes
max(min_free, 512 MB), so the escalation point is back; above it the 5s
cache still costs fewer syscalls than the old 128-call ladder. fits_on_disk
gets the same rule by passing required_bytes, so a merge that does not fit
re-checks rather than sitting on a stale sample.
Windows: fits_on_disk on a missing path was asserted to be Unknown, but
GetDiskFreeSpaceEx resolves up to the containing drive and succeeds — as
common::disk_usage's own test documents. Dropped; the branch is a two-line
else and is not portably reachable.
Also renames an_optimization_is_sized_against_the_disk_not_the_quota, which
needed explaining to be understood.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: report the quota config in telemetry
Reads it from the quota manager rather than the settings, so it is the
config the node is actually enforcing: a peer that missed a consensus
update reports what it is applying, not what the cluster agreed on.
Gated on global access, the same access `GET /quotas` requires, and left
out of `PeerTelemetry` — a quota is per-node state, so each peer reports
its own.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: regenerate OpenAPI, and cover the quota in the telemetry key sets
Two CI failures from the previous commit.
Referencing QuotaConfig from TelemetryData moves its definition earlier
in `components/schemas`, because TelemetryData is generated ahead of
QuotaStatus. Regenerated rather than hand-patched, so the schema is a
pure move.
test_telemetry_detail asserts the exact set of top-level telemetry keys.
The quota is reported at every level, including 0 — it is three scalars,
it is the default the endpoint serves, and it is what explains an update
being rejected — so both key sets gain it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: remove `max_disk_usage_percent` from strict mode
Disk is a node-wide resource, so a per-collection percentage of it never
meant anything a caller could act on: the limit describes how full the
*node* is, and which collection the write happens to target has nothing
to do with it. The global quota is where it belongs.
It shipped in 1.18.2 without documentation, so this drops it outright
rather than deprecating. Removal is soft in every direction: StrictModeConfig
has no `deny_unknown_fields`, so a client still sending it gets it ignored
rather than a 400, and the same struct deserializes the persisted collection
config, so collections created on 1.18.2+ keep loading. Proto field 22 is
reserved so the number is never reused.
The e2e test becomes a quota test — the fixture and the timing are the
interesting parts and they carry over unchanged; only how the threshold is
configured differs.
`max_resident_memory_percent` was documented and stays for now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor: enforce the strict mode memory limit outside the quota
`max_resident_memory_percent` was folded into the quota as an override,
which meant the quota check had to know about strict mode, and retiring
the setting would mean unpicking `EffectiveLimit` and `LimitSource` from
the resolution logic.
It is now a check of its own in `verification/mod.rs`, next to the strict
mode checks it belongs with, borrowing only the measurement from the quota
manager — which stays the node's single reader of process memory, so both
checks still share one reading. Deleting the setting later is deleting one
function and its one caller.
`QuotaManager::check_update` takes no arguments and consults the quota
alone. A collection can still only tighten the limit for itself, because
its own check runs in addition rather than in place of the quota's, and
each rejection now names the config that has to change without having to
carry a `LimitSource` to say so.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor: enforce the quota on the update path, not in strict mode
The quota check sat inside `check_strict_mode_toc_batch` only because that
was the one place holding the collection's strict mode config. It doesn't
need one any more, and the placement had a real cost: coverage depended on
each handler remembering to ask for a strict mode check, and four of the
internal update RPCs do — `sync_internal`, which moves the most bytes onto
a node, does not.
It now runs in `Collection::update_from_client` and `update_from_peer`,
which every update passes through. `update_from_client` checks ahead of the
shard split, so an operation is accepted or refused whole rather than
landing on some shards and being refused by others.
Classification moves with it, from ~10 `consumes_memory` impls on request
DTOs to one exhaustive `CollectionUpdateOperations::consumes_quota`. The
internal enum has variants — raw upserts, conditional upserts, the syncs —
that have no client-facing request type, so per-DTO impls structurally
could not classify them.
Shard-transfer syncs stay excluded, as they are today: a transfer is sized
up once before it starts, and refusing its batches partway abandons work
that is nearly done only for it to restart from the beginning.
Index and named-vector creation reach shards through consensus, past this
check — a peer must not refuse what the cluster agreed to — so they keep
their pre-consensus check, now against the quota manager directly.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: deprecate `max_resident_memory_percent` in strict mode
Same reason the disk threshold went: memory is node-wide, so a
per-collection percentage of it caps how full the *node* is, which has
nothing to do with which collection is being written to. The node-wide
quota caps it once for everything.
Unlike the disk threshold this one shipped documented, in 1.18.0, so it
keeps working — as a limit a collection can tighten for itself, never lift
— and gets the usual markers: `#[deprecated]` on both Rust structs,
`[deprecated = true]` on proto field 21, and `deprecated: true` in the
OpenAPI schema, which schemars derives from the attribute.
The note names 1.21 as the removal. Recording a version matters here: the
audit in docs/plans/overdue-deprecations.md found that this repo has never
written a removal deadline down, and members of the 1.15.0 deprecation
batch are still in tree.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: reconcile the quota readers with #9891#9891 landed effective (cgroup) figures in telemetry while this branch was
making QuotaManager the single reader of memory and disk. Two collisions,
neither of which git sees.
`segment::utils::mem::total_memory_bytes` is now a shared accessor with a
5s TTL, so a cgroup resize is picked up. The quota module had its own
`OnceLock` copy that froze the value at startup — exactly what #9891 set
out to fix — so it delegates to the shared one instead.
Telemetry's new `disk_size` called `common::disk_usage::disk_usage`
directly. That reader lost its TTL cache on this branch when the caching
moved into the quota manager's meter, so it would have taken an uncached
`statvfs` on every telemetry request, and it put a second disk reader back
in the tree. It goes through `QuotaManager::disk_capacity_bytes` now,
sharing the reading the quota check already takes. Verified it still
reports the storage filesystem, matching `df`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor: split the quota manager by what each half does
`manager.rs` had grown to 450 lines holding three separate jobs: owning
the config and its file, taking the readings, and comparing one against
the other.
- `manager/store.rs` — the `Store` enum, `QUOTA_CONFIG_FILE`, and config
validation, which is now the store's own business rather than something
every caller has to remember to do first.
- `manager/measure.rs` — every reading, and `DiskFit`. The "nothing else
calls `statvfs` or reads process RSS" claim is now checkable by looking
at one file.
- `manager/enforce.rs` — `check_update` / `check_capacity` and the
threshold comparison.
- `manager/mod.rs` — the struct, its construction, and the config
accessors: what a reader needs to see first.
Tests move with their subject. No behaviour change: `set_config` used to
validate before delegating to the store, and now the store validates on
write, which is the same order of operations from the outside.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: count copy-on-write deletes toward the quota
Dropping a vector or a payload key does not free anything on its own:
copy-on-write rewrites the point to produce the version without that
field, so storage grows first and is only reclaimed once the optimizer
gets to it. Gating those as if they were reclaiming space let a full node
keep taking writes that make it fuller.
Deleting whole points stays exempt. That is the one operation that has to
work on a node at its limit, or there is no way back under it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: say that the quota's reported usage is per node
`GET /quotas` returns one cluster-wide config and one set of utilization
figures, which reads as though both describe the cluster. They do not:
memory and disk are node-local, so `usage` is whatever the peer that
served the request is seeing, and a peer under its limit says nothing
about the others.
Also corrects `resident_memory_percent`, which claimed to be a share of
total system memory. It is a share of the memory available to the
process, which under a cgroup is the limit rather than the host's RAM.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: treat a node over its quota as a failed replica, not a bad request
A quota rejection described the request as invalid (400) and was classified
non-transient, which is how the replica set recognises errors that every
replica would produce alike. A quota is the opposite: the input is fine and
the answer depends on which machine you ask. On the default `wait=false`
path that combination silently dropped the write — `update.rs` only
deactivates transient failures when nothing completed — leaving the replica
Active and permanently missing data its co-replicas had.
It is now `InsufficientStorage`, transient, HTTP 507 / gRPC
`ResourceExhausted`. So a node that is out of room is handled like one that
is offline:
- last active replica, or every replica over quota: nothing could take the
write, and the client is told the cluster is out of room.
- more than one replica: the full node is deactivated through the same path
a dead peer takes, and the update stands if enough replicas accepted it.
`check_capacity` already keeps recovery off that node until it has room.
The check also moves off `update_from_client`, which applied the
coordinator's own limit to the whole operation even when it held no replica
of the shards being written. Each replica set now gates its own local write
and records the refusal as a failure of this peer, so a node only ever
answers for itself.
`ResourceExhausted` is shared with rate limiting, and the reverse conversion
mapped it straight to `RateLimitExceeded` — a forwarded rejection came back
as 429. Statuses now carry a marker so the two stay distinguishable across
the wire.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: report quota pressure per node, and across the cluster
A quota is node-local, so finding out which node has hit one meant asking
each of them in turn — and nothing at all showed up in monitoring.
`/metrics` gains a `quota_exceeded` gauge for the local node. It is emitted
only while the quota is enabled: with it off the value would be a constant
0 that says nothing about the node, and an alert built on it would go quiet
rather than fire if someone disabled the quota. Telemetry's `quota` field
carries the same verdict alongside the config, since that is where the
metric is derived from.
`GET /quotas` now answers for the whole cluster. A new `GetQuotaUsage` RPC
on the internal `QdrantInternal` service returns what one peer is using,
and the handler fans it out to every known peer in parallel. Peers that do
not answer are left out rather than failing the request — the nodes that
are out of room are exactly the ones most likely to time out, and a partial
answer still names them. Outside distributed mode the field is absent
rather than a map of one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: report the quota metric per resource
`quota_exceeded` was one flag for the whole node, which does not say what
to go and fix — disk is freed by deleting or optimizing, memory by
unloading. It now carries a `resource` label:
quota_exceeded{resource="memory"} 0
quota_exceeded{resource="disk"} 1
A resource with no limit gets no series at all, for the same reason the
metric is absent while the quota is disabled: a series that can never reach
1 reads as healthy and would quietly carry an alert that cannot fire.
`QuotaManager::exceeded` returns the per-resource verdict, with `None` for
a resource this node does not cap. Telemetry reports the same breakdown,
since the metric is derived from it. The peer usage RPC keeps a single
flag — it sits next to both percentages, so it only has to answer "is this
peer refusing writes".
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: drop a no-op error conversion the linter caught
`check_global_access` already returns a `StorageError`, so mapping it
through `StorageError::from` converted the type to itself and tripped
`clippy::useless_conversion`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: hold a tripped quota until usage clears a release margin
A resource resting on its limit crosses it in both directions on the noise
between two readings, and each crossing is expensive: the node refuses a
write, its replica is deactivated, usage dips, recovery starts sending a
whole shard copy back, and the arriving data pushes it over again. The loop
sustains itself, and every lap costs a shard transfer.
A limit now trips at its configured value but only clears once usage has
fallen 5 percentage points below it, so the crossing has to be real. The
margin is floored at 1%, since a limit smaller than the margin would
otherwise be impossible to fall back under and would strand the node.
The verdict is carried on the manager rather than recomputed, which makes
it the thing reporting shows: expect `exceeded` to be set while the
utilization next to it is already back under the limit. Rejections say so
too, rather than claiming a limit that is no longer exceeded:
Disk usage is at 87% of total capacity. It reached the configured limit
of 90% and has to fall below 85% before this node takes writes again.
Changing the config clears the verdicts. New limits are a deliberate act,
and should not be held back by the margin of a limit that no longer exists.
Both resources are now evaluated on every check instead of stopping at the
first failure, so a verdict is never left behind reporting a reading that
has since been superseded.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: make the quota release margin configurable
5 points is a guess about how noisy a deployment's usage is, which is not
something one number can be right about: a node whose disk moves in
gigabyte steps needs a wider margin than one that creeps, and an operator
who wants the old flip-on-every-reading behaviour should be able to ask
for it.
`release_margin_percent` joins the rest of the quota config, so it seeds
from `QDRANT__STORAGE__QUOTAS__RELEASE_MARGIN_PERCENT`, replicates through
consensus, and changes with `PUT /quotas`. Defaults to 5 and is filled in
when a request omits it, so it always answers with the margin actually in
force rather than leaving the caller to assume one. `0` releases as soon as
usage is back under the limit.
`QuotaConfig` grows a hand-written `Default` for it, since deriving one
would have quietly defaulted the margin to 0 and disabled the hysteresis
for anyone constructing a config in code.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor: leave the release margin unset by default, and hold verdicts in atomics
`release_margin_percent` is `null` unless someone sets it, rather than
materialising 5 into every config. A quota written today then does not pin a
number a later release may want to revise, and `{"enabled": false}` still
round-trips as itself. `QuotaConfig::limits` resolves it, next to `enabled`,
so enforcement never sees the unset case.
The verdicts move from a `Mutex<QuotaExceeded>` to one `AtomicBool` per
resource. They are judged independently and nothing reads them as a pair, so
the lock only added contention to the path every update takes; a verdict
that races a concurrent check is re-decided by the next one from a fresh
reading.
That also drops the tri-state. Only "was this over its limit" has to survive
between checks — whether a resource is enforced at all follows from the
config and the reading, so it is derived when reporting rather than stored.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: drop a comment arguing with a design that was never here
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Remove deprecated search/recommend/discover endpoints from OpenAPI
Remove deprecated REST API endpoint definitions from the OpenAPI
generator. These endpoints were deprecated in v1.13.3 (`f4ced2567`,
#5907, 2025-01-30) in favor of the universal `/points/query` endpoint:
- POST /points/search
- POST /points/search/batch
- POST /points/search/groups
- POST /points/recommend
- POST /points/recommend/batch
- POST /points/recommend/groups
- POST /points/discover
- POST /points/discover/batch
Also removes the corresponding request types from the schema generator
and updates the expected API count in the consistency check.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Migrate OpenAPI integration tests to /points/query
The deprecated /points/search, /points/recommend and /points/discover
endpoints (along with their /batch and /groups variants) were removed
from the OpenAPI spec, which caused validation failures in the Python
integration test harness.
This commit migrates the affected tests to the universal /points/query
endpoint:
- Delete tests dedicated to the deprecated endpoints:
test_recommend.py, test_discover.py, test_multicollection_reco.py,
test_recommendation_multivector.py
- Refactor remaining tests to call /points/query (and /query/batch,
/query/groups), translating request bodies (vector -> query / using,
positive/negative -> query.recommend, target/context -> query.discover)
and unwrapping the new result.points response shape.
- Drop equivalence assertions against the now-removed legacy endpoints.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Relax non-empty assertions in migrated recommend/discover tests
The previous migration added `len(...) > 0` assertions to tests that
previously only checked equivalence between the deprecated and new
API. These assertions are too strict because the parametrized
`query_filter` cases legitimately produce empty result sets.
Drop the `> 0` assertion and rely on `request_with_validation` to
verify the response is well-formed and HTTP OK.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Migrate remaining OpenAPI tests off deprecated search endpoints
Tests added to dev after the original migration was written still call
/points/search and /points/recommend/groups through
`request_with_validation`, which resolves the endpoint against the
OpenAPI spec and therefore breaks once the endpoint is not in the spec:
- test_turbo4_storage.py, test_sparse_idf_corpus.py, test_validation.py:
translate /points/search to /points/query (vector{name,vector} ->
query + using, result -> result.points).
- test_group.py: drop the /points/recommend/groups half of the
lookup_from validation test in favour of the query equivalent.
test_sparse_idf_corpus.py's test_query_api_supports_idf_corpus goes
away: with the helper on /points/query every test in the file now
exercises what it asserted.
Also record why test_recommend_group cannot assert on its groups: it
uses every point in the collection as a recommend example, so all of
them are excluded and the result is legitimately empty.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Regenerate openapi.json without the deprecated search endpoints
Drops the 8 deprecated paths and the request schemas that only they
referenced: Search/Recommend/Discover request (+Batch, +Groups) types
and their exclusive dependencies (NamedVector, NamedSparseVector,
NamedVectorStruct, UsingVector, RecommendExample, ContextExamplePair).
Regenerated output is a strict subset of the previous spec, and every
remaining $ref still resolves.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Deprecate the search/recommend/discover RPCs in gRPC
The REST counterparts have carried `deprecated: true` since v1.13.3 and
are now gone from the OpenAPI spec, while the gRPC RPCs never got any
deprecation annotation at all. Mark all 8 with `option deprecated = true`
so generated clients warn, and point each doc comment at its `Query`
replacement.
tonic puts `#[deprecated]` on the generated client methods only; the
server trait gets the doc comment alone, so our own `impl` is unaffected.
The RPCs keep serving traffic — this is annotation only.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Restore the deleted recommend/discover suites on /points/query
The earlier migration deleted these four files outright, but the
query-side tests it left behind are all shallow smoke tests
(`len(result) > 0`, `"points" in result[0]`). The deleted ones carried
invariants with no query-API equivalent anywhere, so deleting them was a
real loss of coverage rather than de-duplication:
- test_recommend.py: default strategy equals average_vector; batch
results identical to sequential singles across six request shapes;
best_score with only negatives yields all-negative scores; best_score
with a single positive orders identically to a nearest query; raw
vectors as examples equal ids as examples.
- test_discover.py: context-only scores are all <= 0; target-only orders
identically to a nearest query but scores differently; with a fixed
context the integer part of the score is stable while the decimal part
moves, and vice versa with a fixed target; batch equals singles;
lookup_from by id equals by vector.
- test_multicollection_reco.py: cross-collection lookup_from, plus
wrong-vector-size, unknown-collection and unknown-vector rejections.
- test_recommendation_multivector.py: the same recommend invariants over
a max_sim multivector collection, which the query suite never covered.
Only test_recommend_missing_lookup_from_collection_with_raw_vector is
dropped as genuinely redundant — test_query.py's
test_query_missing_lookup_from_collection covers query, query/batch and
prefetch.
Two request-shape differences the translation had to absorb:
- Giving no examples at all is 422 (a RecommendInput validation rule),
where the legacy API reported 400 from the query itself. A malformed
example, such as an empty vector, is still 400.
- DiscoverInput requires the `context` key and accepts only an explicit
null to mean "no context", so target-only discover must spell it out.
The legacy API let it be omitted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Security release addressing CVE-2026-59884, CVE-2026-59885, and CVE-2026-59886.
Equivalent to #9973 for the dev branch.
Co-authored-by: Cursor <cursoragent@cursor.com>
POST /cluster/recover can return 200 while raft silently drops the
snapshot request when no leader is known yet, leaving the test stuck
on a missing collection until timeout.
Co-authored-by: Cursor <cursoragent@cursor.com>
Wait for green on write (and read after recover_read) so collection and
partial snapshots are not taken mid-indexing. Otherwise a leftover
appendable segment survives partial merge and breaks manifest equality.
Co-authored-by: Cursor <cursoragent@cursor.com>
#9013 skipped the manual replicate_shard when a transfer was already
visible, but the recovery loop can still start one between that check
and the POST. Accept 400 "already involved in transfer" as success so
the remaining wait assertions still cover recovery.
Co-authored-by: Cursor <cursoragent@cursor.com>
Right after the no_sync snapshot recovery, the recovered replica serves local
reads immediately, but a remote read to it can transiently fail for a short
window. The read path then falls back to the other replica in hash order on
just the requesting peer, so a single routing token momentarily resolves to
different replicas across peers (observed as {A, B, B}), failing the
determinism assertion.
Wait until token-routed reads are stable across all peers for every token the
test asserts on before measuring, so the transient post-recovery fallback
window is passed. Pure test-side change; routing behaviour is unchanged.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test: make test_upload_snapshot robust to shard placement balance
The final assertion in recover_from_uploaded_snapshot assumed a perfectly
balanced shard placement (peer 0 having exactly 2*n_replicas remote shards).
Shard placement across peers is not guaranteed to be balanced, so this made
the test flaky (e.g. peer 0 ended up hosting all shards locally, leaving only
3 remote replicas instead of 4).
Instead, verify the full replica layout is healthy: peer 0 observes every
replica through its local + remote shards, so assert that all replicas are
Active and every shard has exactly n_replicas copies across the cluster.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test: fetch cluster info once for shard validation
Read local and remote shards from a single /cluster response so both lists
come from the same cluster revision, per review feedback.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
A transfer source that misses the transfer abort (e.g. while partitioned
or paused) keeps its local shard wrapped in a proxy. When such a peer can
only catch up via consensus snapshot, snapshot application re-creates
payload indexes with an update operation that the stale forward proxy
forwards to a transfer target which may no longer have the shard. The
resulting precondition error fails snapshot application and stops the
consensus thread ("No target shard N found for update"), leaving the
peer unable to ever catch up.
Snapshot application now explicitly cleans up transfers that are no
longer registered in consensus: the transfer task is stopped and the
proxy is reverted via the new `ShardReplicaSet::discard_proxy_local`,
which is infallible, never contacts the remote, and forgets queued
updates (replica states in the same snapshot already reflect the
transfer outcome).
The consensus test reproduces the incident: pause the transfer source
mid-transfer, restart the other peers so the aborted transfer can only
be learned via snapshot, and verify the source recovers.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat: slice filtering condition for sliced scroll and deterministic sampling
Add a `slice` filter condition selecting points where
`stable_hash(point_id) % total == index`. The hash is SipHash-2-4 with a
zero key over canonical id bytes (8 LE bytes for numeric ids, 16 RFC 4122
bytes for UUIDs) — a frozen public contract, independent of the internal
resharding ring hash, reproducible by clients to predict membership.
For a fixed `total`, slices are disjoint and cover all points, enabling
parallel scroll streams (ES sliced-scroll style) and reproducible sampling
that composes with any other filter condition.
- REST: `{"slice": {"total": N, "index": R}}`; gRPC: `SliceCondition` in
the condition oneof (tag 8)
- Evaluated per point via id_tracker external-id lookup; no payload index
needed; cardinality estimated as `points / total` with no primary clause
- `total >= 1` enforced by NonZeroU32 at parse time, `index < total` by
validation in both REST and gRPC paths
- Hash contract locked by test vectors independently reproduced with a
reference SipHash-2-4 implementation
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* tests: minimal OpenAPI test for slice filter condition
Scrolls all slices of a fixed total over numeric + UUID ids asserting
disjointness and full coverage, checks must_not inversion, and pins the
two rejection paths (422 for index >= total, 400 for total = 0). Requests
and responses are validated against the regenerated OpenAPI spec by the
test harness.
Note: the spec cannot itself reject total = 0 client-side — the Condition
anyOf falls through to the permissive Filter schema, as with any invalid
condition — so rejection is asserted via the server response.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* Fix resharding, on queries filter shards on all shard selectors
* Add failing consensus test: search during resharding with shard keys (#9880)
Reproduces a known bug: after resharding is initialized on a custom
sharded collection with a shard key, searches (with and without the
shard key selector) fail with "does not have enough active replicas",
because the new resharding shard is included in reads before it has
an active replica.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* Exempt explicit shard id selection from resharding read filter
Explicit shard id selection is only used by internal per-shard
operations (local shard API, internal gRPC reads), including the
resharding driver reading back migrated points from the new shard.
These must reach the resharding shard before it becomes visible to
user-facing selectors, and filtering them also made per-shard reads
return silently empty results on peers lagging on hashring commits.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Explicitly set resharding filtering per match branch
---------
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>