mirror of
https://github.com/qdrant/qdrant.git
synced 2026-09-29 09:27:53 -05:00
edge-docs-diff
235
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
76a8647728 |
Follow-up tests for #10349: pinned findings on the merged code (#10661)
* test: pin persisted proxy segments follow-up findings * fix: tests * Respect `up_to` in flushing pending proxy changes * Fix unproxy divergence, centralize logic in single shared function * Don't pass locked segments we don't use * Close the post-swap window losing acknowledged proxy changes `finish_optimization` left a window between `swap_new` and the end of the function where the propagated proxy changes were durable in no reachable place. The proxies had left the holder, so a flush pass no longer saw them as unsaved work and acknowledged the WAL past what their pending changes logs persisted, while their source files were still on disk contradicting the newer state. Their `ack_pin` was only registered at the very end, and the optimized segment, the only durable home of those changes, had no version file yet, so `normalize_segment_dir` deleted it on the next load. A failure or crash in that window lost every change past the proxies' logs. Close it from both sides, reordering only: - Save the optimized segment's version file right after the flush that made the propagated changes durable, before the swap. That flush is what `SegmentBuilder::build` postponed the version file for, so the segment is loadable from the swap onwards. - Register the proxies' deferred destruction (and with it their `ack_pin`) under the same write lock that evicted them, so no flush pass can observe the holder without either the proxies or their pin. Collecting the deferred point ids moves up with the registration, as it borrows the swapped-out proxies. The deferred destruction still cannot run before the manifest is synced: `locked_proxies` holds the segments alive until the end of the function, so `try_drop_data` retries until then. * Rename the optimization test hook after the window it guards The hook no longer sits before the version save, it marks a failure anywhere in the window after the optimized segment was swapped in. Rename it and the test accordingly, and restate the test doc as the invariant that window must uphold rather than the bug it used to describe. * Pin WAL ack while creating snapshot * Patch test that was stuck * Initialize necessary feature flags in tests * Correctly propagate changes in two stages, lock updates on second stage * Use existing WAL ack pinning infrastructure * test: pin unproxy phase 2 propagation failure losing acknowledged changes * test: pin WAL ack pin at zero suppressing clock persistence * Fix propagate and unproxy data consistency error on failure * Store clocks before checking WAL ack pin * fix: wait for the flush worker when stopping it in tests * fix: linter --------- Co-authored-by: timvisee <tim@visee.me> |
||
|
|
20229f99ba |
[combined-storage] Combined storage write (#10669)
* HNSWIndex::build(): add `inline_vectors` arg Let the caller decide whether to use `inline_vectors` format. * SegmentBuilder::build: finalize GraphInline vector storage Instead of old "graph-with-vectors plus regular vector storage", keep only the graph-with-vectors storage. * Gate the GraphInline segment build behind a feature flag |
||
|
|
3eefe44572 |
Fix never-ending optimization loop with vectors:{memory:cached} (#10664)
* Add `test_optimizers_should_settle` * Fix `"memory": "cached"` endless optimization loop |
||
|
|
b2b509cf73 |
Tests for #10349: coverage gaps and pinned findings (#10445)
* shard: persist proxy pending changes on flush, stop holding back WAL ack Hook the pending changes component into the proxy segment's flush: the proxy flusher first persists the buffered operations into the pending changes log, then passes the flush along to the wrapped segment. The proxy's `persistent_version` now covers what the log durably holds on top of what the wrapped segment persisted itself. That is what lifts the WAL cap proxies imposed so far. `flush_all` compares each segment's version against its persistent version; a proxy used to report only the wrapped segment's persisted version while its own version climbed with every buffered operation, so the WAL could never be acknowledged past the point the proxy was created at, and a restart replayed all of it — potentially very expensive operations, such as an update by filter, all over again. With the buffered state durable on disk the generic rule acknowledges the full version, and a restart recovers it from the log instead. Dropping a proxy's data drops the component first, which waits for any in-flight pending changes flusher so it cannot append to the segment directory while that is being deleted. Update the proxy flush test to the new semantics, add a segment holder test asserting the acknowledged version advances past a proxied delete, and update the ack pin rationale in `finish_optimization`: the pin is still needed after the proxies leave the holder, it just snapshots a persistent version that now includes the log. * Persist wrapped segment before pending changes Prevents raising version of proxy segment too early * collection: test crash recovery through persisted proxy changes End-to-end test of the persisted pending changes: wrap every segment of a local shard in a proxy, delete points so the deletes are only buffered, flush, and assert the acknowledgeable version covers them. Then acknowledge the WAL up to that version, drop the shard without ever propagating the proxies, and load it again: the deletes are gone from the WAL and must come back through the pending changes logs. The delete under test is deliberately not the last WAL entry, as the acknowledge never passes the last entry and that one is always replayed. * test: cover untested pending proxy changes paths * test: pin persisted proxy changes findings * Fix bad merge * Remove corrupt length test, we cannot detect if last entry was corrupt Remove a test that asserts an entry that isn't the last cannot have a corrupt length. We cannot reliably detect whether the invalid length was the last entry or not, because nothing else tells us how many entries we expect in the file. At the same time we don't expect random bit flips. So I removed the test. * Simulate segments flush to clear pending changes log file * Update test, also assert proxy segment version * Fix bad merge * Fix blocked test * Enable necessary feature flags in tests (2/2) --------- Co-authored-by: timvisee <tim@visee.me> |
||
|
|
cc7a209c76 |
Persist proxy segment changes across restart, don't stall WAL ack's (#10349)
* segment: move proxy pending change types into segment crate Move the types describing the changes a proxy segment buffers — point deletes (`ProxyDeletedPoint`), payload index changes (`ProxyIndexChange`, `ProxyIndexChanges`) and vector name changes (`IntendedVector`, `ProxyVectorNameChanges`) — from `shard::proxy_segment` into a new `segment::pending_changes` module. Pure move, no behavior change: the proxy segment re-exports them from their old location. Having them in the segment crate lets both the proxy segment and the segment load path share them, in preparation for persisting pending proxy changes to disk and replaying them on restart. * segment: add PendingChange describing a persisted proxy operation Add the `PendingChange` enum with one variant per operation type a proxy segment buffers — point delete, payload index change, vector name change — each carrying the operation version it was issued with. This is the shape in which pending proxy changes are persisted to disk. Derive serde on it and on the buffered change types it embeds, so entries can be serialized into a log file and read back. `PartialEq` on those types lets a persisted batch be matched against the in-memory pending buffer after a flush. * segment: add PendingChanges component persisting proxy changes to a log Add `PendingChanges`, the component that manages the operations a proxy segment buffers for one proxy layer, and persists them to disk so they no longer only live in memory. It keeps the same per-type buffers the proxy segment served its reads from (point deletes, payload index changes, vector name changes), plus a single registration-ordered buffer of everything not yet persisted. `flusher()` writes that buffer into an append-only log file inside the wrapped segment's directory: `pending_changes.log` for the inner most proxy layer, with the layer number as a suffix for each layer above it. Appends follow the mutable ID tracker: all new entries are serialized into one buffer and written with a single call on an append-mode file, then fsynced, so a crash can only leave a torn entry at the very end. Loading truncates such an entry — its operations were never durable and thus never acknowledged in the WAL — but fails hard on a malformed entry in the middle, which cannot be explained by a torn append. The component tracks the highest operation version the log covers. Every registered operation at or below it is either durable in the log or was a no-op that does not need recovery; a flusher advances it to the proxy's version even when there is nothing to write. The pending buffer is deliberately not cleared when the proxy propagates its changes to the wrapped segment, as that only makes them durable once the wrapped segment flushes. Replaying an entry twice is a version-gated no-op. A log file left behind by a previous proxy on the same segment is adopted by `open()`: new entries are appended after it and its highest version is taken over, while its entries are not loaded into the buffers as they are already applied to the segment. `load()` also reconstructs the buffers, for callers that do want the buffered state. * segment: replay persisted pending proxy changes onto a segment on load Add `recover_pending_changes`, to be called when a segment is loaded on restart, before regular WAL replay. If the segment directory holds pending changes log files, the proxies that wrote them did not propagate their buffered state into the segment before the process stopped. Instead of reconstructing the proxies, replay all logged operations directly onto the segment: inner most proxy layer first, each file in append order, through the regular version-gated segment operations (`apply_change`). Entries the segment already applied are silently skipped, so a stale file is harmless. The segment is force-flushed before the files are removed; a crash in between merely replays the files once more. * segment: test PendingChanges component Cover the pending changes component: registering and flushing each operation type and reconstructing the buffers from the log, log file naming per proxy layer and gap-tolerant listing, covering the proxy version without entries, operations registered while a flusher is captured, flushers of a dropped component, torn-tail truncation versus mid-file corruption, adoption of an existing log, and replaying logs onto a real segment: fresh, stale (already applied), multi-layer, and vector name changes. * segment: include pending changes logs in segment snapshots Register the pending changes log files of a segment in its snapshot: add them to `snapshot_files` next to the segment state and version files, existence-guarded, and to the segment manifest as unversioned files. Full, partial and streamed snapshots therefore all carry them. The recovery side needs no changes: a restored segment is loaded like any other, which replays and removes the logs. * shard: back proxy segment pending changes by PendingChanges component Replace the proxy segment's separate `deleted_points`, `changed_indexes` and `changed_vector_names` fields with a single `PendingChanges` component. Reads keep going through the same per-type buffers, now behind accessors; writes go through the component's `register_*` methods, which additionally queue every operation for persistence. Opening the component is fallible, as it adopts a pending changes log a previous proxy may have left in the wrapped segment's directory, so `UnsyncedProxySegment::new` now returns a result. Wrapping another proxy opens the next proxy layer up, writing to its own dedicated log file. No behavior change yet: the proxy still flushes and reports persistence exactly as before, nothing is written to the log. * shard: persist proxy pending changes on flush, stop holding back WAL ack Hook the pending changes component into the proxy segment's flush: the proxy flusher first persists the buffered operations into the pending changes log, then passes the flush along to the wrapped segment. The proxy's `persistent_version` now covers what the log durably holds on top of what the wrapped segment persisted itself. That is what lifts the WAL cap proxies imposed so far. `flush_all` compares each segment's version against its persistent version; a proxy used to report only the wrapped segment's persisted version while its own version climbed with every buffered operation, so the WAL could never be acknowledged past the point the proxy was created at, and a restart replayed all of it — potentially very expensive operations, such as an update by filter, all over again. With the buffered state durable on disk the generic rule acknowledges the full version, and a restart recovers it from the log instead. Dropping a proxy's data drops the component first, which waits for any in-flight pending changes flusher so it cannot append to the segment directory while that is being deleted. Update the proxy flush test to the new semantics, add a segment holder test asserting the acknowledged version advances past a proxied delete, and update the ack pin rationale in `finish_optimization`: the pin is still needed after the proxies leave the holder, it just snapshots a persistent version that now includes the log. * shard: propagate proxy changes when unwrapping on optimizer cancel When an optimization is cancelled or fails, `unwrap_proxy` puts the wrapped segments back into the segment holder. Propagate the changes buffered in each proxy into its wrapped segment first, as the snapshot unproxy path already does, instead of dropping them with the proxy. The pending changes log is deliberately left in place when unwrapping: deleting it before the wrapped segment has flushed the propagated changes would not be crash safe. It is cleaned up on restart and when the segment directory is dropped, and a new proxy on the same segment adopts and appends to it; replaying a stale file is safe because all operations are version gated. * shard: test persisted proxy pending changes Test the proxy segment against its persisted pending changes: buffered changes survive dropping the proxy without propagation and are replayed onto the segment when it is loaded again; unwrapping leaves the log in place and a new proxy on the same segment adopts and appends to it; layered proxies each persist into their own log file and a restart replays both; and a persisted log is part of the segment manifest and snapshot. * collection, edge: recover persisted proxy changes on segment load Replay the pending changes logs left behind by proxy segments onto each segment when a shard loads its segments, right after consistency repair and before the payload index rebuild, vector name reconciliation and WAL replay. Proxy state that made it to disk no longer holds back the WAL acknowledge, so this is where it must be recovered from. Proxies are not reconstructed: the segment holder starts with plain segments carrying the replayed operations, and the logs are removed once the segment flushed them. * collection: test crash recovery through persisted proxy changes End-to-end test of the persisted pending changes: wrap every segment of a local shard in a proxy, delete points so the deletes are only buffered, flush, and assert the acknowledgeable version covers them. Then acknowledge the WAL up to that version, drop the shard without ever propagating the proxies, and load it again: the deletes are gone from the WAL and must come back through the pending changes logs. The delete under test is deliberately not the last WAL entry, as the acknowledge never passes the last entry and that one is always replayed. * segment: make replaying persisted proxy changes on load an explicit mode Add `PersistedProxyChanges` to state whether persisted pending proxy changes are replayed onto a segment when it is loaded. `Replay`, the default, recovers them and removes the logs as before. `Ignore` leaves both the segment and the log files untouched and logs at debug level that replaying was skipped; it is for segment files that mirror those of another writer, where replaying would make the local copy diverge from what the writer's manifest describes. All callers pass `Replay` for now, no behavior change. * collection: do not replay persisted proxy changes on partial snapshot recovery Partial snapshots are recovered by read replicas in a read/write segregation setup. A read replica must not mutate its segments, so it cannot replay the persisted proxy segment changes on load and must ignore them instead: its segment files are a local copy of the writer's that must stay a faithful mirror of them, as later partial snapshots are diffed against what the writer's manifest describes. Replaying would mutate the segment files and remove the logs, making the copy diverge. Thread the replay mode through `LocalShard::load` as a dedicated `PersistedProxyChanges` argument, derived from the recovery type: `RecoveryType::Full` replays as before, `RecoveryType::Partial` ignores the persisted changes and leaves the logs in place. Regular shard loads replay. Extend the crash recovery test with an ignoring load first: the delete under test must not come back and the logs must survive, before a replaying load recovers it. * Persist wrapped segment before pending changes Prevents raising version of proxy segment too early * Fix comment * Fix crash window, only ready optimized segment after propagating changes The optimizer renamed a newly built segment into segments_path and wrote its version file before finish_optimization propagated the proxies' buffered changes into it. A crash in that window left the segment restart-loadable but stale, permanently losing or resurrecting points. Defer the version file save until finish_optimization has fully reconciled proxy changes into the segment, including the post-swap dedup pass, so it stays invisible to restart and snapshot recovery until then. SegmentBuilder::build() gains a `ready` flag; load_segment gains `ignore_missing_version` for the one caller reloading before that point. Incidentally also closes the crash-unsafe cancellation-orphan cleanup gap noted in #9217, since a cancelled build is discarded on restart the same way. * Force flush optimized segment, otherwise we may lose proxy changes * Don't force flush after replay, defer deleting log files until flush * Include persisted proxy changes log file in segment manifest * Add random ID to proxy log files, prevent instance conflicts * Rename proxy log file, always include level * Delete proxy log file on unproxy, defer until next flush cycle * Fix truncation * Reformat * Lock persisted segments behind runtime feature flag * Enable necessary feature flags in tests * Fix linters |
||
|
|
0d2da625c4 |
Add cancellable reads for read-only edge shards (#10646)
* [AI] Add cancellable read trait for read-only edge shards * Move edge read cancellation tests into separate module |
||
|
|
81bb80a5d3 |
Expose id tracker memory placement in collection config (#10597)
Add `id_tracker: { memory: cold | pinned }` to CollectionParams,
CollectionParamsDiff and CreateCollection (REST + gRPC `IdTrackerParams`),
mirroring `payload: { memory }`. `cold` builds the disk-resident id tracker,
`pinned` the in-RAM immutable one. Unset keeps the current behavior: the
`serverless_compatible` feature flag decides.
The requested placement is persisted as an optional `id_tracker_memory` on
SegmentConfig (skipped when unset, so existing configs are unchanged); the
segment builder resolves it through `SegmentConfig::id_tracker_memory_placement`
instead of reading the feature flag directly.
The config mismatch optimizer rebuilds non-appendable segments whose effective
placement differs from the requested one. Appendable segments are skipped: they
always use the mutable tracker and get the current config when indexed.
`cached` is rejected by validation: the disk mapping reader has no
populate-on-open path.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
||
|
|
93e91e1da4 |
Do not claim an unfinished operation when flushing (#10577)
* fix: do not claim an unfinished operation when flushing A flush pass can capture a segment between the separately locked steps of one update operation. Persisting it under that operation's version marks the segment clean while the rest is still in memory, so every later pass skips it and the WAL acknowledge moves past the operation. Clamp what a flush claims to the last fully applied operation. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Fix rustfmt in alias_mapping test after merging dev --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com> |
||
|
|
d4f09f5d89 |
Fix MMR pagination with offsets (rebase of #10502 onto dev) (#10567)
* Fix MMR pagination with offsets (cherry picked from commit |
||
|
|
1e56bbfabb |
Stamp SegmentManifestState::Retiring with the swap time (#10576)
A segment retired by the serverless indexer's swap is still held by every read-only follower that has not live_reloaded past the swap yet. Deleting it "on a later mark" gave those readers no bounded grace period: on a busy shard the next mark follows the swap within seconds. `Retiring` now carries `retired_at` (unix seconds of the swap), so the GC can require a minimum age before it deletes the directory. It serializes as a tagged object like `Optimizing`; the bare `"retiring"` string is no longer accepted. Only the serverless indexer writes this state, and the manifest is behind a feature flag, so no deployed data carries the old form. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
51cf88e7de |
[combined-storage] Integrate VectorStorageType::GraphInline reading (#10515)
* Placement accessors on VectorDataConfig * Wire up the GraphInline storage type * Let the HNSW index reuse the storage's links handle |
||
|
|
a6c9891538 | Add iteration count and time in trying to read lock log line (#10542) | ||
|
|
b2d67cffef |
Fix proxy propagate vector name versions optimizer (#10531)
* Add test, apply all changes after optimization * Artificially bump operation version after optimization propagation |
||
|
|
16cec09de5 |
Fix proxy segment not propagating some named vector changes (#10507)
* Move version reading into debug assertion, it's cheap anyway * Add test, named vector change before index change must still land * Fix named vector change not being applied, raise to segment version * Describe artificially raising version centrally in more detail |
||
|
|
73831260e6 |
Remove msgpack (#10505)
* Remove MessagePack (rmp-serde) The WAL switched from msgpack to CBOR in v0.3.5 (2021-07-11), so v0.3.4 is the last version that wrote msgpack entries. Drop the read fallback kept for those entries, plus the remaining test and bench usages. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Drop unused fs4 dependency from collection Not referenced anywhere in the crate. Still used by wal and common, so the workspace entry stays. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9687da6c41 |
Use more direct calls (#10347)
* Direct call for shard transfer method and keys * Reuse cardinality estimate in sparse plain search * Avoid recounting available points in segment size info * Avoid cloning segment config when updating quantization * Avoid cloning search request for load profile * Direct call for counting read-only segments * Avoid re-reading point range for values count * Direct call to check replica states when initializing collection * Direct call to look up transfer on restart * Direct call for shard replicas after snapshot recovery * Direct call for local replica states in health check * Direct call for payload index schema keys when applying state * Direct calls for sharding method and key mapping when creating shard key * Direct call to check if peer has shards * Direct call for sharding method and keys when dropping shard key * Avoid cloning collection params for group by ordering * Avoid cloning collection params in local shard search * Direct call for peer address when sending Raft messages * Direct call for peer address in who_is * Avoid cloning remote query batch request * Avoid cloning operation in queue proxy update * Avoid cloning gRPC search groups request * Fetch cluster status once in cluster telemetry * Direct call to validate transfer exists on finish * Direct call for sharding method when dropping shard key * Avoid cloning peer address map when listing peers * Avoid cloning peer address map when adding peer to known * Avoid cloning shard key mapping when routing writes with fallback * Avoid cloning shard key mapping when checking resharding start * Avoid cloning gRPC recommend groups request * Avoid cloning operation when retaining forwarded point IDs * Direct call for counting collections in telemetry * Direct call to validate transfer exists on recovery * Direct call for shard IDs by shard key * Direct call for shard keys * Direct call to check if peer has shards in consensus * Direct call for replica state on transfer recovery * Direct call to check for active replicas when routing writes with fallback * Direct call to validate transfer exists on abort |
||
|
|
47e858c94e |
Remove unused (Sparse)VectorStorageType::Empty (#10431)
These were added for named-vector CRUD (
|
||
|
|
f999bc93eb |
[combined-storage] Derive inline-storage warnings from the optimizer's vector config (#10430)
* Refactor: Untangle SegmentOptimizerConfig * Derive inline-storage warnings from the optimizer's vector config |
||
|
|
1e095ed470 |
fix: reject mismatched dense dims in recommend average (#10374)
* fix: reject mismatched dense dims in recommend average Stop silently truncating oversized negative examples during average_vector merge. Validate dense dimensions within each example group and between positive/negative averages before zip-merge. Fixes #10369 * Simplify: keep only the merge-time dimension check The zip truncation in merge_positive_and_negative_avg is the only place an oversized negative can silently pass the downstream dimension check; within-group mismatches already grow the average to the max length and fail the segment-entry check. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * style(query): make recommendation conversion explicit * test: assert recommendation dimension errors Issue: #10369 Make the regression test verify the exact WrongVectorDimension payload for mismatched recommendation vectors. --------- Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com> Co-authored-by: generall <andrey@vasnetsov.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3e82341980 |
Update-only writer: leave optimizing targets alone and create fresh appendable segments (#10416)
* feat: create appendable segments when the write target is optimizing or the shard is empty * review: SegmentManifestState::is_writable, caller-supplied temp dir, uuid from token - `SegmentManifestState::is_writable` with a full match replaces the ad-hoc `matches!` in the manifest enumerator. - `ListedSegment` is destructured in `open` so every field is accounted for. - `create_appendable_from` is test-only; `create_appendable` is the API. - `create_appendable` builds the scratch segment in a caller-supplied local `temp_path` (conventionally `<shard>/temp_segments`) instead of the system temp dir, and takes the uuid from the build token instead of parsing the path. Upload speed of `copy_dir_via` is tracked in #10433. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: generall <andrey@vasnetsov.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
ae0b075c74 |
Skip redundant ID resolutions (#10333)
* perf: skip retrieval in scroll when no payload or vectors are requested Every scroll variant went through SegmentsSearcher::retrieve to build its records, even when neither payload nor vectors were asked for. That is a has_point lookup per id per segment plus a version and id resolution per hit, only to yield records holding nothing but the id. The universal query API always scrolls this way and fetches payload separately afterwards. Build the bare records from the ids directly in that case. The retrieve could only have dropped ids deleted in between, which the update lock held across the scroll rules out. * perf: fetch payload and vectors in the leaf of plain query requests A query without prefetches and without rescoring is served by a single leaf search or scroll whose result is returned as is. The planner still built that leaf without payload or vectors and filled them in afterwards through SegmentsSearcher::retrieve, which resolves every result id in every segment again: the same cost #10312 removed from the search API, paid once more at the end of each query. Let the leaf carry the requested payload and vectors instead, so the segment attaches them to the results it already holds by offset, and clear the root plan so the fill step is skipped. Prefetch leaves and rescored roots (MMR) are unchanged. As with the search API, this fetches payload for each segment's candidates rather than for the merged top `limit` alone. * Use new_empty function * fix: fetch payload and vectors in scroll leaves only (#10384) A search leaf hydrates every segment's local top-k before merging, so `with_payload` there multiplies payload I/O by the segment count — the regression #6279 fixed and `test_payload_io_read_is_within_limit[query]` guards. Scroll leaves retrieve once for the merged page, so they keep fetching directly; search leaves stay bare and the root plan retrieves for the final result. Claude-Session: https://claude.ai/code/session_01SUWh5PqUeSefrUqUwXxU3E Co-authored-by: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b1a1c00059 |
Fix proxied changes dropped when an optimization fails (#10364)
* Add SetFlushInterval op to the model tester Changes the collection's flush_interval_sec mid-run through the same path update_collection takes (persist the optimizer-config diff, then recreate the optimizers in the background). The model is untouched: what it perturbs is the flush cadence, so how much of the workload is still WAL-only when a restart hits, plus the worker stop/start race in on_optimizer_config_update. Kept in FORCE_OFF for now: with the optimizer on it makes stale point state visible within a few ops of the config change. Narrowed to recreate_optimizers_background, see the comment on Swarm::BASE. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011AjmS5GFeGztfP3JqutnXj * Keep SetFlushInterval enabled in the swarm Drops it from FORCE_OFF so the divergence it surfaces is reachable without --enable-force-off (which would also enable the broken vector-name ops). The evidence moves from the FORCE_OFF comment onto the op's own doc. The two optimizer-on harness gates now fail whenever the swarm draws the op. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011AjmS5GFeGztfP3JqutnXj * Propagate proxied changes when unwrapping proxies on optimization failure unwrap_proxy puts the wrapped segments back into the segment holder, so the changes recorded on the proxy while the optimization ran (deleted points, index and vector-name changes) have to reach the wrapped segment first. They did not, so every point deleted or overwritten during the optimization kept its pre-optimization copy live next to the new copy in the write segment, and reads saw both: counts too high, scroll and search returning the stale copy. The snapshot unproxy path already does this; the optimizer failure path was the only place putting a wrapped segment back without it. It is reachable whenever the shard outlives the cancellation, in particular an update_collection that recreates the optimizers while an optimization is in flight. Lock order is holder-then-updates, matching try_unproxy_segment: updates-then- holder-write deadlocks against the snapshot path. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011AjmS5GFeGztfP3JqutnXj * Drop the stale failure note from the SetFlushInterval doc The divergence it described is fixed in this branch. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011AjmS5GFeGztfP3JqutnXj * test as well with 0s as flushing interval * Fail optimization unwrapping when proxy propagation fails Losing proxied deletes and index changes is data corruption, so return the error instead of logging it: no proxy is unwrapped and the changes stay served by the proxies. The cancelled-segment cleanup moves ahead of unwrap_proxy so the orphan is still removed when that error fires. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CpoaAtGbAQAEuxEi8ScHc5 * Drop the model tester --flush-interval-sec flag SetFlushInterval covers the interval now, so the run starts at the shipped 5s default (fixture::INITIAL_FLUSH_INTERVAL_SEC, still traced in the header) and the ops move it from there. Also documents what 0 does now that it is a generated value. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CpoaAtGbAQAEuxEi8ScHc5 * Fail snapshot unproxying when proxy propagation fails Both paths logged the error and unwrapped anyway, dropping the deletes and index changes that never reached the wrapped segment. Same reasoning as unwrap_proxy in the optimizer. try_unproxy_segment hands the lock back and leaves the proxy installed, the failure mode its doc already describes: the caller keeps it in `proxies` and unproxy_all_segments retries the propagation right after. unproxy_all_segments returns before touching the holder, so the temp segment the surviving proxies write into stays in place (remove_segment_if_not_needed only checks whether it is empty and appendable, not whether a proxy still references it). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CpoaAtGbAQAEuxEi8ScHc5 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e40b3bd799 |
Stop growing appendable segments past max_segment_size (#10027)
* Steer writes away from appendable segments at max_segment_size * pick a write target that stays under the configured size cap, instead of growing an appendable segment past it * clamp the deferred points threshold to max_segment_size, treating a zero cap as uncapped * apply the same cap when replaying the WAL, so recovery matches live updates * plumb the cap through the update worker and cover the silent-failure gaps Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Extract the per-segment capacity check into a helper Lets has_appendable_segment_with_capacity short-circuit on the first segment below the cap instead of collecting every eligible ID. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
83fe47c90a |
feat: add a dedicated min operator to score formulas (#10296)
Follow-up to #10287, which added `max`. Expressing a minimum still required spelling out `(a + b - |a - b|) / 2`, the sign flip of the max identity — drop the `neg` and you silently get a maximum instead. It also only works for two operands and mentions each one twice, so the scorer walks every sub-tree twice per candidate point. The pair is what makes clamping expressible: {"max": [0.0, {"min": [1.0, "$score"]}]} `min` mirrors `max` throughout, and both guard helpers introduced in #10287 already took an `operator: &str`, so they are reused unchanged: an empty operand list is rejected at parse time rather than folding to +infinity, and the Edge FFI rejects it at construction time. The result needs no `is_finite` check, since `min` cannot produce a non-finite value from finite inputs. The unindexed-field walker shares one arm for `Max | Min` as the bodies are identical, with a test pinning `min` separately so a later split cannot silently drop it. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b88becd3b5 |
docs: parameter names in doc comments that the signatures do not have (#10290)
11 names across 7 files. Renames that did not reach the comment above them (further_searches for further_results, query_context for segment_query_context, block_ranges for local_block_ranges, op for operation twice, request for requests, max_threads for max_kmeans_threads), and 3 arguments that were removed from a signature and left documented (is_on_disk, collection_params, search_runtime_handle with timeout). Documentation only, no behaviour change. |
||
|
|
e2d42462fa |
feat: add a dedicated max operator to score formulas (#10287)
* feat: add a dedicated max operator to score formulas
Expressing a maximum in a score formula required spelling out the
arithmetic identity `(a + b + |a - b|) / 2`. That is easy to get wrong
(the `/ 2` is load-bearing), only works for two operands, and mentions
each operand twice, so the scorer evaluates every sub-tree twice per
candidate point.
`max` is variadic, mirroring `sum` and `mult`:
{"max": ["$score", {"mult": [0.5, "popularity"]}]}
Unlike `sum` and `mult`, `max` has no identity element for the empty
case, so an empty operand list is rejected at parse time rather than
folding to -infinity and scoring every point with a non-finite value.
The check lives in `ExpressionInternal::parse_and_convert`, which every
entry point passes through, and the Edge FFI additionally rejects it at
construction time to match how that crate validates elsewhere.
The result needs no `is_finite` check: unlike `log10`, `exp`, `div`,
`sqrt` and `pow`, `max` cannot produce a non-finite value from finite
inputs, so it follows the existing `sum` convention.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: cover max error propagation and datetime operands
An operand that fails must fail the whole expression rather than being
passed over in favour of a finite sibling. Covered with the failure both
before and after the finite operand: `mult` short-circuits on zero and
so can skip evaluating later operands, and this pins down that `max`
must not grow a similar shortcut that would swallow an error.
Also covers `max` over datetime operands, which reach the scorer through
a separate conversion to seconds, so that "score by whichever timestamp
is newer" is verified rather than assumed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
ba76dd6188 | Remove stale comment (#10279) | ||
|
|
f3ffc65531 |
fix(shard): flush CoW destinations before the payload-index pre-build flush (#10201)
* fix(shard): flush CoW destinations before the payload-index pre-build flush create_field_index force-flushes each segment before building an index on it (flush-before-build, #9767), one segment at a time, outside flush_all's all-segment lock capture and copy-on-write dependency ordering. That flush durably advances a CoW source past the delete halves of its pending moves. The appendable-first iteration order usually flushes the destination before the source, but not always: a destination proxy-wrapped by a running optimization is classified non-appendable and can skip its flush entirely through the already_indexed short-circuit (the proxy reports the field as present), and a move landing mid-pass is ordered behind nothing. Once the source flushes, the move's WAL entry stops being replayable: the pre-image is durably deleted while the only current copy sits in the unflushed destination, and a graceful close then loses the point. This is the root cause of the nightly model-testing reload divergence (#10095), traced end-to-end in CI runs 31583878492 and 31583871346: cow move op 5197 into a freshly proxied destination, index op ~5252 flushing every source past it while skipping the proxy, destination reloading at 5181, replay declining with 'No point with id'. The fix mirrors flush_all's invariant at the only per-segment flush site: before flushing a segment, flush the destinations of its pending flush_dependency edges (one hop suffices, destinations are appendable and never CoW sources). Destination guards are taken before the flush lock to keep the documented [segment locks -> flush lock] ordering. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(shard): regression test for the CoW-destination flush in create_field_index Reproduces the #10095 loss shape deterministically: a pending copy-on-write move out of a non-appendable source, a destination whose own pre-build flush is skipped by the already_indexed short-circuit, then a holder-wide create_field_index. Verified failing with the dependency-aware flush neutralized (destination stays behind the move while the source flushes past it) and passing with it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor(shard): move the CoW-aware single-segment flush into SegmentHolder Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
153d92e215 |
Fix quota disk-limit test on near-empty tmpfs (#10256)
limits_only_apply_while_the_quota_is_enabled assumed the system temp dir is on a filesystem at least 1% full, the smallest configurable disk limit. On a tmpfs /tmp the usage floors to 0% and the limit never trips. Measure the tempdir's filesystem first and fall back to a tempdir in the crate directory when it is under 1% used. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e32d3fbf89 |
Add acosh expression to formula query (#10231)
Unary inverse hyperbolic cosine, parallel to sqrt/ln/exp/log10, in REST, gRPC, and edge (FFI + Python) interfaces. Inputs below 1 produce the same NonFiniteNumber error as an invalid sqrt or ln. Closes #10186 Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
86b9330628 |
transfer: send raw payloads, behind feature flags (#10066)
A raw point can carry its payload as the byte blob it is stored as, mirroring `PointStructRaw.raw_payload` on the internal gRPC API. The blob travels from the sending node into the receiving node's WAL untouched, so the sender never parses the payload it read and neither node builds a protobuf value tree for it. It is parsed exactly once, where the operation is unpacked for apply (`process_point_operation`), because that is the first place the parsed form is actually needed: `set_full_payload` goes through the payload index, which cannot be updated from bytes. The gRPC boundary therefore only checks the encoding tag and rejects a point that sets both payload fields, the way the enclosing request already rejects both `points` and `raw_points`. Moving the parse onto the apply path makes its error classification load-bearing, so a malformed blob is reported as `OperationError::MalformedPayloadBlob` — the payload sibling of `MalformedVectorBlob`, mapped to `CollectionError::BadInput` for the same reason: a bad blob that reached the WAL has to be skipped on replay instead of crash-looping recovery. Three consequences of the blob living that long are handled explicitly rather than by convention: - `decode_payload_raw` takes the blob only once it has parsed, so a failure leaves the point holding it instead of holding neither representation. - `upsert_points_raw` and `sync_points_raw` refuse a point that still carries a blob. They read the parsed payload, so such a point would otherwise be stored with no payload at all, and a `debug_assert!` would not catch it in release. - `is_equal_to` compares blob to stored blob as bytes. A differing encoding costs a redundant upsert on sync, never a skipped one. The `raw_payload_transfer` bench measures the trade, per 100-point batch (one transfer batch) at payloads of ~200 B / ~700 B / ~7 KB: - Sender, storage bytes to wire: 16x / 37x / 113x faster. This is where the whole win is — no parse of the blob that was read, no value tree built. - WAL encode: 5x / 11x / 25x faster, writing a byte string instead of a map. - Receiver, wire to applicable point: 1.09x / 1.10x / 1.06x. Near neutral, as it swaps walking a prost value tree for a JSON parse. - Wire bytes: ~6% smaller. WAL bytes: 10-32% *larger*, because the blob is JSON while a parsed payload is written as a compact CBOR map. The WAL growth is accepted rather than fixed: decoding earlier to win those bytes back costs a second full deserialization, and would leave the receiving side with a `payload_raw` that is never populated. Making the blob itself compact belongs in the payload storage encoding (`RawPayloadEncoding` is the extension point for it), not here. Two flags, both off by default and both sender-only (nodes accept raw points and raw payloads regardless), read where the transfer batch is prepared: - `transfer_raw_points` transfers every collection as raw points, not only those whose vector storage would drift in a decode-encode round-trip. - `transfer_raw_payloads` ships the blob a raw read hands out; without it the prepared batch decodes it back into the parsed payload, and the wire message is exactly what it is today. Neither is enabled by `all`: a node only accepts them once it runs a version that understands them, so they can only be switched on a release later. Nothing enforces that yet — the transfer has no peer-version gate. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6cc6980309 |
fix: do not panic when the global quota manager is installed twice (#10185)
`TableOfContent::new` installs the process-global quota manager, and `set_global` treated a second install as a startup-order bug worth a `debug_assert!`. A test binary runs all of its tests in a single process, and every test that builds a table of contents installs the manager again, so all but the first one panic. Building several tables of contents in one process is legitimate for test harnesses, so the second install is no longer fatal. A node that installs twice still reports it loudly through the existing error log. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
03a09ef51f |
fix: route append-only point deletion through delete_point_internal (#10199)
Alternative to #10188. Instead of skipping check_consistency_and_repair's storage cleanup entirely on append-only segments, and keeping a separate delete_point_tombstone_only helper that callers must remember to pick, fold the append-only decision into delete_point_internal itself: - delete_point_internal now branches on is_append_only_delete() internally: tombstone-only (id tracker drop only) when true, full payload/field-index clear + id-tracker drop otherwise. - delete_point_tombstone_only is removed; both call sites (ordinary delete_point, and check_consistency_and_repair's cleanup of dangling versions found by fix_id_tracker_inconsistencies) now just call delete_point_internal, so the repair path gets correct append-only behavior for free instead of needing its own explicit check. - version_tracker.set_payload (payload-storage version, used for partial snapshots) moves inside delete_point_internal too, gated behind the same branch: it's only bumped when payload storage is actually touched. Took the opportunity to also thread it through an explicit op_num: Option<..> parameter, since check_consistency_and_repair's repair pass has no real op_num to associate the change with. This keeps "how to delete a point" a property of segment state rather than something every caller has to branch on externally, which is what let the original bug slip through check_consistency_and_repair in the first place. Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
6633205704 |
[UpdateOnly] create Blobstore-backed storages append-only under a feature flag (#10154)
* [UpdateOnly] create Blobstore-backed storages append-only under a feature flag A new `append_only_storages` feature flag, enabled by `serverless_compatible`, switches every Blobstore creation site — the payload storage, the appendable field indexes (numeric, map, geo, full-text) and the sparse vector storage — to the append-only Logstore mode. One shared helper maps each site's Gridstore layout to its Logstore counterpart, carrying the page size and compression over; blocks and regions have no append-only equivalent. Only creation consults the flag: an existing storage keeps its persisted mode, both modes are always readable, so flipping the flag never strands data. Two changes make the flag usable rather than booby-trapped: `Logstore::delete_value` now succeeds trivially where nothing is stored, as mutable mode does, and errors only for a stored value. The ordinary write paths delete defensively — an index clears a slot before filling it, an empty value is stored as a deletion — and only ever hit occupied slots when something is genuinely mutated in place. A segment derives `append_only_storages` from the persisted payload storage mode when it opens — not from the flags, which may have changed since it was created — and it forces append-only mutation semantics on itself: every mutation clones to a fresh slot, and the same-operation slot-reuse shortcut is disabled, since the second step of a multi-step write would rewrite a payload row those storages cannot rewrite. The end-to-end test runs as its own binary (feature flags are process-global) with `serverless_compatible` on: the segment comes out holding Logstore storages, and upserts, updates of existing points, multi-step same-operation writes, deletes, an index build over existing points, a flush and a reload all run against them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * [UpdateOnly] assume the flag pairing instead of deriving it, trim the docs Per review: `append_only_storages` without `append_only_mutations` is not a state to defend against — `init_feature_flags` forces the pairing, and the same-operation slot-reuse check reads the flag directly. That deletes the segment-side derivation: the `append_only_storages` segment field, the persisted-mode read at open, and the `is_append_only` accessor chain through `Blobstore`, `PayloadStorageImpl` and `PayloadStorageEnum`. Docstrings and comments trimmed to the guarantees; how the write paths use them is their own business. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * [UpdateOnly] restructure the creation config around mode-neutral options `CreateOptions` in the blobstore crate holds what a caller actually decides — page size, block size, compression — and `into_config(append_only)` turns them into the config of either mode, each taking the fields it can express. The segment-side `storage_config` supplies only the mode, from the feature flag. That removes the misnomer chain the previous cut left behind: nothing named gridstore returns a config that might not be one, and no call site builds a `GridstoreConfig` just to have its fields repacked. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * [UpdateOnly] keep Logstore strict; fix the callers that deleted nothing Per review, `Logstore::delete_value` goes back to an unconditional error: a delete reaching an append-only storage is a caller bug to fix, not a case to absorb. The callers that issued vacuous deletes are fixed instead: - The numeric and geo indexes only delete from the storage when their in-memory index actually held values at the slot — the two are written in lockstep, so an empty slot has nothing stored either. The map and text indexes already worked this way. - The sparse storage skips the delete for keys at or past its end, where nothing was ever stored. Each removed call was wasted work in mutable mode too. The e2e test now also drives a numeric index and the sparse storage against append-only mode, and asserts that deleting a stored sparse vector fails. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * [UpdateOnly] drop the same-op slot reuse; upserts write the whole point at once The append-only path never needs a multi-step point write: the one real multi-stepper was the shard's upsert — `upsert_point` followed by a payload step under one operation number — and it now goes through `upsert_moved_point`, which writes vectors and payload as one operation and one slot. With that, the same-operation slot-reuse carve-out in `handle_point_mutate` has nothing to carry: on an append-only segment every mutating step clones to a fresh slot, unconditionally, and the `append_only_storages` special case disappears with it. The version gate skips only on strictly newer versions, so a caller that still multi-steps stays correct — it pays a slot per step. `PointToUpsert` now exposes the point's parts — raw vectors, decoded vectors, payload — and both write paths are provided from them: `upsert_into` hands the parts to `upsert_moved_point`, `write_moved` adapts them to the copy-on-write move callback. The two hand-written `upsert_into` bodies and the follow-up payload helper are gone. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Regenerate OpenAPI for the `append_only_storages` feature flag Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * one extra debug assertion * fmt --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f4ad4f4c25 |
chore(deps): drop dead dependencies in edge-path crates (#10109)
- blobstore: move `dataset` to dev-dependencies (test/bench only) - shard: remove unused `fs4` - segment: move `tap` to dev-dependencies (test/bench only) - sparse: move `tempfile` to dev-dependencies (test only) Removes the `dataset -> reqwest -> hyper/tower/h2` root from the `edge` dependency graph. |
||
|
|
0a6cb3b4cf |
[Raw payloads]: read payload as stored bytes in retrieve_raw (#10040)
* segment: read payload as stored in retrieve_raw `retrieve_raw` already hands back vectors as stored; let the caller ask for the payload the same way, so a reader that only relocates a point parses nothing. `RawPayloadFormat` states what the caller wants — no payload, parsed, or as stored — and replaces the `WithPayload` argument, which could express a key selection that a raw read cannot serve anyway. [`MaybeRawPayload`] states what came back, which can differ from the request in one direction only: a payload storage that keeps payloads parsed cannot answer `Raw` with a blob, and now says so instead of encoding a payload for a reader that would parse it straight back. The raw path reaches the blobstore through `read_payloads_maybe_raw`, mirroring `read_payloads` down the payload storage and payload index traits, so it keeps the batched read. Every caller asks for `Parsed`, so this changes no behaviour: the copy-on-write move and the sync comparison need the parsed payload anyway, and the shard transfer switches over with the feature flag that ships the blob to another node. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * segment: always hand out the stored payload blob from retrieve_raw Review follow-up: instead of telling `retrieve_raw` in which form to return the payload, it always returns it as stored and a caller that needs the parsed form decodes it itself. - Drop `RawPayloadFormat` and the payload parameter it replaced: no production caller ever asked for anything but the whole payload, and a selector cannot be applied to an opaque blob anyway. - Drop `MaybeRawPayload` / `MaybeRawPayloadRef`: only `InMemoryPayloadStorage` could produce the parsed variant, and no segment can be built with that storage (`PayloadStorageType` is `Mmap` or `InRamMmap`, both blobstore-backed). `SegmentRecordRaw` carries a plain `Option<RawPayload>`. - `PayloadStorageRead::read_payloads_maybe_raw` becomes `read_payloads_raw` and hands out `Option<&[u8]>`. The in-memory storage keeps payloads parsed, so it encodes on read, producing the bytes an on-disk storage would have written. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * api: decode a received raw payload with the shared decoder `decode_payload` at the gRPC boundary matched on the encoding and parsed the blob itself, duplicating `RawPayload::decode`. Add the inbound conversion from the wire type and let the one decoder do the reading, so another encoding has a single place to be taught. The conversion also rejects an encoding number no variant maps to, which prost would otherwise hand out as the default encoding — a blob from a node that writes payloads some other way must not be read as JSON. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * simplification --------- Co-authored-by: Ivan Pleshkov <ivan.pleshkov@qdrant.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7364cc42ef |
feat(edge): add query_batch for batched planned queries (#10100)
* feat(edge): add query_batch for batched planned queries Expose the planned-query batch path as a public API so multiple independent queries can share one planning pass over leaf searches and scrolls. Wired through EdgeShardRead, FFI, and Python bindings. Co-authored-by: Cursor <cursoragent@cursor.com> * perf(edge): push batched query vectors down to segments `query_batch` planned the whole batch at once but then executed every leaf search on its own: one query context, one fan-out over all segments, and one single-vector `Segment::search_batch` call per leaf. Execute the batch as a batch instead: - `EdgeReadView::search_batch` builds the query context once, visits the segments once, and hands each segment the leaves that agree on everything but their query vector as a single multi-vector `search_batch` call. `search` is now a thin wrapper over a one-element batch. - Move `SearchType`/`BatchSearchParams` from `collection`'s segments searcher into `shard`, next to `CoreSearchRequest`, and add `group_search_batches` so both the collection and the edge read path share one grouping implementation. Edge computes the grouping once and reuses it per segment. - `search_matrix` now issues its per-sample nearest queries through `query_batch`; they share filter, limit and vector name, so the whole sample is scored in one batched search per segment instead of one full segment pass per sampled point. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: generall <andrey@vasnetsov.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8db152b270 |
Make tonic optional in shard and clarify feature dependencies in Cargo.toml (#10069)
|
||
|
|
aa6c5d8403 |
Global quota API (#10035)
* feat: global quota API Memory and disk are node-wide resources, so configuring their thresholds per collection through strict mode makes little sense. Move them behind a single cluster-wide `QuotaManager`. The quota config is seeded from `storage.quotas` in the settings (and so from env vars), overridden by `quota.json` in the storage directory, and updated cluster-wide through a new `SetQuotaConfig` consensus operation which rewrites that file on every peer. Raft snapshots carry it too, so a peer that joins by snapshot picks it up. Quotas are enforced wherever the strict mode memory and disk checks used to run, but no longer gated behind `strict_mode.enabled`: a value set in an enabled strict mode config still wins per resource, the quota is the default. Rejections name both the condition that tripped and the config that governs it. `GET /quotas` reports the config plus current utilization to global read users; `PUT /quotas` replaces it for global manage users. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: cover the quota endpoints in the API consistency checks `test_all_rest_endpoints_are_covered` and the OpenAPI endpoint count both break on any new REST endpoint. Add `GET`/`PUT /quotas` to `ACTION_ACCESS` with their JWT access tests, and bump the expected API count. The quota endpoints stay out of `REST_ENDPOINT_WHITELIST`: that list is for data-plane endpoints reported per-endpoint in metrics. Also add a Raft snapshot CBOR compatibility test — snapshots are exchanged between peers of different versions during a rolling upgrade, so `quota_config` must be absent-tolerant in both directions. Review feedback: persist through `SaveOnDisk`, which already implements the write-before-swap protocol this was doing by hand; validate the config at both persistence boundaries, since a hand-edited quota file or a config arriving through consensus does not pass the REST handler's validation, and a `0%` limit would reject every update forever. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: assert seeding a quota from invalid settings persists nothing Follow-up to review feedback claiming `SaveOnDisk::load_or_init` writes the init value before it is validated. It does not — only `SaveOnDisk::new` persists — but the property matters: were seeding to persist first, invalid settings would leave a `quota.json` that fails validation on every subsequent start, and the node could only be recovered by deleting it by hand. Pin it down with a test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: make QuotaManager the single reader of memory and disk The quota checks measured memory and disk themselves, while the optimizer and the WAL disk watcher each called `fs4::available_space` behind their own ad-hoc caches. Fold all of it into QuotaManager: it owns the readings, the freshness policy, and the limits they are compared against. Moves the module to `lib/shard`, since the optimizer sits below `storage` and has to reach it; `storage::quota` re-exports it, so consensus, the `/quotas` API and StorageConfig are unchanged. The manager is installed as a process singleton by TableOfContent, ahead of loading any collection. - Callers hand in QuotaLimits overrides instead of a StrictModeConfig, and an override can now only tighten. A collection-level admin could raise `max_disk_usage_percent` past a cluster-wide limit that needed global manage rights to set; ties resolve to the quota so the rejection names the knob that actually has to change. - Measurements are cached for 5s, but a reading at or above its limit is never reused: a rejected client retries, and freeing the resource has to take effect on the next request rather than a TTL later. - `fits_on_disk` sizes an optimization against physical free space only, never the configured limits. Optimizations are what free a full disk, so the quota must not be what stops one. - `percent_of` widens to u128 instead of saturating the multiply, which under-reported utilization (failing open) above ~184 PB. - StorageConfig::quotas is optional; absent means no quota is enforced. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: don't recover dead replicas onto a node at a resource limit Recovering a dead replica pulls a whole copy of its shard onto this node. If it is already at its memory or disk quota that transfer cannot finish, and starting it only pushes the node further past the limit. Skip it and reconsider on a later sync, once the resource frees up. Adds QuotaManager::check_capacity for work that lands bytes here without being an update. Unlike fits_on_disk the configured limits do apply: taking on a replica is not what frees a full node, so there is no deadlock to avoid by letting it through. The check is hoisted out of the per-shard loop because a node over its limit re-measures on every call, so checking per dead shard would cost a statvfs each. It is free when no quota is configured. Also trims the comments across the quota module, which had grown well past what the code needs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: drop trivial and duplicated quota tests Six tests removed, ~140 lines, with no loss of coverage: - a_rejection_names_the_knob_that_has_to_change asserted that a format! contains its own literals; the message is covered end-to-end by the override test and by test_global_quota.py. - a_node_over_its_quota_has_no_capacity_to_take_on_a_replica was 30 lines for check_capacity, a one-line delegation to check_update the test above it already calls. - a_rejecting_measurement_is_never_served_from_the_cache duplicated the meter test, which proves the same rule with an injected reader instead of inferring it from the real filesystem. - free_space_is_reported_without_enforcing_anything covered a one-line accessor, and its point is what the fits_on_disk test is for. - The two resolve tests and the three meter tests each collapse into one. DiskFit::Unknown keeps its coverage as two lines inside the fits_on_disk test rather than its own fixture. The snapshot compat pair becomes one test: the second only asserted cluster_metadata.is_empty(), which says nothing about quotas — the real check was the deserialize, now an expect that states it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: re-measure free space as the disk fills, and drop a Windows-only assert Two CI failures, both from this branch. e2e test_low_disk: the DiskUsageWatcher I replaced escalated to checking on every call once free space fell below 512 MB. Folding it into the quota manager lost that — available_bytes passed no limit, so a reading was reused for the full 5s however little space was left. On a disk filling as fast as that test fills it, 5s blind is enough to actually run out and the WAL write dies instead of returning "No space left on device". available_bytes now takes a watch_below level and never reuses a reading under it, which is what the old ladder was expressing. The watcher passes max(min_free, 512 MB), so the escalation point is back; above it the 5s cache still costs fewer syscalls than the old 128-call ladder. fits_on_disk gets the same rule by passing required_bytes, so a merge that does not fit re-checks rather than sitting on a stale sample. Windows: fits_on_disk on a missing path was asserted to be Unknown, but GetDiskFreeSpaceEx resolves up to the containing drive and succeeds — as common::disk_usage's own test documents. Dropped; the branch is a two-line else and is not portably reachable. Also renames an_optimization_is_sized_against_the_disk_not_the_quota, which needed explaining to be understood. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: report the quota config in telemetry Reads it from the quota manager rather than the settings, so it is the config the node is actually enforcing: a peer that missed a consensus update reports what it is applying, not what the cluster agreed on. Gated on global access, the same access `GET /quotas` requires, and left out of `PeerTelemetry` — a quota is per-node state, so each peer reports its own. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: regenerate OpenAPI, and cover the quota in the telemetry key sets Two CI failures from the previous commit. Referencing QuotaConfig from TelemetryData moves its definition earlier in `components/schemas`, because TelemetryData is generated ahead of QuotaStatus. Regenerated rather than hand-patched, so the schema is a pure move. test_telemetry_detail asserts the exact set of top-level telemetry keys. The quota is reported at every level, including 0 — it is three scalars, it is the default the endpoint serves, and it is what explains an update being rejected — so both key sets gain it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: remove `max_disk_usage_percent` from strict mode Disk is a node-wide resource, so a per-collection percentage of it never meant anything a caller could act on: the limit describes how full the *node* is, and which collection the write happens to target has nothing to do with it. The global quota is where it belongs. It shipped in 1.18.2 without documentation, so this drops it outright rather than deprecating. Removal is soft in every direction: StrictModeConfig has no `deny_unknown_fields`, so a client still sending it gets it ignored rather than a 400, and the same struct deserializes the persisted collection config, so collections created on 1.18.2+ keep loading. Proto field 22 is reserved so the number is never reused. The e2e test becomes a quota test — the fixture and the timing are the interesting parts and they carry over unchanged; only how the threshold is configured differs. `max_resident_memory_percent` was documented and stays for now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: enforce the strict mode memory limit outside the quota `max_resident_memory_percent` was folded into the quota as an override, which meant the quota check had to know about strict mode, and retiring the setting would mean unpicking `EffectiveLimit` and `LimitSource` from the resolution logic. It is now a check of its own in `verification/mod.rs`, next to the strict mode checks it belongs with, borrowing only the measurement from the quota manager — which stays the node's single reader of process memory, so both checks still share one reading. Deleting the setting later is deleting one function and its one caller. `QuotaManager::check_update` takes no arguments and consults the quota alone. A collection can still only tighten the limit for itself, because its own check runs in addition rather than in place of the quota's, and each rejection now names the config that has to change without having to carry a `LimitSource` to say so. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: enforce the quota on the update path, not in strict mode The quota check sat inside `check_strict_mode_toc_batch` only because that was the one place holding the collection's strict mode config. It doesn't need one any more, and the placement had a real cost: coverage depended on each handler remembering to ask for a strict mode check, and four of the internal update RPCs do — `sync_internal`, which moves the most bytes onto a node, does not. It now runs in `Collection::update_from_client` and `update_from_peer`, which every update passes through. `update_from_client` checks ahead of the shard split, so an operation is accepted or refused whole rather than landing on some shards and being refused by others. Classification moves with it, from ~10 `consumes_memory` impls on request DTOs to one exhaustive `CollectionUpdateOperations::consumes_quota`. The internal enum has variants — raw upserts, conditional upserts, the syncs — that have no client-facing request type, so per-DTO impls structurally could not classify them. Shard-transfer syncs stay excluded, as they are today: a transfer is sized up once before it starts, and refusing its batches partway abandons work that is nearly done only for it to restart from the beginning. Index and named-vector creation reach shards through consensus, past this check — a peer must not refuse what the cluster agreed to — so they keep their pre-consensus check, now against the quota manager directly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: deprecate `max_resident_memory_percent` in strict mode Same reason the disk threshold went: memory is node-wide, so a per-collection percentage of it caps how full the *node* is, which has nothing to do with which collection is being written to. The node-wide quota caps it once for everything. Unlike the disk threshold this one shipped documented, in 1.18.0, so it keeps working — as a limit a collection can tighten for itself, never lift — and gets the usual markers: `#[deprecated]` on both Rust structs, `[deprecated = true]` on proto field 21, and `deprecated: true` in the OpenAPI schema, which schemars derives from the attribute. The note names 1.21 as the removal. Recording a version matters here: the audit in docs/plans/overdue-deprecations.md found that this repo has never written a removal deadline down, and members of the 1.15.0 deprecation batch are still in tree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: reconcile the quota readers with #9891 #9891 landed effective (cgroup) figures in telemetry while this branch was making QuotaManager the single reader of memory and disk. Two collisions, neither of which git sees. `segment::utils::mem::total_memory_bytes` is now a shared accessor with a 5s TTL, so a cgroup resize is picked up. The quota module had its own `OnceLock` copy that froze the value at startup — exactly what #9891 set out to fix — so it delegates to the shared one instead. Telemetry's new `disk_size` called `common::disk_usage::disk_usage` directly. That reader lost its TTL cache on this branch when the caching moved into the quota manager's meter, so it would have taken an uncached `statvfs` on every telemetry request, and it put a second disk reader back in the tree. It goes through `QuotaManager::disk_capacity_bytes` now, sharing the reading the quota check already takes. Verified it still reports the storage filesystem, matching `df`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: split the quota manager by what each half does `manager.rs` had grown to 450 lines holding three separate jobs: owning the config and its file, taking the readings, and comparing one against the other. - `manager/store.rs` — the `Store` enum, `QUOTA_CONFIG_FILE`, and config validation, which is now the store's own business rather than something every caller has to remember to do first. - `manager/measure.rs` — every reading, and `DiskFit`. The "nothing else calls `statvfs` or reads process RSS" claim is now checkable by looking at one file. - `manager/enforce.rs` — `check_update` / `check_capacity` and the threshold comparison. - `manager/mod.rs` — the struct, its construction, and the config accessors: what a reader needs to see first. Tests move with their subject. No behaviour change: `set_config` used to validate before delegating to the store, and now the store validates on write, which is the same order of operations from the outside. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: count copy-on-write deletes toward the quota Dropping a vector or a payload key does not free anything on its own: copy-on-write rewrites the point to produce the version without that field, so storage grows first and is only reclaimed once the optimizer gets to it. Gating those as if they were reclaiming space let a full node keep taking writes that make it fuller. Deleting whole points stays exempt. That is the one operation that has to work on a node at its limit, or there is no way back under it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: say that the quota's reported usage is per node `GET /quotas` returns one cluster-wide config and one set of utilization figures, which reads as though both describe the cluster. They do not: memory and disk are node-local, so `usage` is whatever the peer that served the request is seeing, and a peer under its limit says nothing about the others. Also corrects `resident_memory_percent`, which claimed to be a share of total system memory. It is a share of the memory available to the process, which under a cgroup is the limit rather than the host's RAM. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: treat a node over its quota as a failed replica, not a bad request A quota rejection described the request as invalid (400) and was classified non-transient, which is how the replica set recognises errors that every replica would produce alike. A quota is the opposite: the input is fine and the answer depends on which machine you ask. On the default `wait=false` path that combination silently dropped the write — `update.rs` only deactivates transient failures when nothing completed — leaving the replica Active and permanently missing data its co-replicas had. It is now `InsufficientStorage`, transient, HTTP 507 / gRPC `ResourceExhausted`. So a node that is out of room is handled like one that is offline: - last active replica, or every replica over quota: nothing could take the write, and the client is told the cluster is out of room. - more than one replica: the full node is deactivated through the same path a dead peer takes, and the update stands if enough replicas accepted it. `check_capacity` already keeps recovery off that node until it has room. The check also moves off `update_from_client`, which applied the coordinator's own limit to the whole operation even when it held no replica of the shards being written. Each replica set now gates its own local write and records the refusal as a failure of this peer, so a node only ever answers for itself. `ResourceExhausted` is shared with rate limiting, and the reverse conversion mapped it straight to `RateLimitExceeded` — a forwarded rejection came back as 429. Statuses now carry a marker so the two stay distinguishable across the wire. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: report quota pressure per node, and across the cluster A quota is node-local, so finding out which node has hit one meant asking each of them in turn — and nothing at all showed up in monitoring. `/metrics` gains a `quota_exceeded` gauge for the local node. It is emitted only while the quota is enabled: with it off the value would be a constant 0 that says nothing about the node, and an alert built on it would go quiet rather than fire if someone disabled the quota. Telemetry's `quota` field carries the same verdict alongside the config, since that is where the metric is derived from. `GET /quotas` now answers for the whole cluster. A new `GetQuotaUsage` RPC on the internal `QdrantInternal` service returns what one peer is using, and the handler fans it out to every known peer in parallel. Peers that do not answer are left out rather than failing the request — the nodes that are out of room are exactly the ones most likely to time out, and a partial answer still names them. Outside distributed mode the field is absent rather than a map of one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: report the quota metric per resource `quota_exceeded` was one flag for the whole node, which does not say what to go and fix — disk is freed by deleting or optimizing, memory by unloading. It now carries a `resource` label: quota_exceeded{resource="memory"} 0 quota_exceeded{resource="disk"} 1 A resource with no limit gets no series at all, for the same reason the metric is absent while the quota is disabled: a series that can never reach 1 reads as healthy and would quietly carry an alert that cannot fire. `QuotaManager::exceeded` returns the per-resource verdict, with `None` for a resource this node does not cap. Telemetry reports the same breakdown, since the metric is derived from it. The peer usage RPC keeps a single flag — it sits next to both percentages, so it only has to answer "is this peer refusing writes". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: drop a no-op error conversion the linter caught `check_global_access` already returns a `StorageError`, so mapping it through `StorageError::from` converted the type to itself and tripped `clippy::useless_conversion`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: hold a tripped quota until usage clears a release margin A resource resting on its limit crosses it in both directions on the noise between two readings, and each crossing is expensive: the node refuses a write, its replica is deactivated, usage dips, recovery starts sending a whole shard copy back, and the arriving data pushes it over again. The loop sustains itself, and every lap costs a shard transfer. A limit now trips at its configured value but only clears once usage has fallen 5 percentage points below it, so the crossing has to be real. The margin is floored at 1%, since a limit smaller than the margin would otherwise be impossible to fall back under and would strand the node. The verdict is carried on the manager rather than recomputed, which makes it the thing reporting shows: expect `exceeded` to be set while the utilization next to it is already back under the limit. Rejections say so too, rather than claiming a limit that is no longer exceeded: Disk usage is at 87% of total capacity. It reached the configured limit of 90% and has to fall below 85% before this node takes writes again. Changing the config clears the verdicts. New limits are a deliberate act, and should not be held back by the margin of a limit that no longer exists. Both resources are now evaluated on every check instead of stopping at the first failure, so a verdict is never left behind reporting a reading that has since been superseded. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: make the quota release margin configurable 5 points is a guess about how noisy a deployment's usage is, which is not something one number can be right about: a node whose disk moves in gigabyte steps needs a wider margin than one that creeps, and an operator who wants the old flip-on-every-reading behaviour should be able to ask for it. `release_margin_percent` joins the rest of the quota config, so it seeds from `QDRANT__STORAGE__QUOTAS__RELEASE_MARGIN_PERCENT`, replicates through consensus, and changes with `PUT /quotas`. Defaults to 5 and is filled in when a request omits it, so it always answers with the margin actually in force rather than leaving the caller to assume one. `0` releases as soon as usage is back under the limit. `QuotaConfig` grows a hand-written `Default` for it, since deriving one would have quietly defaulted the margin to 0 and disabled the hysteresis for anyone constructing a config in code. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: leave the release margin unset by default, and hold verdicts in atomics `release_margin_percent` is `null` unless someone sets it, rather than materialising 5 into every config. A quota written today then does not pin a number a later release may want to revise, and `{"enabled": false}` still round-trips as itself. `QuotaConfig::limits` resolves it, next to `enabled`, so enforcement never sees the unset case. The verdicts move from a `Mutex<QuotaExceeded>` to one `AtomicBool` per resource. They are judged independently and nothing reads them as a pair, so the lock only added contention to the path every update takes; a verdict that races a concurrent check is re-decided by the next one from a fresh reading. That also drops the tri-state. Only "was this over its limit" has to survive between checks — whether a resource is enforced at all follows from the config and the reading, so it is derived when reporting rather than stored. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: drop a comment arguing with a design that was never here Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
75293e5990 |
[Payload bytes] decode raw payload directly (#10032)
* Deserialize raw payload directly to Payload during conversion * Treat empty payloads equally * Fix edge |
||
|
|
75385df69f |
Remove dead code (#10030)
* Remove dead code * Remove unused dependencies * `allow(dead_code)` -> `expect(dead_code)` * ast-grep: rule-tests/*-test.yml => tests/*-test.yml For brevity. * ast-grep: forbid allow(dead_code) |
||
|
|
ec1aec5f4c |
Payload bytes gRPC (#9949)
* Prepare GRPC for raw payload bytes * Make RawPayload a separate protobuf message |
||
|
|
0a16a62f99 |
feat: io_uring setting to control which components use the io_uring backend (#10008)
* feat: `io_uring` setting to control which components use the io_uring backend A few components have both an mmap and an io_uring variant reading the very same files: the immutable dense vector storages, the single-file TurboQuant storage, and the mmap payload storage. Until now the choice was a side effect of `async_scorer` — a vector-search knob — plus, for the payload storage, a feature flag that was parked off because io_uring is ~2x slower than mmap when the data fits the page cache (#9310, #9409). Add `storage.performance.io_uring`, optional, with two modes: - unset (default): unchanged behaviour. The vector storages keep following `async_scorer`; the payload storage stays on mmap. - `disabled`: no component uses io_uring. - `auto`: a component uses io_uring when its memory placement is `cold` (data is left on disk, so reads hit the disk and there is something to gain), its feature flag allows it, and the kernel supports io_uring. Components meant to sit in RAM keep using mmap. The decision lives in one place, `segment::common::io_uring::use_io_uring`, so the openers no longer each reach for the async-scorer global. Kernel support is now probed up front through `is_io_uring_supported()` instead of opening a file and falling back on error. `async_payload_storage` now defaults to on: it no longer decides anything by itself, it only lifts the ban, and the payload storage no longer follows `async_scorer` at all — so turning it on cannot silently move an existing `async_scorer: true` deployment onto the slower path. Which backend a component ended up on depends on the config, the placement and the kernel at once, so report it in `SegmentInfo`: `vector_data[name].io_backend` and `payload_storage_io_backend`, both `"mmap" | "io_uring"`, absent for components that have no such choice. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Trim comments, drop trivial tests Two tests were only restating their own implementation: `test_mode_round_trip` round-tripped the encode/decode pair next to it, and `test_io_uring_config` checked that serde deserializes a two-variant enum. The mode matrix test stays, it is the one that pins the semantics. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Flatten `IoBackend` in OpenAPI, derive `JsonSchema` for `IoUringMode` Per-variant doc comments on a plain string enum make schemars emit a `oneOf` of anonymous single-value objects instead of a flat `enum`. Move the variant descriptions into the enum doc, as `Memory` and friends already do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Update lib/segment/src/vector_storage/turbo/turbo_vector_storage.rs Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com> * Update lib/segment/src/types.rs Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com> * Update lib/segment/src/types.rs Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com> * Update lib/segment/src/types.rs Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com> * Update lib/segment/src/types.rs Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com> * upd openapi schema * Update lib/common/common/src/flags.rs Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com> * Require kernel io_uring support in the async-scorer fallback `use_io_uring` returned `get_async_scorer()` verbatim when the `io_uring` setting is unset, so an enabled async scorer on a kernel without io_uring opened the io_uring storage, failed, and fell back to mmap with an error log per segment. Gate that branch on `is_io_uring_supported()` too, like `Auto` already is, so the component just stays on mmap. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * upd openapi schema --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com> |
||
|
|
ba3043b343 | ConditionChecker::check_batched: use in hnsw (#9845) | ||
|
|
446d140c2d |
Slice filtering condition: sliced scroll / deterministic sampling (#9899)
* feat: slice filtering condition for sliced scroll and deterministic sampling
Add a `slice` filter condition selecting points where
`stable_hash(point_id) % total == index`. The hash is SipHash-2-4 with a
zero key over canonical id bytes (8 LE bytes for numeric ids, 16 RFC 4122
bytes for UUIDs) — a frozen public contract, independent of the internal
resharding ring hash, reproducible by clients to predict membership.
For a fixed `total`, slices are disjoint and cover all points, enabling
parallel scroll streams (ES sliced-scroll style) and reproducible sampling
that composes with any other filter condition.
- REST: `{"slice": {"total": N, "index": R}}`; gRPC: `SliceCondition` in
the condition oneof (tag 8)
- Evaluated per point via id_tracker external-id lookup; no payload index
needed; cardinality estimated as `points / total` with no primary clause
- `total >= 1` enforced by NonZeroU32 at parse time, `index < total` by
validation in both REST and gRPC paths
- Hash contract locked by test vectors independently reproduced with a
reference SipHash-2-4 implementation
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* tests: minimal OpenAPI test for slice filter condition
Scrolls all slices of a fixed total over numeric + UUID ids asserting
disjointness and full coverage, checks must_not inversion, and pins the
two rejection paths (422 for index >= total, 400 for total = 0). Requests
and responses are validated against the regenerated OpenAPI spec by the
test harness.
Note: the spec cannot itself reject total = 0 client-side — the Condition
anyOf falls through to the permissive Filter schema, as with any invalid
condition — so rejection is asserted via the server response.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
0982e8699c |
feat: segment manifest optimizing state with lease (#9873)
* feat: segment manifest optimizing state with lease * fix: clippy * fix: exhaustive manifest state matching * fix: merge manifest rebuilds under the write lock * refactor: named state predicates, drop redundant enumerator test Review follow-up: move the enumerator's filter into SegmentManifestState::is_usable, name the preserving predicate is_optimizer_mark, delete the enumerator test that re-tested serde. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: generall <andrey@vasnetsov.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c196d2eb1a |
Benches: use SmallRng instead of ChaCha12-based generators (#9887)
* Benches: use SmallRng instead of ChaCha12-based generators All benchmarks used StdRng or rand::rng() (ThreadRng), both backed by the ChaCha12 block cipher in rand 0.10. Benchmarks do not need crypto-strength randomness, and several draw random values inside the timed closure, so cipher work was included in the measurement itself. Switch every bench target to SmallRng (Xoshiro256++), and key the HNSW graph cache and sparse index cache by RNG algorithm so stale caches built from the old generator are not reused against newly generated vectors. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Benches: replace free-function rand::random with local SmallRng Addresses review: rand::random draws from the thread RNG (ChaCha12), including inside the timed loop of the pq score benchmark. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
247dbbbb07 |
Return sorted Vec from SegmentEntry::vector_names (#9870)
The HashSet bought nothing: names are unique by construction (map keys) and all callers only iterate. A sorted Vec skips the hashing and set allocation, and makes the iteration order deterministic. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8cbd061b84 |
Deduplicate plain and raw upsert/sync drivers via point traits (#9877)
* Deduplicate plain and raw upsert drivers via PointToUpsert trait upsert_points and upsert_points_raw were ~60-line clones differing only in how the point struct writes itself into a segment. Extract a private PointToUpsert trait with the two variation points — upsert_into (in-place write) and write_moved (CoW-move record transform) — implemented for PointStructPersisted and PointStructRawPersisted, and fold the chunked driver into a single generic upsert_points_impl. The public functions keep their names and signatures as thin wrappers. The duplicated upsert_with_payload/upsert_raw_with_payload tails collapse into one shared set_full_or_clear_payload helper, which also carries the single has_point debug assertion. No behavior changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Deduplicate plain and raw sync drivers via PointToSync trait sync_points and sync_points_raw were ~80-line clones of the same 5-step algorithm, differing only in the retrieval call (retrieve vs retrieve_raw) and the stored-record type compared against. Extend the upsert approach with a PointToSync subtrait carrying an associated StoredRecord type, retrieve_stored, and is_equal_to (delegating to the existing inherent methods), and fold the drivers into a single generic sync_points_impl. Step 5 calls upsert_points_impl directly; PointToUpsert and upsert_points_impl become pub(super) to be visible within points/. Public signatures unchanged. No behavior changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fe7ca8443a |
Split shard update module into per-operation submodules (#9876)
lib/shard/src/update.rs grew to 2200+ lines. Split it into an update/ directory by operation kind, moving code verbatim: - mod.rs: process_* dispatch entry points + re-exports (public API and crate-internal paths are unchanged) - points/: upsert.rs (plain, conditional and raw), delete.rs (by id and by filter), sync.rs (plain and raw) - vectors.rs, payload.rs, field_index.rs: per-kind apply functions - helpers.rs: shared filter-based point selection (incl. the deferred points corner case) and check_unprocessed_points - tests.rs: the test module, unchanged Only additions are per-file imports, re-export lists and two pub(super) visibility bumps for now-cross-module helpers. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |