Commit Graph
192 Commits
Author SHA1 Message Date
Tim Visée 9f01222ccf Integrate new bitflags structure (#10123)
* Add `FlagsMode::from_feature_flags`, the mode for newly created flags

Compact in serverless-compatible deployments, dynamic otherwise. Only
creation consults it; opening existing flags detects their mode from
disk.

* Support the compact mode in the read-only flags types

Add `ReadOnlyFlags`, the mode-dispatching union of the two read-only
counterparts, serving the shared `RoaringFlagsRead` surface. Teach
`InMemoryBitvecFlags` to detect the mode it opens; its compact
`reload_appended` decodes the whole (small) file, as the format has no
random access.

* Create flags through mode selection in storages and indexes

Vector storage deleted flags and the bool/null indexes now open through
`open_or_create` with the mode from the feature flags: serverless
deployments create compact flags, dedicated ones keep creating dynamic
flags, and existing flags are opened in their detected mode either way.

* Read flags in either mode in the read-only bool and null indexes

`ReadOnlyFlags` shares the `RoaringFlagsRead` surface and the lifecycle
signatures of the roaring type it replaces, so the swap is a type
rename.

* Add TODO to not lock bitmask structure during flush

* `MutableStoredBitmask::save` returns the number of bytes written

Zero when the skip-clean save wrote nothing. Lets wrappers charge the
actual write to a hardware counter.

* Refuse to open compact flags in a dynamic-mode directory

Creating the compact file next to dynamic files would leave a directory
of both modes behind, which every later open rejects — refuse up front
instead. Both production callers already rule the case out through
`FlagsMode::detect`, so this only removes a foot-gun for future callers.

The open-or-eagerly-create logic moves into
`open_or_create_compact_mask`, shared with the update-only writer next.

* Rewrite `UpdateOnlyStoredFlags` onto the compact bitmask

The update-only flags writer now writes the compact mode — a single
roaring-encoded `compact_flags.dat` through `MutableStoredBitmask` —
instead of rewriting the whole padded dynamic file pair every batch. A
flush with no effective changes now writes nothing at all, where the old
writer rewrote the full mask on any `set`.

This also fixes opening serverless-created segments: the old open
eagerly wrote a `status.dat` into directories the writable side had
created in the compact mode, leaving files of both modes behind and
poisoning the directory for every later open.

A directory already holding dynamic-mode flags is refused loudly rather
than kept current or migrated; rebuild the segment to migrate its flags.
Migration may come later.

Drops the now-dead `InMemoryBitvecFlags::into_bitvec` and
`DynamicFlagsStatus::new`, and demotes `file_size_for` to private.

* Run edge tests with serverless feature flags

The edge fixtures ran with default feature flags, building leader shards
with dynamic-mode flags — a configuration edge never serves in
production, and one the update-only flags writer now refuses. It also
hid that the writer poisoned compact directories: no test exercised
update-only writes over a serverless-created shard.

Feature flags are process-global and first-init-wins, so every fixture
in the binary initializes the same serverless set; the manifest test
folds into it, since serverless implies `write_segment_manifest`.

* Don't use sequencial mode for one shot reads
2026-08-18 17:51:21 +02:00
Andrey VasnetsovandClaude Fable 5 087c29289f Edge: seed appendable segment on load from existing segments' indexes (#10257)
`ensure_appendable_segment` used a shard-root `payload_index.json` that
nothing in edge ever wrote, so a shard loaded with only immutable
segments got a bare appendable segment and the appendable chain stayed
unindexed until a merge happened to include an indexed segment.

Build the segment directly and seed it with the union of
`get_indexed_fields()` over the loaded segments, the same reconciliation
the optimizer performs for its CoW segment. Drops the shard-root file
from edge.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 15:51:14 +02:00
Luis Cossío 57e7389f91 [LiveReload] Prepare segment preload (#10221)
* genericize live_reload fs parameters

* impl live_preload for ReadOnlySegment

* split edge refresh into preload and apply passes

* only rotate file infos after successful reload
2026-08-17 21:43:24 -04:00
Andrey VasnetsovandClaude Opus 5 81e9fb3e75 [UpdateOnly] Honor upsert update_mode in the batch writer (#10236)
* [UpdateOnly] Honor upsert update_mode in the batch writer

The writer rejected every `UpsertPointsConditional`. Accept the ones whose
condition is empty — `insert_only` and `update_only` — since existence is
the whole gate they need, and locating a batch's points already answers it.

The gate is evaluated per mutation at its position in the fold, so an
`insert_only` upsert sees a point an earlier operation of the same batch
created, matching a leader that resolves each operation only after the ones
before it were applied. A conditional upsert may therefore not discard the
mutations it follows.

Rejecting an upsert also means never reading the point it would have
overwritten: `needs_stored_point` asks whether the first mutation that
applies to an existing point discards it, so an `insert_only` batch pays
nothing for the ids that are already taken.

A conditional upsert carrying a real filter is still rejected — evaluating
one needs payload indexes the writer never fetches.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [UpdateOnly] State contracts in the update-mode docstrings

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [UpdateOnly] Trim the update-mode diff

Drop `--update-mode` from edge-shard-update: the modes are covered by unit
and end-to-end tests, and the flag cost a wrapper enum, a conversion and a
parameter threaded through both run paths. The tool still reports rejected
points, which the exhaustive match requires.

Inline `always_applies` into its one caller, drop the two test-batch
wrappers over `conditional_batch`, drop the `update_only` fold test whose
truth table two other tests already assert, and shorten two over-long
comment blocks.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 14:05:13 +02:00
Andrey VasnetsovandClaude Fable 5 e32d3fbf89 Add acosh expression to formula query (#10231)
Unary inverse hyperbolic cosine, parallel to sqrt/ln/exp/log10, in REST,
gRPC, and edge (FFI + Python) interfaces. Inputs below 1 produce the same
NonFiniteNumber error as an invalid sqrt or ln.

Closes #10186

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 10:47:20 +02:00
Andrey VasnetsovandClaude Opus 5 bb43c9d057 gitignore: ignore lib/edge/publish/target (#10237)
`amalgamate.py` builds the generated crate in place, leaving a Cargo target
directory next to the sources. `/examples/target` was ignored but the
publish crate's own was not, so it showed up as ~32k untracked files that a
path-wide `git add` picks up.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 10:20:44 +02:00
Andrey VasnetsovandClaude Fable 5 5f8cef9ebd [UpdateOnly] Writer over object storage (#10214)
* Drop the vestigial UniversalWrite bound from the update-only writer

Neither writer kind performs in-place writes: AppendableSegment is built on
UniversalAppend, and DeleteOnlySegment tombstones via whole-mask atomic_save
(UniversalWriteFileOps), which UniversalAppend's supertrait already carries.
The bound is a leftover from the DiskIdTracker-based iterations that mutated
the deleted mask in place.

With it gone, UpdateOnlyEdgeShard::apply_batch is instantiable with the
object-store-appendable CachedBlobFile.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* edge-shard-update: --apply writes the batch, over object storage too

Open the object-storage backends through CachedBlobFs/CachedBlobFile
instead of the read-only DiskCacheFs handle, so the shard is appendable in
both modes, and add --apply: generate the same schema-derived batch and run
apply_batch instead of preview_batch. Dry run stays the default and the
generation is shared, so the preview cannot drift from what an apply would
do. AwsConfig::native_append is exposed as --native-append for
AiStor/RustFS-style endpoints; the Cached* types join io_bridge_object_store's
re-export of the io_bridge stack.

Applying to a leader-produced shard currently fails with a clean refusal —
its appendable segment's payload storage was created in mutable mode, which
the append-only writer rejects — the known segment-bootstrap gap, next in
line.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* CachedBlobFile: latency tracing for append_bytes

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* UpdateOnlyEdgeShard: sequential batches through one writer

Writers open once at shard open, next to the lookup segments they resume
from. apply_batch hands the writer back on success, live-reloading the
lookup half of every segment the batch wrote to (new
LookupSegment::live_reload, mirroring the read-only segment's); on error
the writer is consumed, since its lookups may no longer describe the
durable state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* edge-shard-update: --interactive mode, sequential batches on one writer

After each applied batch, prompt on stdin for the next round's ids and
apply them through the writer apply_batch handed back — no shard
re-open — with op-num (and seed) incremented per round.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* CachedBlobFile: create the missing object on an offset-0 rewrite append

The caller-side rewrite path (part-copy S3 stores below the direct-append
threshold) validated the offset against the mirror length, whose
initialization HEAD-requests the remote and surfaced NotFound for an
object that does not exist yet. Direct-append backends (GCS compose,
native append) already create the object on an offset-0 append; the
rewrite now reads a missing remote as length zero so its whole-object PUT
does the same, and a non-zero offset against a missing object reports an
offset conflict.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* UpdateBatchOutcome: per-point records of retired slots

Each applied point now carries a PointApplyRecord: what happened to it
(stored/deleted/skipped/missing) and which slots it vacated where —
tombstoned per segment, or superseded in place for the old write-target
copy of a stored point. Built in the same loop that decides
tombstone-vs-supersede, so the report cannot drift from the writes.

edge-shard-update logs one line per point after the applied summary,
telling a fresh insert from an overwrite and naming the segments the old
copies were deleted from.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 20:06:00 +02:00
Andrey VasnetsovandClaude Fable 5 06ffcb881f Add CachedBlobFile: cached reads + write-through appends for object stores (#10206)
* Add CachedBlobFile: cached reads + write-through appends for object stores

Combine a DiskCache mirror (reads) with a BlobFile remote handle (appends)
into CachedBlobFile/CachedBlobFs, the appendable universal-IO citizen for
object stores. Appends perform the remote mutation inline and are durable
at Ok: a native write-offset append in AppendMode::Native (with a soft
limit on appends per object), or a whole-object rewrite in
AppendMode::Rewrite for stores without native append. After a successful
append the mirror length is advanced without extra IO; appended blocks
fault in from the remote on first read.

The multipart UploadPartCopy rewrite path (prefix >= 5 MiB) and the
rewrite-required error classification are left as todo!() pending the
AsyncRewrite backend capability.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Backend-advertised AppendMethod; reactive appended-block cap recovery

Replace CachedBlobFile's stored AppendMode with AsyncAppend::supported_append:
the backend advertises Native or PartialUpload, and append takes a matching
AppendRequest variant, rejecting the ones it does not support. The multipart
UploadPartCopy todo moves into the S3 backend's PartialUpload arm.

Drop the native_appends soft-limit counter: it is per-handle in-memory state
that resets on every restart, so it can never be the correctness mechanism
and persisting it would not make it authoritative either. The store is the
authority: hitting its appended-block cap now surfaces as the new
UniversalIoError::AppendRewriteRequired (S3 400 TooManyParts), and
CachedBlobFile recovers with a whole-object rewrite. Unrecognized errors
stay hard errors instead of silently triggering rewrites.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Per-store append strategies; server-side rewrites for plain S3 and GCS

Replace the single AppendContext struct with an enum of strategy objects,
one per store capability, each owning its append logic:

- NativeAppend: the signed write-offset PutObject (S3 Express, MinIO
  AiStor; AwsConfig::native_append declares it for AiStor-like endpoints,
  s3_express implies it).
- PartCopyAppend: plain S3 — appends land as one atomic multipart rewrite
  whose prefix parts are server-side UploadPartCopy requests; nothing but
  the appended data crosses the network. object_store keeps such
  provider-specific calls out of its portable surface, so the requests are
  hand-signed like the native append.
- ComposeAppend: GCS — the appended data is uploaded as a temporary
  neighbor object and composed onto the destination server-side,
  conditional on the observed generation (a real compare-and-swap).

AppendMethod is replaced by AppendSupport, which tells the caller the only
thing it needs: when the store takes a direct append. Always (native, and
compose: no part minimums, no block cap), AboveThreshold (part-copy: the
copied prefix lands as non-last multipart parts, >= 5 MiB each), or Never.
CachedBlobFile drops its hardcoded MIN_COPY_PREFIX and rewrites locally
only below the backend-advertised threshold; AppendRequest::Rewrite now
means only "append and rebuild as a single blob" — the appended-block cap
recovery.

The append module is split one file per strategy, with a shared
SignedRequestContext transport and a test-only HTTP stub.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* DiskCache tracks the remote object's etag

Seeded from the new known_etag open extra (OpenExtra::with_known_etag),
refreshed from FileInfo on schedule_reopen, and settable directly for
callers that mutate the remote out of band.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Remove AppendRequest enum; appended-block cap recovery moves into the backend

AsyncAppend::append takes plain (path, offset, data). A native S3 store
that rejects an append with TooManyParts now falls back to the part-copy
rewrite inside the dispatcher, instead of surfacing AppendRewriteRequired
to CachedBlobFile for a second Rewrite request. The Rewrite variant was
handled identically to Append everywhere except that one native path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Escalate to download+rewrite when the store rejects a part-copy rewrite

The cap-recovery rewrite is chosen by the store's returned error, not a
client-side threshold: a part-copy attempt rejected with EntityTooSmall
(typed as UniversalIoError::AppendEntityTooSmall, parsed from the S3
error <Code>) falls back to downloading the sub-part-minimum prefix and
PUTting the whole object back, guarded by a prefix-length offset check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Fix S3 Express appends: zonal endpoint + s3express SigV4 service

Hand-issued appends targeted the standard endpoint and signed as "s3",
so every append to a directory bucket got 404 NoSuchBucket, masked as
AppendOffsetConflict by the 404 mapping. Derive the zonal
{bucket}.s3express-{az}.{region} base from the mandatory --{az}--x-s3
bucket suffix (mirroring object_store's private derivation), carry the
SigV4 service name in SignedRequestContext, and treat a 404 as a
conflict only for NoSuchKey or bodiless responses — NoSuchBucket stays
a loud error guarding the endpoint derivation. extract_xml_tag moves up
to the context module and now tolerates tag attributes and
pretty-printed bodies.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Server-side etag precondition on appends; BlobFile loses UniversalAppend

AsyncAppend::append carries an expected_etag that S3 part-copy rewrites
attach as x-amz-copy-source-if-match (412 -> AppendEtagMismatch, a new
typed error) and download_rewrite checks against the GET's own etag;
native write-offset PUTs and GCS compose ignore it. BlobFile appends
only through the inherent etag-aware append_bytes now — CachedBlobFile
calls it directly with its DiskCache-tracked etag — and BlobFs's
mutating ops become inherent, delegated from CachedBlobFs, per the
standing TODOs. The append conformance battery runs over the
CachedBlobFs stack, via new direct constructors that share one backend.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Drop unfulfilled too_many_arguments expectation

rewrite_parts has exactly seven parameters — at the clippy threshold,
not over it — so the lint never fires and the expect fails CI under
-D unfulfilled-lint-expectations.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 18:26:49 +02:00
Ivan PleshkovandClaude Opus 5 86b9330628 transfer: send raw payloads, behind feature flags (#10066)
A raw point can carry its payload as the byte blob it is stored as, mirroring
`PointStructRaw.raw_payload` on the internal gRPC API. The blob travels from the
sending node into the receiving node's WAL untouched, so the sender never parses
the payload it read and neither node builds a protobuf value tree for it.

It is parsed exactly once, where the operation is unpacked for apply
(`process_point_operation`), because that is the first place the parsed form is
actually needed: `set_full_payload` goes through the payload index, which cannot
be updated from bytes. The gRPC boundary therefore only checks the encoding tag
and rejects a point that sets both payload fields, the way the enclosing request
already rejects both `points` and `raw_points`.

Moving the parse onto the apply path makes its error classification load-bearing,
so a malformed blob is reported as `OperationError::MalformedPayloadBlob` — the
payload sibling of `MalformedVectorBlob`, mapped to `CollectionError::BadInput`
for the same reason: a bad blob that reached the WAL has to be skipped on replay
instead of crash-looping recovery.

Three consequences of the blob living that long are handled explicitly rather
than by convention:

- `decode_payload_raw` takes the blob only once it has parsed, so a failure
  leaves the point holding it instead of holding neither representation.
- `upsert_points_raw` and `sync_points_raw` refuse a point that still carries a
  blob. They read the parsed payload, so such a point would otherwise be stored
  with no payload at all, and a `debug_assert!` would not catch it in release.
- `is_equal_to` compares blob to stored blob as bytes. A differing encoding costs
  a redundant upsert on sync, never a skipped one.

The `raw_payload_transfer` bench measures the trade, per 100-point batch (one
transfer batch) at payloads of ~200 B / ~700 B / ~7 KB:

- Sender, storage bytes to wire: 16x / 37x / 113x faster. This is where the whole
  win is — no parse of the blob that was read, no value tree built.
- WAL encode: 5x / 11x / 25x faster, writing a byte string instead of a map.
- Receiver, wire to applicable point: 1.09x / 1.10x / 1.06x. Near neutral, as it
  swaps walking a prost value tree for a JSON parse.
- Wire bytes: ~6% smaller. WAL bytes: 10-32% *larger*, because the blob is JSON
  while a parsed payload is written as a compact CBOR map.

The WAL growth is accepted rather than fixed: decoding earlier to win those bytes
back costs a second full deserialization, and would leave the receiving side with
a `payload_raw` that is never populated. Making the blob itself compact belongs in
the payload storage encoding (`RawPayloadEncoding` is the extension point for it),
not here.

Two flags, both off by default and both sender-only (nodes accept raw points and
raw payloads regardless), read where the transfer batch is prepared:

- `transfer_raw_points` transfers every collection as raw points, not only those
  whose vector storage would drift in a decode-encode round-trip.
- `transfer_raw_payloads` ships the blob a raw read hands out; without it the
  prepared batch decodes it back into the parsed payload, and the wire message is
  exactly what it is today.

Neither is enabled by `all`: a node only accepts them once it runs a version that
understands them, so they can only be switched on a release later. Nothing
enforces that yet — the transfer has no peer-version gate.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 11:28:50 +02:00
Andrey VasnetsovandClaude Fable 5 b43f70a6b6 [UpdateOnly] tombstone points in immutable segments via whole-mask rewrite (#10196)
* [UpdateOnly] tombstone points in immutable segments via whole-mask rewrite

DeleteOnlySegment::tombstone_points marks the retired slots in the
segment's deleted-points bitmask (id_tracker.deleted, shared by the
immutable and disk-resident tracker formats) and replaces the file
whole via atomic_save — the one mutation that works on backends
without random-offset writes. Both read-only trackers already
live-reload this file by opening a fresh handle and diffing, so the
rewrite needs no read-side changes.

The mutation cycle lives in StoredBitSlice::atomic_update: read the
stored bits (or start from a caller-provided seed), apply the update,
save atomically; a closure error writes nothing. The seed comes from
the read phase by analogy to AppendableIdTrackerState:
LookupSegment::writer_state now returns WriterIdTrackerState, whose
DeleteOnly variant carries the deleted mask when the tracker already
holds it in memory — always for the immutable tracker, only if
materialized for the disk-resident one, which deliberately avoids
loading the full deleted set.

Tombstoning needs no more of the backend than reads plus atomic_save,
so DeleteOnlySegment's bound drops to
UniversalRead<Fs: UniversalWriteFileOps>.

Unlike the writable trackers' drop(), the slot's version is not zeroed
(the versions file is in-place-mutated, which object stores cannot do):
deletion authority in these formats is the bit — every lookup filters
through it — and a stale version on a tombstoned slot is the same state
a crash between drop-bit and drop-version leaves, which
fix_inconsistencies already absorbs as storage cleanup.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Close the temp-file handle in tests that atomically replace it

NamedTempFile holds the file open for its lifetime, and Windows refuses
the rename in atomic_save while any handle is open. into_temp_path()
closes the handle and keeps the deletion guard.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-12 09:56:23 +02:00
Daniel Boros 638f8aad63 edge-tool: bound upload and upsert memory, skip the WAL, fix generated payload paths (#10189)
* fix: bound edge-tool resource use and fix generated payload paths

* fix: keep sibling array elements when merging generated payload paths
2026-08-11 22:12:45 +02:00
Andrey VasnetsovandClaude Sonnet 5 ca20151659 Add edge-tool: CLI for creating, seeding, optimizing, and uploading local edge collections (#10159)
* Add edge-tool: CLI for creating, seeding, optimizing, and uploading local edge collections

Mirrors the style of lib/edge/tools/shard_update and shard_query: `create` builds a
minimal EdgeShard on disk (dense/sparse vectors, quantization presets including
turbo4, payload indexes, target segment count), `upsert` seeds it with random points
matching its live schema, `optimize` runs the shard optimizers, and `upload` pushes
the resulting directory to S3/GCS. Useful for quickly spinning up test collections
without a running Qdrant server, then promoting them to object storage.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* edge-tool: initialize feature flags, enable serverless_compatible, fix --sparse ambiguity

Initialize the global feature-flag OnceLock at startup (with serverless_compatible
set, cascading write_segment_manifest/append_only_mutations/compact_bitmask/
append_only_storages) so runs no longer spam "Feature flags not initialized!" and
collections are created in the serverless-compatible format.

Also splits --sparse into a plain boolean flag plus a repeatable --sparse-name:
clap's optional-value parsing for the old `--sparse [NAME]` form silently
swallowed a following positional PATH as the sparse vector's name whenever
--sparse was the last flag before it (e.g. `create --dense 1024 --sparse ./col`).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* edge-tool: fix --sparse=NAME to require_equals instead of a separate flag

The --sparse/--sparse-name split from the previous commit lost the ability to
name a sparse vector with --sparse itself. Restore a single --sparse[=NAME]
flag, but with require_equals(true): clap then only binds a value via
--sparse=NAME, never via a following bare token, so it stays safe next to the
trailing PATH positional in every position (bare --sparse, --sparse=NAME, or
multiple --sparse=NAME occurrences) without reintroducing the ambiguity that
made --sparse swallow PATH as its value.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* edge-tool: remove --segments from create, it has no effect there

EdgeOptimizersConfig::default_segment_number only feeds MergeOptimizer as a
merge-down ceiling (reduce segment count when it exceeds the target); unlike
the main collection's LocalShard::build_local, EdgeShard::new never loops to
pre-create N appendable segments. A freshly created collection always starts
at exactly 1 segment, so passing --segments to `create` was silently a no-op.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* edge-tool: add --indexing-threshold-kb to create

Unlike --segments (removed previously), the indexing threshold is a parameter
IndexingOptimizer actually consults on every optimize() run: segments larger
than it get an HNSW index built. Verified end-to-end (create with a 1KB
threshold, upsert 2000 points, optimize) that it produces an hnsw-indexed
segment where it would otherwise stay plain.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* edge-tool: add --clean to upload, wiping the destination prefix first

Lists every object under DESTINATION and deletes it via ObjectStore::delete_stream
before uploading, so re-uploading a collection recreated with a different shape
(different segment UUIDs) doesn't leave the old segment's files behind.

Verified against the local S3 proxy: uploaded one collection, then a second,
differently-shaped one to the same prefix without --clean (29 objects, stale
leftovers from the first); re-uploading the second with --clean correctly
dropped it back to exactly its own 19 files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-11 15:54:27 +02:00
Andrey VasnetsovandClaude Sonnet 5 309c64945a [UpdateOnly] wire every component into the appendable segment (#10152)
* [UpdateOnly] wire every component into the appendable segment

`AppendableSegment::store_points` sheds its `todo!()`: the id tracker claims a
fresh slot per point, every component writes its data at those slots — each
named vector storage, the payload storage, the payload indexes — and only then
do the versions cover them, the step that makes the points visible to readers.
A crash anywhere in between leaves claimed, unpublished slots, which the next
writer to open the segment retires. Each vector comes from whichever half of
`FullyQualifiedPoint` holds it: the batch's decoded vectors win over the bytes
carried from the point's previous slot, and a name in neither still takes its
slot as a vector the point does not have.

The store components open lazily, on the first `store_points`. A batch that
only deletes writes nothing but the mappings log, so it never pays for those
opens — and it keeps working against segments whose payload storage was created
in mutable mode, which the append-only writers refuse and which is all any
leader builds today.

The writer now also remembers what it stored, so `tombstone_points` skips a
point this very batch wrote instead of retiring its fresh slot; the caller can
hand over every slot a stored point used to occupy without holding that rule.

`UpdateOnlySegmentEnum::open` takes the segment config, which is where the
writer learns which vector storages exist.

The end-to-end edge tests now run stores the whole way through: located and
resolved through the `LookupSegment`s, appended by the writer, and read back
through an ordinary follower — a new point with its payload, a rewrite winning
over the old copy, a replayed batch skipping on the published versions, and a
second writer resuming every component where the first ended. The leader still
writes its payload storage in mutable mode, so the tests recreate it empty in
append-only mode, standing in for segment creation wiring that does not exist
yet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [UpdateOnly] drop the stored-ids guard from `tombstone_points`

The caller already never asks to retire a point its batch stored — it has to
hold that rule regardless, since `preview` mirrors it to count outcomes — so
the writer-side set was redundant state, and it made `tombstone_points`
silently drop requests instead of honoring a stated contract. The contract is
now stated: only points the batch deletes go here, because a delete addresses
the external id and would take a stored point's fresh slot along with the
stale one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [UpdateOnly] gate the store tests off Windows

The leader's writable storage preallocates chunk files, and the append-only
writer cuts them back to end at the data — its append offset is a
compare-and-swap token, so a file longer than the data would make every append
conflict. That cut replaces the file, which Windows refuses while the writer's
own `LookupSegment`s hold it memory-mapped; on Linux the old inode simply lives
on under the mappings. Nothing to fix in the writer: Windows cannot shrink a
mapped file, and the production target is object storage, where neither
preallocation nor mmap exists.

The delete tests keep running everywhere; the store tests move into a
`#[cfg(not(windows))]` module together with the imports and helpers only they
use, so the Windows build carries no unused-import warnings. Cross-checked with
`--target x86_64-pc-windows-msvc`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [UpdateOnly] wire the quantized overlay into StoreComponents

Opens UpdateOnlyQuantizedVectors alongside each dense, non-multivector,
non-Turbo4-datatype vector's raw storage, when the segment's quantization
config supports incremental appends (Binary/Turbo). Multivector and Turbo4
combinations are out of scope (see UpdateOnlyQuantizedVectors' own doc
comment) — such a vector simply has no quantized overlay entry and stays
searchable exactly through its raw storage alone, same as before.

store_points keeps the overlay's row count in exact lockstep with the raw
storage: every point takes a row in both, in the same order, at the same
id (start_slot + offset) — a decoded vector encoded for real, a
Raw-bytes-carryover blob decoded back to f32 per its actual storage
datatype (mirroring QuantizedVectors::create_impl's use of
PrimitiveVectorElement::quantization_preprocess for the same purpose on
the non-update-only path), and a Missing vector as an all-zero placeholder.
Skipping a row for the latter two cases would silently misalign every
later quantized lookup — scoring one point's vector against another's
quantized copy — so this mirrors the raw storage's own "every point takes
its slot" rule exactly rather than only handling the common decoded case.

UpdateOnlyQuantizedVectors now retains its resolved QuantizedVectorsConfig
(exposed via quantization_config()/dim()) rather than discarding it after
opening storage, since a reopened overlay's persisted config is the source
of truth for how to decode carried-over bytes — not necessarily identical
to whatever live config the caller has to hand. Its now-unused flusher()
is dropped: like every other update-only storage in this stack, a write is
already durable when append_many/upsert_vector returns.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 11:23:01 +02:00
Andrey VasnetsovandClaude Opus 5 1d18e46551 [UpdateOnly] split UpdateOnlySegment into its lookup and writer phases (#10142)
* [UpdateOnly] split UpdateOnlySegment into its lookup and writer phases

Applying a batch runs in two phases that agree on almost nothing, and
`UpdateOnlySegment` was both: `resolve.rs` used every field, `append.rs`
used none of them and could not — a `ReadOnlyPayloadStorage` has no append
path. The `fs` field existed only for the writes that were never wired up.

Split along that line:

* `LookupSegment` (was `UpdateOnlySegment`) is the read phase. Every segment
  of a shard is opened as one, on read-only bounds, and the phase above them
  aggregates. Loses the dead `fs` field.
* `DeleteOnlySegment` and `AppendableSegment` are the write phase, one
  segment each, `UpdateOnlySegmentEnum` over the two. Opened for one batch
  and dropped with it, matching the append-only components, which buffer
  nothing across calls.

The phases meet at `SegmentWriterState`, produced by
`LookupSegment::writer_state` and consumed by `UpdateOnlySegmentEnum::open`.
It carries the mappings-log tail an appendable writer resumes from, which
`UpdateOnlyAppendableIdTracker::new` requires to come from one and the same
read of that log. The writer kind follows the id-tracker format that was
loaded, not the segment config: the format decides how a point is retired.

That difference makes `tombstone_points` take both ids, `(external, slot)`;
an immutable segment marks the slot in its deleted-points bitmask, an
appendable one records a retirement for the id in its mappings log. The
appendable half is implemented — deletes now run end-to-end. `store_points`
and the immutable bitmask remain `todo!()`, still waiting on the append-only
storages and field indexes.

Two bugs surfaced while wiring it up:

* A point stored into the write target must not have its old slot retired
  there: appending records a mapping that supersedes it, and retiring the id
  on top would take the new slot with it.
* A second `apply_batch` through one writer resurrected deleted points. It
  resumed the log from the `mappings_end` its own first batch had moved past,
  and appending there cut that batch's entries off. Refused now; lifting it
  means reloading the segments after a batch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [UpdateOnly] one batch per writer, enforced by the type system

Cleanup pass over the phase split.

`apply_batch` now takes `self`. It could only ever serve one batch — the
segments are read when the writer opens, and that read is both what a batch
resolves against and what its writers resume from — and the runtime guard
enforcing that cost a flag, its doc, two imports, a hand-maintained
`writes_anything` condition, an error branch and a test. Consuming the writer
makes the second call a compile error instead.

Also:

* drop `LookupSegment::uuid`, which nothing ever read, along with the two
  parameters and the argument that fed it;
* `AppendableSegment::tombstone_points` was a copy of the tracker's own
  `retire_pending_inserts`; both now go through `delete_points`;
* fold the duplicated "segment disappeared mid-batch" error into
  `LookupSegmentHolder::get`, and restore `write_target_uuid` as an `Option`,
  which is what two of its three callers wanted;
* one fixture helper for the writer tests instead of three copies;
* state the mappings-log co-read invariant once, with pointers, instead of
  three times.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [UpdateOnly] cut the writer surface down to what it does

* `flush()` is gone from both writers and the enum. Both bodies were `Ok(())`
  and would stay that way: the id tracker persists what it writes before
  returning, and the deleted-points bitmask writer does not exist yet. The
  ordering it looked like it enforced — new slots durable before the tombstones
  retiring the old ones — falls out of call order, since every write is durable
  when it returns. Bring it back with the first storage that buffers.
* `SegmentWriterState` was an enum of one unit variant and one payload, which
  is `Option`. `writer_state()` returns `Option<AppendableIdTrackerState>`, and
  `None` reads as what it means: no mappings log to resume, so a delete-only
  writer.
* `LookupVectorData` wrapped a single `Arc<AtomicRefCell<_>>`; the map holds it
  directly now.
* `appendable` joins the five `pub` fields around it, and `is_appendable()`
  goes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 00:42:54 +02:00
Till Bungert 2ef887bae1 fix(edge): skip pool install for single search thread (#10118) 2026-08-07 12:28:28 +02:00
7364cc42ef feat(edge): add query_batch for batched planned queries (#10100)
* feat(edge): add query_batch for batched planned queries

Expose the planned-query batch path as a public API so multiple
independent queries can share one planning pass over leaf searches
and scrolls. Wired through EdgeShardRead, FFI, and Python bindings.

Co-authored-by: Cursor <cursoragent@cursor.com>

* perf(edge): push batched query vectors down to segments

`query_batch` planned the whole batch at once but then executed every leaf
search on its own: one query context, one fan-out over all segments, and one
single-vector `Segment::search_batch` call per leaf.

Execute the batch as a batch instead:

- `EdgeReadView::search_batch` builds the query context once, visits the
  segments once, and hands each segment the leaves that agree on everything
  but their query vector as a single multi-vector `search_batch` call.
  `search` is now a thin wrapper over a one-element batch.
- Move `SearchType`/`BatchSearchParams` from `collection`'s segments searcher
  into `shard`, next to `CoreSearchRequest`, and add `group_search_batches`
  so both the collection and the edge read path share one grouping
  implementation. Edge computes the grouping once and reuses it per segment.
- `search_matrix` now issues its per-sample nearest queries through
  `query_batch`; they share filter, limit and vector name, so the whole
  sample is scored in one batched search per segment instead of one full
  segment pass per sampled point.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 11:59:40 +02:00
Srimon Danguria 8db152b270 Make tonic optional in shard and clarify feature dependencies in Cargo.toml (#10069) 2026-08-05 15:37:36 +02:00
Andrey VasnetsovandClaude Opus 5 21db2f3ff9 Bump edge packages (Python + Rust + FFI) to 0.8.0 (#10098)
Minor version bump of the edge packages from `0.7.2` to `0.8.0`.

- `lib/edge/python/Cargo.toml`: `qdrant-edge-py` 0.7.2 -> 0.8.0
- `lib/edge/publish/amalgamate.py`: `VERSION` constant bumped
  (`qdrant-edge` on crates.io)
- `lib/edge/ffi/Cargo.toml`: `qdrant-edge-ffi` bumped, kept in sync with the
  other two as its header comment requires
- `lib/edge/publish/ast-grep-rules.yaml`: inline package version comment
  updated (it was stale at 0.7.1)
- `Cargo.lock`: regenerated version entries

Follows the same pattern as #9252.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 14:55:05 +02:00
dependabot[bot] 0b357b58c7 build(deps): bump syn from 2.0.118 to 3.0.2 (#10071)
Bumps [syn](https://github.com/dtolnay/syn) from 2.0.118 to 3.0.2.
- [Release notes](https://github.com/dtolnay/syn/releases)
- [Commits](https://github.com/dtolnay/syn/compare/2.0.118...3.0.2)

---
updated-dependencies:
- dependency-name: syn
  dependency-version: 3.0.2
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-04 09:53:01 +02:00
9f3c07b00d feat: UpdateOnlySegment / UpdateOnlyEdgeShard batch writer skeleton (#10021)
* feat: `UpdateOnlySegment` / `UpdateOnlyEdgeShard` batch writer skeleton

Mirror image of the read-only pair, for the serverless updater: a
shard/segment whose public surface is writes only, built for batches of
many tiny operations against remote, append-only storage.

Implemented:

* `UpdateOnlySegment<S>` with a deliberately narrow open — id tracker,
  payload storage and one storage per named vector, all cold. No vector
  index, no quantized vectors, no payload index on the segments the
  writer only reads from.
* `SegmentUpdateView`, the shared home of resolution logic, generic over
  the component traits (`VectorDataStorageRead` is a `VectorDataRead`
  without the index, so a segment that opens no index can produce the
  view). Batched `locate_points` / `point_versions` /
  `read_stored_points`.
* `UpdateOnlyEdgeShard<S>::apply_batch`: fold the batch to one entry per
  point, locate the points, read only the ones that cannot be resolved
  from the batch alone, materialize `FullyQualifiedPoint`s, append them
  and tombstone the slots they replace.

`todo!()`, pending the append-only components on the roadmap (appendable
`DynamicStoredFlags` and `ChunkedVectors`, an appendable payload
blobstore and field indexes): `store_points`, `tombstone_points`,
`flush`, and creating the first appendable segment.

Filter-selected operations, point sync, conditional upserts and the
schema-level operations are rejected up front rather than silently
skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: codespell implementor → implementer in SegmentUpdateView docs

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs: trim update-only writer docstrings to guarantees

Less verbose throughout: state each function's contract — ordering,
absent-value behavior, preconditions, durability — and drop narration
about where types are used or why alternatives were rejected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor: fold SegmentUpdateView into UpdateOnlySegment as inherent methods

The view was premature: it had exactly one producer, and its trait
bounds bought an unexercised option. Resolution (locate / versions /
read raw) now lives as inherent methods on UpdateOnlySegment, still
generic over the backend. A shared view can be extracted when a second
producer appears, e.g. batched CoW moves out of regular segments.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor: split edge update_only batch.rs into a module

Pure move: mutation.rs (PointMutation fold + materialize), plan.rs
(UpdateBatchPlan operation intake), tests.rs. PointUpdates::new/push
narrowed to pub(super).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: parallel per-segment batch reads + tombstone every copy of a point

locate_points and read_stored_points visit segments in parallel on a
dedicated edge-update rayon pool (build_search_pool generalized to
build_segment_pool with a thread-name prefix).

locate_points now keeps every slot a point occupies, not just the
newest copy: a rewrite or delete retires all of them. Tombstoning only
the newest slot would let an older duplicate left by an interrupted
move outlive the point — and resurrect it after a delete.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor: point_versions returns a map keyed by internal id

The id tracker's batch read is keyed by internal id already; returning
AHashMap drops the positions_of reverse-lookup adapter. Absent key =
unwritten slot, defaulted to version 0 at the caller.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: accept a deferred threshold when opening UpdateOnlySegment

Groundwork for an external rebuilder working the same directory: the
cutoff loads slots at or above it into the appendable id tracker's
deferred track (same appendable-only filter as ReadOnlySegment). It
hides nothing from the writer — resolution runs WithDeferred, so every
point still locates at its latest slot. The edge shard passes None
until the rebuilder coordination exists.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: preview_batch — resolve a batch without writing anything

apply_batch and the new preview_batch share one resolution stage
(resolve_batch: locate, read, materialize into per-point PointActions),
so a dry-run reports exactly what an apply would do. Plus
segment_configs(): per-segment configs with the write target marked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: prefetched + parallel segment opens for the update-only writer

UpdateOnlySegment::open now mirrors ReadOnlySegment::open: a
per-segment CachedFs primed by preopen, config parsed once and handed
to open_via. The edge shard opens segments in parallel on its pool,
keeping fail-hard semantics. With Populate::No throughout, prefetches
transfer no data-file content — only configs, the id tracker and the
deleted flags, whose opens consume them whole anyway.

Also: ReadOnlyAppendableIdTracker::preopen now tolerates the
not-yet-created mappings/versions files of an empty appendable segment,
matching its open's contract — previously unreachable because followers
skip appendable segments on error, while the writer must open them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: edge-shard-update — dry-run batch upserts against a shard

Counterpart of edge-shard-query for the write path: opens an
UpdateOnlyEdgeShard over a local directory or S3/GCS object storage,
generates random points shaped by the shard's own schema (segment
config + payload-index schema), and logs what applying them would do —
locations, versions, actions, tombstones — via preview_batch. Nothing
is written: the write half is still todo!().

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: box PointAction::Store to appease clippy::large_enum_variant

A resolved point is ~384 bytes while every other variant is empty.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: adapt to dev's dead-code sweep (#10030)

Restore NamedVectors::remove_ref — removed as dead on dev, but the
batch fold's DeleteVectors arm is now its first caller. Drop the
allow(dead_code) on segment::update_only (no longer needed) and switch
the writer's unread fs field to expect(dead_code), per the new
ast-grep rule.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: root <111755117+qdrant-cloud-bot@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-31 09:10:42 +02:00
xzfc 75385df69f Remove dead code (#10030)
* Remove dead code

* Remove unused dependencies

* `allow(dead_code)` -> `expect(dead_code)`

* ast-grep: rule-tests/*-test.yml => tests/*-test.yml

For brevity.

* ast-grep: forbid allow(dead_code)
2026-07-30 20:59:31 +00:00
Daniel Boros 714b61e9a9 feat: pin per-shard search pool to a core (#10029)
* feat: pin per-shard search pool to a core

* fix: plumb search_pool_core through bindings

* feat: expose search_pool_core in python bindings

* fix: validate search pool core before pinning

* chore: trim comments
2026-07-30 20:41:48 +02:00
dependabot[bot] c16763e0f5 build(deps): bump uniffi from 0.31.2 to 0.32.0 (#10001)
Bumps [uniffi](https://github.com/mozilla/uniffi-rs) from 0.31.2 to 0.32.0.
- [Changelog](https://github.com/mozilla/uniffi-rs/blob/main/CHANGELOG.md)
- [Commits](https://github.com/mozilla/uniffi-rs/compare/v0.31.2...v0.32.0)

---
updated-dependencies:
- dependency-name: uniffi
  dependency-version: 0.32.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-28 11:41:58 +02:00
Sasha Denisov 94b0ec1ab6 test(edge-ffi): assert Match.Any/Except land in the correct AnyVariants arm (#9991)
The AnyVariants redesign's tests only checked `is_ok()`. Strengthen them to
pin the actual contract now that the strings-XOR-integers constraint is
type-enforced:

- destructure the converted `SegmentMatch` and assert the values land in the
  matching engine `AnyVariants` arm, in order (catches a mis-wired `From`);
- cover the previously-missing `Except` x `Strings` corner;
- assert `into_iter().collect()` into the engine's `IndexSet` dedups
  (first-seen order), matching the engine's own construction;
- pin empty-set validity for both `Any` (matches nothing) and `Except`
  (matches everything) — the one semantic the docstring promises and the old
  XOR guard used to reject.

Uses `let`-else rather than a wildcard match arm to satisfy the crate's
`wildcard_enum_match_arm` lint on engine enums. 3 tests -> 7; suite 122 -> 126.
2026-07-28 08:56:15 +02:00
Sasha Denisov 4469e9b2f6 refactor(edge-ffi): model Match.Any/Except as a typed AnyVariants one-of (#9988)
Match::Any/Except took `{ strings: Option<Vec<String>>, integers: Option<Vec<i64>> }`
with a runtime XOR check — two of the four representable states were invalid
(both-none, both-set -> InvalidArgument), and the prior field defaults made the
all-invalid empty call the easiest thing to type in Kotlin/Swift.

Replace with a typed sum type `AnyVariants { Strings | Integers }`, matching the
engine's own `AnyVariants`, the gRPC `oneof`, the crate's existing
`ValueVariants`, and every official Qdrant client (Python union, Java/Go/Rust
typed constructors, JS `string[] | number[]`). "Exactly one" is now enforced at
compile time; the two InvalidArgument paths and the field defaults are gone.
Empty sets remain valid (Any -> matches nothing, Except -> matches everything),
mirroring the engine.

122 crate tests pass; clippy clean; regenerated bindings expose
`Match.Any(AnyVariants.Strings(...))`.
2026-07-27 15:19:57 +02:00
Sasha Denisov 2498bfc6bf fix(edge-ffi): default every Option field to None so optionals are skippable (#9987)
`#[uniffi(default = None)]` was applied to request Record fields but not to
enum-variant fields (Query / Match / Fusion, payload-index params) or the
response Records — an oversight, not a UniFFI limitation: uniffi 0.31 fully
supports defaults on enum-variant fields (verified end-to-end by regenerating
the bindings and confirming `= null` lands, e.g. `Query.Nearest.using`).

Annotate every remaining Option<T> field (enum variants + Records) so the
generated Kotlin/Swift bindings default them to null/nil. Consumers can now omit
any optional field — e.g. `Query.Nearest(vector)` instead of
`Query.Nearest(vector, using = null)` — and adding a new optional field to a
Record or enum variant stays source-compatible (named-argument callers).

Constructor/method arguments already use
`#[uniffi::constructor/method(default(...))]` and are unchanged.

Verified: 123 crate tests pass; regenerated bindings have zero optional fields
without a default.
2026-07-27 13:19:07 +02:00
be543561e5 Add Logstore and Blobstore wrapper (#9673)
* Gridstore: introduce storage operating mode in config

Add a mode field to the gridstore config, selecting between the dynamic
mode (current behavior, the default) and the upcoming serverless mode.
The mode is specified through StorageOptions on creation, persisted in
config.json, and read back first when opening so the correct variant can
be selected automatically. Configs written before this field existed
deserialize as dynamic.

For now, selecting the serverless mode returns an error; the variant
itself is added in follow-up commits.

* Gridstore: move dynamic implementation into dedicated module

Mechanical move of the current Gridstore implementation into
gridstore/dynamic.rs as DynamicGridstore. The public Gridstore struct
becomes a thin wrapper holding a mode variant enum, propagating every
call into the selected variant. For now the enum only has the dynamic
variant; the serverless variant is added in follow-up commits.

No logic changes to the dynamic implementation itself: only visibility,
the config parameter now passed into open (the wrapper reads it first to
select the mode), and open_or_create staying on the wrapper.

* Gridstore: add serverless tracker

Add the append-only mapping tracker for the serverless storage mode.

The tracker file is a plain array of 16-byte mapping entries without any
header: the number of mappings is defined by the exact file length, and
the entry index is the point offset. The file starts empty and only ever
grows by appending, existing bytes are never rewritten. Mappings must be
set in monotonically increasing point offset order; skipped offsets are
backfilled as zeroed entries which decode as None.

New mappings are buffered in memory and appended with a single write per
flush. A flush with a stale target is a no-op so bytes are never written
twice. A torn trailing entry (file length not a multiple of the entry
size) is ignored when reading and truncated away when opening writable.

Unlike the dynamic tracker, the file is read and written directly with
positional file IO instead of memory mapping, as serverless environments
do not handle memory mapped files well.

* Gridstore: add serverless storage variant

Add the append-only gridstore variant for serverless deployments, which
restrict IO to appending to files: existing bytes can never be
rewritten, and IO is expensive so as few files as possible are used.

The variant stores all value data in a single page file next to the
serverless tracker and the storage config, three files in total. Both
data files start empty and only ever grow by appending; there is no
preallocation, no used-block bitmask and no gap/region bookkeeping.
Values are appended at put time at the next block aligned offset, with
the zero padding included in the write so it lands exactly at the end
of the file. Mappings are buffered and appended to the tracker with a
single write per flush, after the page file is synced, so a mapping on
disk never points at data that is not durable.

Values cannot be updated or deleted, and must be put at monotonically
increasing point offsets; violations are rejected before any data is
written. Files are read and written directly, never memory mapped.

The mode is selected through StorageOptions on creation and picked up
automatically from the persisted config when opening.

* Gridstore: serverless support in reader and view

Extend the read-only GridstoreReader and the GridstoreView with the
serverless mode, keeping both public types unchanged: like the writable
Gridstore they now hold a mode variant internally, selected
automatically from the persisted config when opening.

The serverless reader holds the tracker and page directly and reads the
files positionally, without memory mapping. A live reload re-reads the
mapping count from the exact tracker file length (there is no size
header), ignoring a torn trailing entry, and never truncates as it is
read-only. Value reads always go directly to the file, so newly
appended data is readable without remapping anything.

* Gridstore: document storage operating modes

* Gridstore: review fixes for the serverless mode

Hardening and cleanup from a review pass over the new serverless
storage variant:

- Batch the reader side iteration like the writer already did, instead
  of materializing tracker mappings for the full range in one go, which
  could transiently allocate gigabytes on large storages.
- Recover the append cursors when a positional write fails partway:
  truncate the file back to the tracked length so a retried append or
  flush never rewrites bytes that already landed in the file.
- Validate page addressability before appending value data, a rejected
  put must not grow the page file.
- Cross-check tracker and page consistency when opening: mappings that
  reference value data past the end of the page file (e.g. after a
  partial copy or restore) now fail fast instead of surfacing as
  opaque read errors per point.
- Reject value pointers into any page other than page 0 on the
  serverless read path with PageNotFound, matching the dynamic mode
  contract, instead of silently reading from a wrong location.
- Refresh the reported storage size on reader live reload even when no
  new mappings were flushed, unflushed value data may have grown the
  page file already.
- Validate configs read from disk: a corrupt config with zero sized
  blocks, pages or regions is now rejected when opening instead of
  panicking on a division by zero later.
- Classify rejected serverless puts as UnsupportedOperation, consistent
  with rejected deletes, so they don't surface as user-facing
  validation errors at the segment level.
- Deduplicate the compression dispatch into Compression::compress and
  Compression::decompress, and the serverless file create/open patterns
  into shared direct IO helpers, so the two modes and files can't
  silently drift apart.

* Gridstore: cover both operating modes in mode-agnostic tests

Parameterize the gridstore tests that exercise mode-agnostic behavior
over both the dynamic and serverless mode with rstest, using a
single and bulk put/get roundtrips, storage files, basic persistence,
corrupt config rejection, batched read congruence, reader live reload,
and the different block sizes.

Mode specific expectations branch inside the tests: expected file
names, storage size semantics (whole blocks vs exactly packed bytes),
value pointer layout (page spill over vs a single packed page), and
gaps (created by deletes in dynamic mode, by skipped puts in serverless
mode). Dynamic-only internals assertions are kept behind a mode check.

Tests around updates, deletes, page spanning, block reuse and other
dynamic-only behavior intentionally stay dynamic; the serverless
specific format invariants remain covered by the dedicated serverless
tests.

* Gridstore: port serverless specific tests from sibling branch

Source the serverless specific test cases that the
serverless-gridstore-updates branch added, adapted to the dedicated
variant implemented here (distinct file names, headerless tracker
with 16 byte entries, a single packed page without trailing padding,
and rejected re-puts):

- writes only ever append: tracker and page files only grow and
  previously written bytes stay byte-for-byte untouched
- new mappings land exactly at the end of the tracker file, which
  always covers the exact number of mappings
- mapping gaps are zero-padded on disk and survive reopening
- values are packed back to back at block aligned offsets, the page
  file ends exactly at the last value
- serverless mode never creates nor reports block flag files
- a flusher persists exactly the mappings that existed at its
  creation, later puts stay pending
- a config claiming the wrong mode fails loudly in both directions
  instead of loading the incompatible file format of the other mode

Tests around their mode switching, page spanning and tolerated deletes
don't apply to this design and are intentionally not ported.

* Gridstore: test serverless production risk scenarios

Add tests for the operational aspects that matter before serverless
mode goes to production, each covering a scenario that wasn't
evaluated yet:

- Replayed puts of already persisted offsets (a WAL redo after a
  crash where the flush completed but was never acknowledged) are
  rejected without appending anything, and max_point_offset is the
  exact offset a replay must resume at.
- The accepted crash case of a tracker file extended with zeroed
  bytes: the entries count as permanent None mappings, can never be
  put again, and the storage stays consistent and writable past them.
- The read-only reader never modifies the files: opening over a torn
  tracker tail, reading, iterating and live reloading leave both
  files byte-for-byte untouched.
- A multi-round put/flush/reopen cycle always exposes exactly the
  flushed prefix, with the mapping count matching the exact tracker
  file length and unflushed offsets reusable.
- An append beyond the maximum addressable block offset is rejected
  before writing anything, keeping retried puts from growing the page
  file unboundedly.

* Gridstore: rename serverless mode to append-only, split into module

Rename the mode after its defining characteristic instead of its
deployment target: files only ever grow, existing bytes are never
rewritten. Renames Mode::Serverless to Mode::AppendOnly (persisted as
"mode": "append_only") and the on-disk file names to
append_only_tracker.dat and append_only_page_0.dat. The serverless
deployment motivation stays in the documentation.

Also split the single 2300 line serverless.rs into an append_only
module with dedicated files for the storage, page, view, reader and
tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Use universal IO in Gridstore

* Include upstream preopen logic in new Gridstore variant

* Gridstore: buffer append-only value writes until flush

In append-only mode, put previously wrote the value data to the page
file right away, one write operation per put, while mappings were
already buffered and batch persisted on flush. Buffer value writes the
same way: both the value and its mapping now only land on disk once a
flush cycle executes.

This batches all new value data into a single write operation per
flush, which is significantly more efficient on S3 based storage where
every write is a costly operation. A flush now performs exactly two
writes: one appending all buffered value data to the page file, one
appending all pending mappings to the tracker file, in that order, so
a mapping on disk never points at value data that is not durable.

The page mirrors the tracker's pending mechanism: an in-memory buffer
that is byte for byte the next append (zero padding between block
aligned values included), a watermark captured at flusher creation so
puts made during a flush stay buffered, a stale-flush no-op guard so
appended bytes are never written twice, and truncate-back recovery on
failed writes. Reads transparently serve buffered values from memory.

As a side effect, a crash between flushes now leaves nothing on disk
at all, where the write-through approach left orphaned value bytes in
the page file. The buffered data is held in memory until the next
flush, bounded by the flush cadence.

Universal IO filesystem handles are now required to be Send + Sync, so
the flusher closure can carry one to grow the page file at flush time;
all existing backends already satisfied this.

* Gridstore: rename inner DynamicGridstore to Gridstore

The dynamic variant keeps the Gridstore name; the outer dispatching type
will be renamed to Blobstore in a follow-up. Until then the inner type is
referred to as dynamic::Gridstore to distinguish it from the outer type.

* Gridstore: rename append-only variant to Arenastore

The append-only variant stores all value data in a single ever-growing
page, allocating space by appending, hence: arena store.

* Gridstore: rename outer storage type to Blobstore

The outer type dispatching between the two storage variants is now called
Blobstore, being more generic than Gridstore. This frees up the Gridstore
name, which now exclusively refers to the dynamic mode variant, next to
Arenastore for the append-only variant. Storage components keep using the
outer type, so they now use Blobstore.

The gridstore crate name, GridstoreError, and the persisted names
(config.json mode, payload config storage_type) are unchanged.

* Gridstore: split Gridstore and Arenastore into dedicated modules

The outer module is now blobstore, matching the Blobstore type it
defines. The two storage variants each get their own submodule: the
dynamic Gridstore moves from dynamic.rs into gridstore/ with its reader
and view extracted from the shared files, mirroring the arenastore/
module (previously append_only/) which already had this layout.

* Rename gridstore crate to blobstore

The crate is named after the outer Blobstore storage type it provides.
The gridstore name lives on in the dynamic mode variant. GridstoreError
and the persisted names (config.json mode, payload config storage_type)
are unchanged.

* Arenastore: pack values back to back across multiple pages

Drop the block alignment from the append-only mode: values are packed
byte to byte, without blocks, and the tracker offset is now a plain byte
offset within the page. Blocks and regions are dynamic mode concepts;
their page size constraints no longer apply to append-only configs.

Bring back support for multiple pages. Once appending a value would
grow the current page beyond the configured page size, a new page is
started, bounding the size of and the number of appends to each file:
object stores like S3 Express limit the number of appends per object.
A value larger than the page size gets a page of its own; values never
span pages.

A rollover creates the new, empty page file at put time; the value data
itself stays buffered until the next flush, which appends to each
touched page with a single write, using per-page watermarks captured at
flusher creation. The reader scans for consecutively numbered page
files when opening, validates the most recent mappings against them,
and adopts pages created since on a live reload.

* Blobstore: rename dynamic mode to mutable

Rename Mode::Dynamic to Mode::Mutable, and the persisted config value
with it: config.json now writes "mode": "mutable". There is no
compatibility alias for "dynamic", released versions never wrote the
mode field (a missing field still defaults to mutable), only unreleased
storages did.

The Gridstore type and module names for the mutable variant are
unchanged.

* Fix Edge compilation due to package rename

* Review remarks

* Extract Gridstore preopen into module

* Rename Arenastore files

* Use universal IO for append operations

* Rename GridstoreError to BlobstoreError

The error type belongs to the Blobstore crate and is shared by both the
Gridstore and Arenastore variants, so it follows the crate naming. Also
update the user-facing error messages that referred to the old name.

* Split config into per-variant types

* Rename Arenastore to Logstore

Rename the Arenastore type to Logstore, including the reader, view,
config, module and variant names. The storage file names follow:
log_page_{n}.dat and log_tracker.dat. The persisted mode tag stays
"append_only".

* Move bitmask module into the Gridstore variant

The bitmask tracks free blocks, which only exists in the mutable mode.
Move the module from the crate root into the Gridstore variant that
owns it. It stays re-exported at the crate root because the bitmask
benchmark needs a public path.

* Move pages module into the Gridstore variant

Like the bitmask, the block based pages module is only used by the
mutable mode. Move it from the crate root into the Gridstore variant
that owns it. The Logstore variant has its own page implementation.

* Use universal IO for every Logstore operation

Replace the direct_io module with universal IO in the append-only
tracker, making the whole Logstore go through a universal IO backend
bounded by UniversalRead and UniversalAppend:

- The tracker is generic over the backend now. Reads go through
  UniversalRead with the caller's access pattern, flushes land as one
  atomic append with the same offset compare-and-swap recovery as the
  pages: a retried append after a lost acknowledgement is adopted
  instead of appended twice. A torn trailing entry is still truncated
  away on writable open, through a fresh handle since shrinking is not
  supported through an open one.
- The reader now schedules a prefetch for the tracker file too, it no
  longer bypasses the backend.
- The config write, clear and wipe use the backend file operations
  instead of local filesystem calls, matching the Gridstore variant.

* Batch reads in Logstore read_values

Apply the same batching logic as the Gridstore variant: resolve all
mappings first, then fetch the value data, both through the backend's
read pipeline so async backends can serve the reads in parallel.

The tracker gains a batched lookup mirroring the mutable tracker's
iter, serving pending mappings and out of range point offsets directly
from memory. The pages gain a batched value read; unflushed values are
served from the in-memory buffers, and since values never span pages
each value is a single read without reassembly.

Like in the Gridstore variant, the callback may now be invoked in a
different order than the requested point offsets.

* Better describe logstore live reload ordering

* use enum for options, swap `*Options`<->`*Config` naming

* don't wrap enum in struct

* ditch unused `StorageConfig`, make deserialization more ergonomic

* rename `*Options`->`*Config`

* make `preopen` non-blocking

* fixup! ditch unused `StorageConfig`, make deserialization more ergonomic

* fixup! use enum for options, swap `*Options`<->`*Config` naming

* fixup! don't wrap enum in struct

* fix rebase

* use `populate` param in Logstore

* test: failing repro of stale page after live reload across rollover

A reader that live-reloads between a page rollover and the following
flush adopts the new, still empty page. The previous page is then no
longer the last one and is never reloaded again, so the tail that the
next flush appends to it stays invisible to the reader forever:

    value pointer at byte 100 with length 100 is out of range

AppendOnlyPages::live_reload only reloads the last held page, assuming
earlier pages never change once a newer page exists. But the rollover
creates the new page file eagerly at put time, while the previous
page's buffered tail only lands at the next flush (see
test_rollover_writes_no_value_data_before_flush), so a page can keep
growing on disk after its successor exists.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: reload all pages that grew

* use Fs in `open_or_create`

* fix: publish tracker mappings only after the pages reload

`AppendOnlyTracker::live_reload` observed the mapping count and made it
visible in one step, before `LogstoreReader::live_reload` reloaded the
pages. Every failure path in the page reload -- `list_files`, reopening a
grown page, opening an adopted one, the truncation check -- therefore left
the reader with mappings referencing value data it never loaded, so reads
in the new offset range fail until a later reload happens to succeed. The
edge refresh loop keeps a segment whose reload failed, expecting it to keep
serving its pre-refresh state, which it then does not.

Split observing from publishing: `reload_count` refreshes the handle and
returns the count as a `PendingReload` token, `commit_reload` publishes it.
The reader still observes the tracker first, as the writer persists pages
before the mappings referencing them, but only commits once the pages are
loaded. Reopening without committing is harmless: reads stay bounded by the
unchanged count, and the bytes below it never change.

A partial failure inside the page reload needs no unwinding, pages running
ahead of the tracker is the safe direction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* perf: batch the value reads in Logstore iteration

`LogstoreView::iter_range`, the path behind `Logstore::iter` and
`LogstoreReader::iter`, fetched the mappings for the whole range with a
single read but then read the values themselves one at a time, serially.
Gridstore routes its `iter` through `read_values` and pipelines both stages,
so a full scan of an append-only storage was the one read path without
batching -- one blocking round trip per value on the object store backends
this variant exists for. It is reached by payload storage iteration and by
the payload index build, which scans every payload.

Feed the pointers into `read_batch_values` instead, keeping the single
contiguous tracker read, which is better than the per-offset pipeline
scheduling Gridstore does on that side.

Values are now delivered through the read pipeline, so the callback may be
invoked out of order, as it already could be for Gridstore's `iter` and for
`read_values` in both variants. Both segment callers are order independent.
Tests that happened to rely on the mmap backend completing reads in
scheduling order now sort before comparing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test: don't run the failed-page-reload test on Windows

The test shrinks a page file out of band to make the page reload fail, but
Windows refuses to resize a file while the reader holds it mapped, which it
does by construction here: "the requested operation cannot be performed on a
file with a user-mapped section open". The panic is on the injection itself,
the code under test never runs.

There is no portable injection. Truncating a page the reader holds is what
the check under test detects, so the mapping cannot be avoided; failing the
adopted page open instead needs a listed but unopenable file, and
`local_list_files` descends into matching directories rather than listing
them; failing the directory listing needs the storage directory removed,
which Windows also refuses while pages are mapped.

The storage itself is fine on Windows, its append path grows mapped pages
there and every other Logstore test passes. The logic under test is platform
independent and stays covered elsewhere, with the tracker half of the
guarantee pinned by `test_live_reload`, which runs on every target.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Luis Cossío <luis.cossio@outlook.com>
2026-07-26 20:23:28 +02:00
a2b68597c9 feat: shared qdrant-edge-ffi crate (UniFFI boundary for mobile SDKs) (#9374)
* Add shared qdrant-edge-ffi crate with UniFFI bindings

Introduce a shared FFI crate that wraps Qdrant Edge's core types with
UniFFI attributes so the same Rust source can power both the Swift and
Kotlin bindings. The crate lives at lib/edge/ffi/ and exposes ~60 public
types (EdgeShard, Point, Query, Filter, UpdateOperation, …) plus their
enum / record / sealed-union variants.

- lib/edge/ffi/src/       UniFFI-wrapped domain types (config, types,
                          filter, query, update, error, lib)
- lib/edge/ffi/bindgen/   Separate crate housing the uniffi-bindgen CLI
                          (needed so consumers of the `uniffi` runtime
                          don't have to depend on its CLI feature)
- lib/edge/ffi/uniffi.toml Sets the generated Kotlin package to
                          tech.qdrant.edge.ffi (Swift uses the crate
                          name verbatim)

Every public type carries Rust doc comments that UniFFI propagates to
Swift Quick Help and Kotlin KDoc, so the generated bindings ship with
first-class IDE documentation.

Workspace changes:
- Cargo.toml          Adds the new crates as workspace members and
                      introduces a `release-mobile` profile
                      (thin LTO, codegen-units=1, strip=symbols,
                      panic=abort) for size-conscious mobile builds
- Cargo.lock          Pins uniffi 0.31 and its transitive dependencies

Made-with: Cursor

* fix(edge-ffi): harden FFI boundary, add tests, quantization parity, optimize/HNSW

Builds on @ivan-afanasiev's qdrant-edge-ffi foundation (preceding commit) — takes
it to a tested, safe, reviewable state. Split out of #9359 per maintainer request
so the FFI crate can be reviewed in isolation; Swift/Android SDK PRs stack on top.

Boundary safety (host input → catchable error, never a process abort):
- release-mobile profile switched to `panic = "unwind"` so UniFFI's catch_unwind
  turns a panic into a catchable error (abort would risk WAL/segment consistency
  on an on-device DB).
- Fallible boundary conversions reject bad input (UUID, geo, JSON path, payload
  JSON, contradictory match filters) with InvalidArgument instead of panicking.
- Host-supplied counts bounded: limit/offset (bounded_limit, 1 Mi cap), vector
  size (1..=65536), HNSW params (m/payload_m ≤ 2048, ef_construct 4..=100000,
  max_indexing_threads ≤ 1024) — these drive eager allocation / thread spawning
  at optimize(), so unbounded values would abort uncatchably.

API:
- Quantization parity with the Python Edge SDK: all four strategies
  (Scalar/Product/Binary/Turbo) accepted; HnswIndexConfig + optimize() exposed
  (without optimize() search is brute-force).
- config() is an honest "as-requested" read-back (HNSW + quantization round-trip).
- EdgeError stays branchable (ShardClosed / InvalidArgument / OperationError;
  field is `reason`, not `message`, to avoid the Kotlin Throwable collision).

edge core (required by the boundary):
- EdgeShard::flush is fallible (OperationResult) instead of panicking on lock
  contention; Drop logs a flush error instead of aborting; python flush()? updated.
- scroll.rs drops a with_capacity(limit) pre-alloc a huge limit could turn into an
  allocator abort (defense-in-depth alongside bounded_limit).

Tests (CI: cargo +nightly test -p qdrant-edge-ffi): 4 unit + 22 conversion +
18 integration — persistence, crash-recovery, payload round-trip, concurrency,
delete-reload, search ranking, scroll pagination, quantization accept + config
round-trip, HNSW optimize, boundary rejection.

* fix(edge-ffi): validate geo radius/rings and reject empty field conditions

Three boundary-validation gaps surfaced in review of #9374, all rejected
now with EdgeError::InvalidArgument instead of producing wrong results or
reaching a panic in the geo index:

- GeoRadius: a negative or non-finite radius passed straight through to
  the geo index. Reject !is_finite() || < 0.0.
- GeoLineString rings (exterior + interiors of a GeoPolygon): the segment
  type was built by direct struct literal, bypassing the engine's
  validate_line_string (which only runs on the serde path). A malformed
  ring (<4 points or unclosed) could panic in the geo index on indexed
  payloads. Mirror validate_line_string at the single GeoLineString
  conversion chokepoint, covering both exterior and interior rings.
- FieldCondition: a condition with no predicate set is a silent no-op
  (matches every point). Reject it, mirroring the engine's
  validate_field_condition. This is the engine/gRPC/REST/Python contract
  of "at least one" predicate -- NOT "exactly one"; multiple predicates
  remain valid and AND together. The doc comment is corrected accordingly.

Adds 9 conversion tests (geo radius negative/NaN/infinite/valid, ring
too-few/unclosed/bad-interior, field-condition no-predicate/multiple).
cargo +nightly test -p qdrant-edge-ffi: 53 green.

* fix(edge-ffi): adapt to memory/idf API and fallible info after rebase

Keep FlushMode::Sync from recent segment-holder changes while preserving
fallible flush. Fill newly required memory/idf fields (matching the Python
edge bindings) and propagate EdgeShard::info()'s OperationResult.

Co-authored-by: Cursor <cursoragent@cursor.com>

* ci(edge): disable checkout credential persistence in edge-test

actions/checkout persists GITHUB_TOKEN into .git/config by default. This
workflow runs on pull_request from any branch and then builds and runs
repository-controlled code (Rust/Python examples), which could read or
exfiltrate that token. The job never pushes, so drop the persisted
credentials with persist-credentials: false.

Addresses a CodeRabbit security finding on the edge-ffi PR.

* fix(edge): block on lock in flush instead of failing on contention

flush() used try_lock()/try_read() and returned a 'lock busy' OperationError
when a concurrent update/optimize held the WAL or segment lock. That branch is
the wrong trade-off: callers of flush() expect their data persisted, and the
one spot it would fire on the direct Rust path is exactly when an in-flight
update is holding the WAL lock across its whole operation — i.e. when there is
unflushed data most worth persisting. At the FFI boundary it is moreover dead
code, since the outer Mutex<Option<EdgeShard>> already serializes every call.

Switch to blocking lock()/read(), matching the semantics update() and optimize()
already use on these same locks. flush() stays fallible so a genuine WAL/segment
I/O error is still surfaced rather than panicking. Drop cannot contend (it needs
&mut self, so no &self borrow can hold the locks), so it will not hang.

Addresses a CodeRabbit review nitpick on the edge-ffi PR.

* fix(edge-ffi): harden vector/query boundary from multi-agent review

Addresses findings from a multi-agent review of the FFI boundary:

- Multivector conversions were infallible: an empty outer Vec panicked in
  release (MultiDenseVectorInternal::new_unchecked only debug_asserts) and a
  ragged matrix was silently reshaped against row[0].len(), storing data the
  host never sent. Make NamedVector/Vector -> persisted conversions TryFrom and
  validate the matrix (non-empty, uniform non-zero row width) with the same
  rules as the engine's try_from_matrix, returning InvalidArgument.
- Reject non-finite (NaN/inf) vector components at ingest, mirroring the geo
  is_finite guard. The engine validates only dimensionality, so a poisoned
  component would be stored and later serialized back as JSON null silently.
- unload() now returns Result: it flushes explicitly and surfaces a final fsync
  failure instead of only reaching Drop's log line (no default log sink exists).
  On error the shard stays loaded so the host can retry.
- Cap filter/prefetch nesting depth (MAX_QUERY_NESTING_DEPTH). Condition::Filter
  and nested Prefetch are self-recursive; an unbounded host tree would overflow
  the stack — a SIGABRT that panic=unwind cannot catch. Reject deeper trees as
  InvalidArgument via depth-threaded conversion helpers.
- Replace the '""'-on-serialization-failure fallback in payload/vector JSON
  encoding with .expect (serialization is infallible; fail loud, not silent
  invalid JSON).
- Docs: correct the flush() # Errors (can return OperationError), the edge-core
  flush() caller enumeration, upsert_points/update_vectors # Errors (vector
  validation), and reword the fictional lib/edge/VERSION / version-sync comment
  in ffi/Cargo.toml to reflect that no automated check exists yet.

* test(edge-ffi): add behavior coverage for vectors, filters, query, updates

Adds 18 integration tests closing gaps a multi-agent review flagged (the suite
proved type conversion but not behavior):

Safety (back the new boundary validation):
- multi_vector_invalid_matrices_rejected — empty/ragged/zero-dim multi-vectors
- non_finite_vector_components_rejected — NaN/inf across single/named/multi/sparse
- finite_vectors_accepted_by_upsert_constructor — over-rejection guard
- deeply_nested_filter_rejected_shallow_accepted, deeply_nested_prefetch_rejected
  — depth cap rejects >64, accepts shallow

Behavior:
- filter_restricts_count_scroll_and_search — a filter actually narrows the result
  set across count/scroll/search (not just that conversion succeeds)
- flush_under_concurrent_upserts — flush() blocks under a concurrent update loop,
  no panic, final count == successful upserts
- vector_content_round_trips_through_retrieve_and_search
- cosine_distance_ranks_by_direction, euclid_and_manhattan_rank_nearest_first
- delete_points_by_filter / update_vectors / delete_vectors / delete_payload /
  clear_payload — the five previously-untested update ops
- multivector_round_trips, sparse_vector_round_trips
- rrf_fusion_over_prefetches_returns_fused_set

Suite: 4 unit + 32 conversion + 36 integration = 72, all green.

Not covered (blocked by FFI surface, tracked for follow-up): facet() and OrderBy
scroll both need a payload index, and UpdateOperation exposes no create-index
constructor.

* fix(edge): avoid lost-update TOCTOU in set_vector_hnsw_config

set_vector_hnsw_config read().clone()'d the config, mutated the clone, then
write(|c| *c = cfg) overwrote the whole config. The read lock is released before
the write, so a concurrent config update between the two is silently discarded
(lost-update TOCTOU). Use SaveOnDisk::write_optional to run the fallible mutation
on a clone inside the held lock, returning None on failure to abort persist+swap
without overwriting — the atomic pattern the repo's SaveOnDisk learning
prescribes for fallible config mutations.

Addresses a CodeRabbit critical finding on the edge-ffi PR.

* fix(edge-ffi): adapt read requests to edge::* structs after #9901 rebase

#9901 (Edge: request-specific structures for EdgeShardRead) moved the read API
off the shard/core request types onto edge-owned request structs. Retarget the
FFI conversions accordingly:

- QueryRequest/SearchRequest/CountRequest/ScrollRequest/FacetRequest/Prefetch
  now convert into edge::{QueryRequest,SearchRequest,CountRequest,ScrollRequest,
  FacetRequest,Prefetch} (were ShardQueryRequest/CoreSearchRequest/*Internal).
  Field shapes are identical except score_threshold, which is a plain ScoreType
  (f32) on the edge structs, not OrderedFloat — drop the wrap. Nesting-depth
  guard and bounded_limit validation preserved.
- retrieve() builds an edge::RetrieveRequest and calls the new single-argument
  EdgeShardRead::retrieve (was a 3-arg call).
- Cargo.lock reconciled onto dev's lockfile + the FFI/uniffi deps (dev advanced
  23 commits incl. dependency bumps).

cargo +nightly test --locked -p qdrant-edge-ffi: 4 unit + 32 conversion + 18
integration all green.

* feat(edge-ffi): full engine coverage — restructure, all update ops, all scoring queries, missing shard methods

Restructure the crate to mirror the edge crate's layout: shard.rs owns the
EdgeShard object and lifecycle, ops/ has one file per read operation
(request/response records + conversions + exported method together), and
update construction stays in update.rs. Multiple #[uniffi::export] impl
blocks merge cleanly in the generated bindings.

Interface modernization:
- retrieve() takes a RetrieveRequest record mirroring edge::RetrieveRequest
- config surface moves off the deprecated always_ram/on_disk booleans to a
  Memory placement enum (cold/cached/pinned); read-back resolves legacy
  flags via memory_placement(), None-preserving for HNSW
- EdgeShard.inner is RwLock<Option<...>>: operations take the read half and
  run in parallel; unload/update_from_snapshot take the write half and
  drain in-flight requests
- request/config records carry #[uniffi(default = ...)], so generated
  Swift/Kotlin constructors get default arguments (count exact=true, facet
  limit=10, HNSW 16/100/10000/0, everything optional defaults to nil/null)
- all conversions destructure their source exhaustively; intentionally
  unexposed internal fields are named `field: _` with a why-comment

Full update-operation coverage (Python SDK parity):
- upsert_points gains condition/update_mode (conditional upsert),
  update_vectors gains condition, set_payload gains a JSON-path key
- new: delete_vectors_by_filter, set/overwrite/delete/clear payload
  by-filter forms, overwrite_payload, create/delete_field_index (with a
  PayloadSchemaType enum), create_dense_vector, create_sparse_vector,
  delete_vector_name

Full scoring-query coverage:
- Query::Nearest takes a NamedVector (dense/sparse/multi-vector search)
- new Query variants: Recommend (BestScore/SumScores), Discover, Context,
  Feedback; new ScoringQuery variants: Formula (recursive Expression
  object with validating, depth-capped constructors) and Mmr;
  Fusion::Rrf gains weights

Missing shard methods: query_groups, search_matrix, create, path,
snapshot_manifest, update_from_snapshot (full + partial recovery),
set_hnsw_config, set_vector_hnsw_config, set_optimizers_config.

Two compile-time coverage maps (update.rs, ops/query.rs) exhaustively
match the engine's operation/query enum trees with no wildcard arms, so a
new engine variant fails compilation in this crate until the FFI surface
decision is recorded. A staging passthrough feature keeps them exhaustive
when shard/staging is enabled. The scoring-query map immediately caught
the otherwise-missed Mmr variant.

Tests: 82 total (4 unit + 32 conversion + 46 integration), including new
end-to-end coverage for field-index-enabled facet, conditional upsert,
overwrite/clear-by-filter, recommend/discover/context/MMR, formula
re-scoring, grouped queries, search matrix, vector-name ops, and the
lifecycle additions. Kotlin+Swift binding generation verified.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(edge-ffi): cover the remaining enum boundaries — filter tree, formula, schema, selections

Add coverage maps for the four construction-only enum families that had no
exhaustiveness guard, and expose everything they revealed as missing:

Filter surface (the big one — filter.rs map over Condition/Match/
RangeInterface/AnyVariants/ValueVariants):
- Condition::Slice — deterministic id-space slice filter, with the
  serde-path total/index validation re-applied at the boundary
- Condition::Nested — nested-object array filters, depth-counted against
  the recursion cap
- Match::TextAny / Match::Phrase / Match::Prefix — the three previously
  unreachable text-match modes
- FieldCondition.datetime_range — RFC 3339 datetime ranges (mutually
  exclusive with the float range, matching the engine's single
  RangeInterface slot)
- Filter.min_should
- WithPayload::Exclude — exclude-style payload selection (types.rs map
  over WithPayloadInterface/PayloadSelector/WithVector)
- OrderBy.start_from — Integer/Float/Datetime cursor (mapped in the
  scoring-query coverage map)
- CustomIdChecker recorded as intentionally unexposed (serde-skipped,
  runtime-internal)

Formula (formula.rs map over ExpressionInternal + DecayKind): all 17
expression variants were already constructible; the map now forces a
decision when the engine grows a new one.

Field-index schema and vector-name config (update.rs map extended):
PayloadFieldSchema::FieldType covered per schema type; the per-type
FieldParams forms recorded as not exposed yet; VectorNameConfig
Dense/Sparse tied to their constructors.

Tests: 91 total (+7 conversion, +2 integration: slice partitioning and
payload-exclude retrieval). Kotlin binding generation verified for the
new types.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(edge-ffi): resolve clippy warnings (inline format arg, large enum variant)

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(edge-ffi): address multi-agent review of the full-coverage expansion

Fixes findings from a multi-agent review (each independently re-verified),
targeting the new surface added in the full-engine-coverage expansion. All
verdicts CONFIRMED except query_groups (refuted → no change). 100 tests green.

- formula: cap Expression node count (MAX_FORMULA_NODES = 10_000), not just
  depth. Expression is a host-held Arc; combinators eagerly deep-clone their
  children's inner tree, so reusing one handle as several children
  (sum([e, e]) repeated) grows node count as 2^depth while depth stays under
  the 64 cap — an eager multi-GB clone that aborts the process (uncatchable
  under panic=unwind). Track a saturating size in node() and decay(); the
  depth guard stays (it protects the stack for narrow-but-deep chains).
- sparse: validate host sparse vectors at the boundary via
  validate_sparse_vector_impl (indices.len == values.len, unique indices).
  Without it a length mismatch reached an out-of-bounds panic at insert/scoring
  (opaque caught-panic, not a typed error) and duplicate indices silently
  double-counted. Route both sparse arms through a fallible helper; the
  infallible From<SparseVector> is removed so no path can bypass validation.
- order_value: surface it on ScoredPoint/Record (new OrderValue { Int, Float },
  matching the Python edge SDK). It was dropped behind a comment claiming "no
  order-by-scored surface" — false: order-by query/scroll and OrderBy.start_from
  are exposed, and ordered results carry a constant score and no next_offset, so
  order_value is the only way to resume ordered pagination when payload is off.
- search_matrix: drop from v1. It is an O(n^2) analytics op the reference Python
  edge SDK does not expose; its flat 1-Mi bounds are the wrong shape (one call
  hangs/OOM-crashes the device on a normal shard). Removing pre-publish is free;
  re-adding later with proper caps is non-breaking, the reverse is not.
- tests: node-count reject, sparse len/dup reject, order_value populated +
  start_from resume, and snapshot negative tests (bad path / corrupt archive /
  post-unload surface OperationError, not a panic, and the shard survives).

* fix(edge-ffi): reject oversized formula before the eager clone; close test gaps

Follow-up from a re-review of the fix commit.

- formula: the node-count check ran AFTER the eager `inner` deep-clone, because
  the clone was an argument to `node()` (evaluated before the function body). A
  host could build one accepted handle and reuse it as many children in one call
  (sum(vec![e; N])), materializing N x e.size nodes before the size check could
  reject it — the same uncatchable OOM abort the cap was meant to prevent, in two
  calls. `node()` now takes the children handles plus a build closure and runs
  both caps BEFORE invoking it, so a rejected tree never clones (decay() already
  did this). All combinators route through it.
- tests: decay node-count reject + happy path (decay has its own guard, was
  untested); wide-fanout reject (regression for the ordering fix); query-side
  ScoringQuery::OrderBy -> ScoredPoint.order_value (only the scroll/Record side
  was covered); real multi-page order-by continuation via StartFrom (was a
  single-page degenerate case); sparse values-longer reject; float OrderValue;
  with_payload=false omits payload. Suite: 107 green (4 + 39 + 64).

* feat(edge-ffi): keep search_matrix behind an off-by-default `matrix` feature

Reconsidered the outright removal: the FFI crate is a general UniFFI boundary,
not mobile-only, so non-mobile consumers (desktop/server/Rust) may want the
distance-matrix op. Instead of deleting it, gate the whole `ops/matrix.rs`
module behind a new off-by-default `matrix` Cargo feature.

- The mobile Swift/Kotlin bindings build without the feature, so search_matrix
  stays off the mobile surface (verified: default `cargo test` excludes it and
  its test; `--features matrix` includes both).
- The O(n^2) DoS is documented, not capped here: the op is opt-in and off the
  constrained mobile surface, so bounding sample_size/limit_per_sample for a
  given deployment is the enabling consumer's / SDK layer's responsibility. The
  module doc spells this out; the per-field bounded_limit only stops a lone
  u64::MAX value, not the quadratic compute.
- CI: add an `--all-features` test leg so the feature-gated surface (matrix +
  the pre-existing staging passthrough) can't bit-rot — this also closes the
  build-publish review note that `staging` had no CI coverage.

* docs+test(edge-ffi): tighten matrix doc wording; pin formula node-count boundary

Non-blocking polish from a final all-reviewer pass (6/6 APPROVE):

- matrix.rs / Cargo.toml docs: (1) the per-field bounded_limit caps at
  MAX_RESULT_COUNT (1,048,576), not merely a lone u64::MAX — say so, since 1 Mi
  is itself catastrophic for an O(n^2) op; (2) the DoS bound is owned by the
  opting-in consumer, not an "SDK/wrapper layer" that need not exist for a raw
  UniFFI consumer; (3) the mobile bindgen (a follow-up PR) must build with
  default features to keep the op off the mobile surface — stated as intent, not
  present-tense fact (no swift/android dirs exist yet).
- formula_node_count_exact_boundary: pins the exact MAX_FORMULA_NODES threshold
  (10_000 accepted, 10_001 rejected) — the existing tests jumped to ~16k, so the
  precise cap edge was unverified.

Suite: 108 default / 109 --all-features, all green.

* feat(edge-ffi): make the matrix feature on by default

Flip `matrix` to on-by-default (`default = ["matrix"]`). General (desktop /
server-side / Rust) UniFFI consumers now get `search_matrix` without opting in;
the mobile Swift/Kotlin bindgen builds with `--no-default-features` to drop the
O(n^2) analytics op from the mobile binding surface.

CI now covers all three shapes: default (matrix on), `--no-default-features`
(mobile surface, matrix excluded — guards the crate still compiles/passes
without it), and `--all-features` (matrix + staging). Verified: default 66 /
no-default 65 / all-features 66 integration tests, all green; no Cargo.lock
drift.

* fix(edge-ffi): rustfmt the sparse validator; correct stale matrix-feature docs

From a full all-reviewer pass of the review fixes:

- types.rs: rustfmt the `to_internal_sparse` `.map_err` chain. It failed
  `cargo +nightly fmt --all -- --check` (the rust-lint.yml CI gate) — the
  edge-test legs only run `cargo test`, so it slipped through locally.
- ops/mod.rs + integration.rs: fix two doc comments that still said the `matrix`
  feature is 'off-by-default' after it was flipped on-by-default. Left uncorrected
  they'd mislead the follow-up mobile-bindgen author into skipping
  `--no-default-features` and shipping the O(n^2) op to phones.

Design decisions (recorded): matrix stays on-by-default; the O(n^2) DoS is
documented, not capped — a static cap can't know the target device (compute cost
is device-dependent; the caller owns it via async/timeout). No cap added.

Reviewers: 3 APPROVE, 3 CHANGES-REQUESTED, all CR items were these two doc/fmt
misses plus the recorded design calls. All 3 feature configs green (66/65/66).

* fix(edge-ffi): resolve clippy errors in FFI tests

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(edge): reject conflicting vector-name re-create instead of desyncing config

The segment-level `create_vector_name` is idempotent: re-creating an existing
name is a no-op that leaves storage untouched. `EdgeShard::update` treated any
`Ok` from that no-op as success and unconditionally re-applied the requested
params to the shard config, so a second `create_dense_vector("v", 8, Cosine)`
after `("v", 4, Dot)` left `config()` advertising 8/Cosine while storage kept
4/Dot — a shape the shard then rejects on upsert. The empty name "" (the
default vector) hit the same path against a single-vector shard's primary field.

Reject a conflicting re-create up front (before the WAL, so it can't brick
replay) with a clear error, accept an identical re-create as idempotent, and
only re-apply config when storage actually changed.

* fix(edge-ffi): harden FFI boundary from an adversarial per-search-type review

Boundary-validation and doc fixes surfaced by fuzzing each query type with
executed repros:

- Reject non-finite floats that JSON cannot represent but the raw-f64 FFI can:
  order-by `StartFrom::Float` (a NaN panicked on a float-indexed field and
  silently truncated an integer scan), and formula `decay` midpoint/scale +
  `div` by_zero_default (a NaN evaded the engine's comparison-based range
  checks → debug panic across the boundary / silent all-zero rescore).
- Drop the `key` parameter from `overwrite_payload`/`overwrite_payload_by_filter`:
  the engine has no payload selector on the overwrite path (the server discards
  it too), so a keyed overwrite silently replaced the whole payload.
- Reject a `FieldCondition` carrying more than one predicate: the engine has no
  defined semantics for it (it evaluates one, index-dependent), so passing
  several through diverged silently from a Qdrant server. Callers AND predicates
  via separate `must` conditions.
- Doc fixes: FeedbackCoefficients b/c (exponent/multiplier, not weight/margin),
  RecommendStrategy::BestScore default, div-by-zero behavior, retrieve
  duplicate-ID collapsing, non-atomic partial-snapshot recovery.

Also fix the clippy `--all-targets -D warnings` lints in the test files that
reddened CI's `lint` job (uninlined_format_args, err_expect, disallowed
std::fs::write) and add regression tests for the validations above.

* fix(edge): compare only vector identity fields when detecting re-create conflicts

The conflicting-re-create guard added in the previous commit compared the full
`EdgeVectorParams`/`EdgeSparseVectorParams` via `PartialEq`, but a vector
declared in the initial `EdgeConfig` is stored with `on_disk: Some(false)` (from
`from_vector_data_config`) while a `CreateVectorName` op leaves it `None`. So an
identical re-create of a construction-defined vector — or of any vector after
`set_vector_hnsw_config` — was falsely rejected as a conflict, a regression on
the idempotent path.

Compare only the identity fields the op actually defines (dense: size, distance,
multivector_config, datatype; sparse: modifier, datatype); the storage/tuning
fields it cannot express are set from the config, the optimizer, or
`set_*_config` and must not trigger a false conflict. Adds a regression test
re-creating the config-defined "vec".

* style: apply cargo fmt

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(edge-ffi): expose payload_schema in info() and parameterized index creation

Close the last interface gap in the FFI surface: `info()` previously
elided the engine's `payload_schema`, and `create_field_index` only
accepted a bare type.

- New `payload_index` module mirroring the full `PayloadSchemaParams`
  family (all 8 index types incl. tokenizer/stopwords/stemmer options)
  with bidirectional conversions.
- `ShardInfo.payload_schema` reports each index's type, creation params,
  and indexed point count.
- New `UpdateOperation::create_field_index_with_params` constructor;
  coverage map updated accordingly.
- Boundary normalizations, following the quantization-config precedent:
  deprecated `on_disk` folds into the reported `memory` placement, and
  contradictory integer params (lookup+range both off) reject with
  `InvalidArgument`.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Ivan Afanasiev <ivan.afanasiev@yahoo.com>
Co-authored-by: root <111755117+qdrant-cloud-bot@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 18:26:32 +02:00
Andrey VasnetsovandClaude Fable 5 446d140c2d Slice filtering condition: sliced scroll / deterministic sampling (#9899)
* feat: slice filtering condition for sliced scroll and deterministic sampling

Add a `slice` filter condition selecting points where
`stable_hash(point_id) % total == index`. The hash is SipHash-2-4 with a
zero key over canonical id bytes (8 LE bytes for numeric ids, 16 RFC 4122
bytes for UUIDs) — a frozen public contract, independent of the internal
resharding ring hash, reproducible by clients to predict membership.

For a fixed `total`, slices are disjoint and cover all points, enabling
parallel scroll streams (ES sliced-scroll style) and reproducible sampling
that composes with any other filter condition.

- REST: `{"slice": {"total": N, "index": R}}`; gRPC: `SliceCondition` in
  the condition oneof (tag 8)
- Evaluated per point via id_tracker external-id lookup; no payload index
  needed; cardinality estimated as `points / total` with no primary clause
- `total >= 1` enforced by NonZeroU32 at parse time, `index < total` by
  validation in both REST and gRPC paths
- Hash contract locked by test vectors independently reproduced with a
  reference SipHash-2-4 implementation

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* tests: minimal OpenAPI test for slice filter condition

Scrolls all slices of a fixed total over numeric + UUID ids asserting
disjointness and full coverage, checks must_not inversion, and pins the
two rejection paths (422 for index >= total, 400 for total = 0). Requests
and responses are validated against the regenerated OpenAPI spec by the
test harness.

Note: the spec cannot itself reject total = 0 client-side — the Condition
anyOf falls through to the permissive Filter schema, as with any invalid
condition — so rejection is asserted via the server response.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 17:45:17 +02:00
Andrey VasnetsovandClaude Fable 5 4ff6acaf8c Edge: request-specific structures for EdgeShardRead (#9901)
Replace the mixed read interface (internal types, explicit parameter
enumeration, ad-hoc custom types) with edge-owned request structs, one
file per request in src/requests/. Each has a new() constructor taking
only the required parameters, a no-macro fluent builder in src/builders/,
and a From conversion into the internal request type in
requests/conversions/, grouped by request type. Conversions construct
and destructure with full field lists, so a parameter added on either
side fails compilation instead of being silently dropped.

The old reexport aliases (ScrollRequest = ScrollRequestInternal, etc.)
are replaced by the edge types under the same names; retrieve() takes a
RetrieveRequest instead of a parameter triple. Python bindings wrap the
edge types, and the published Rust examples use the builders.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 12:53:58 +02:00
Andrey VasnetsovandClaude Fable 5 842f701aae refactor: group edge crate internals into edge_shard and read_view modules (#9898)
* refactor: group edge crate internals into edge_shard and read_view modules

Restructure lib/edge/src so the top level only contains the modules that
form the public crate surface. Implementation files move under the type
they implement:

- edge_shard/: the EdgeShard struct with its load/config-resolution
  helpers (previously inlined in lib.rs), plus optimize, shard_read,
  snapshots, and update
- read_view/: the EdgeShardRead/ReadSegmentHandle traits and
  EdgeReadView (previously read_view.rs), plus the per-operation impl
  files count, facet, grouping, info, matrix, query, retrieve, scroll,
  and search; build_search_pool (previously pool.rs) is folded into
  read_view/mod.rs next to par_map_segments, the seam it powers

lib.rs is now a thin facade of module declarations and re-exports. The
public API is unchanged: all previously exported names resolve exactly
as before, verified against the edge tests, the python bindings,
edge-shard-query, and the regenerated publish amalgamation with its
examples.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor: trim EdgeShardRead surface and split read_view/read_only modules

Trait surface:
- Drop query_scroll and rescore_with_formula from EdgeShardRead and the
  EdgeShard inherent wrappers; only the internal query pipeline used
  them, via the pub(crate) EdgeReadView methods that remain.
- Hide the snapshot plumbing (read_segments, search_pool, plus
  config_snapshot/path providers) in a crate-private ReadViewProvider
  trait. EdgeShardRead now declares only user-facing methods and is
  implemented for every provider through a blanket impl, so the
  plumbing is not callable from user code (verified with a negative
  compile test; a private supertrait alone leaves supertrait methods
  callable through generic bounds). An empty sealed marker supertrait
  keeps the trait unimplementable downstream.
- config_snapshot and path stay public: used by edge-shard-query and
  the python bindings.

Module layout:
- read_view/: mod.rs keeps EdgeReadView and build_search_pool;
  ReadSegmentHandle moves to handle.rs, the trait machinery to
  shard_read.rs, and the nine per-operation impl files into ops/.
- read_only/: the follower's ReadViewProvider impl moves out of mod.rs
  into shard_read.rs, mirroring edge_shard/shard_read.rs.

Verified: edge tests, python bindings, edge-shard-query, regenerated
publish amalgamation with all examples compiling and running.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 16:09:04 +02:00
0982e8699c feat: segment manifest optimizing state with lease (#9873)
* feat: segment manifest optimizing state with lease

* fix: clippy

* fix: exhaustive manifest state matching

* fix: merge manifest rebuilds under the write lock

* refactor: named state predicates, drop redundant enumerator test

Review follow-up: move the enumerator's filter into
SegmentManifestState::is_usable, name the preserving predicate
is_optimizer_mark, delete the enumerator test that re-tested serde.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 04:29:03 +02:00
Andrey VasnetsovandClaude Opus 4.8 8888046cd3 Add usage skill file for edge-shard-query tool (#9836)
Standalone usage reference for `edge-shard-query`: backends (S3 / GCS /
uio-grpc), connection and tuning flags, the scroll / search / search-sparse
sub-commands, filtering, and the live-reload diff mode.

Written to stand on its own, so it can be shared as a link without also
sharing the source.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 15:54:33 +02:00
Andrey VasnetsovandClaude Fable 5 a5a0ecc5e9 Add search-sparse sub-command to edge-shard-query tool (#9834)
Sparse nearest-neighbour search over a ReadOnlyEdgeShard, reusing the
same SearchRequest path as dense search with VectorInternal::Sparse.
The query vector is accepted as a JSON object or an index:value pair
list, sorted and validated before use. --vector is required (no random
fallback: the sparse vocabulary is not recoverable from the shard
config), and --hnsw-ef is omitted since it does not apply to the
sparse index.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 15:35:38 +02:00
1d4d6f02da Per-query IDF corpus for sparse vector search (#9661)
* Add per-query IDF corpus for sparse vector search

Let the caller choose, per query, which population sparse IDF statistics
are computed over. `params.idf` is either `"global"` (default, unchanged
behavior) or `{"corpus": <filter>}`, where the corpus filter is
independent of - and usually broader than - the retrieval filter.
Decoupling the two keeps the score scale stable when the retrieval
filter tightens: term importance is measured against a population the
user names, not against whatever subset the filter happens to select.

Design decisions:
- Corpus grammar is restricted to a conjunction (`must`) of `match`
  conditions on payload fields; loosening later is backward compatible.
- Strict mode validates the corpus filter like a read filter
  (unindexed fields rejected).
- `idf` on a vector without the IDF modifier is a validation error,
  never silently ignored.
- An empty corpus yields degenerate but corpus-scoped scores (smoothed
  IDF over N=0), never a fallback to global statistics - in multi-tenant
  collections a fallback would leak term statistics across tenants.

Implementation:
- QueryContext IDF stats are keyed by corpus, so one batch can mix
  requests with different corpora.
- Statistics come from the sparse index: df(term) is counted over the
  query terms' posting lists only, never by scanning stored vectors.
  Small corpora (under ~1/32 of the segment, by cardinality estimate)
  are kept as a sorted id list galloping through posting lists via
  skip_to; large ones as a dense membership mask filled streaming from
  the filtered-points iterator. A misestimated small corpus degrades
  into the mask.
- Exposed uniformly: REST (`params.idf`), gRPC (`IdfParams` message),
  edge python bindings; OpenAPI schema regenerated.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Apply rustfmt

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix clippy manual_is_multiple_of in sparse IDF corpus test.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Allow any filter as IDF corpus

Drop the must+match grammar restriction on the corpus filter. A
restriction enforced only as a validation step over the full Filter
type buys nothing; if a narrower corpus syntax is ever wanted, it
should be a dedicated API-level type instead.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix build: add memory field to SparseIndexConfig in idf corpus test

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: root <111755117+qdrant-cloud-bot@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-14 11:16:47 +02:00
Andrey VasnetsovandClaude Fable 5 43e3d6ea8d [UIO] Request-specific load profile for read-only opens (#9797)
* Introduce request-specific LoadProfile with per-component placement

A read-only shard opened for one known request (the serverless cold-start
path) doesn't have to warm components the request will never touch.
LoadProfile captures that from the request: warm components keep the
persisted-config placement, everything else is parked cold. All placement
decisions live in one place, so the memory placement of a whole segment
under a profile is reviewable in one file.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzbiZFVHooxKVoJX4k6Ais

* Thread LoadProfile through the read-only segment open

ReadOnlySegment::open takes an optional profile; first_preopen and
open_via resolve it into per-component populate overrides so the opens
make the same placement decisions the prefetches did. Pinned components
that materialize on open regardless (quantized RAM storage kinds, the
immutable-RAM sparse index) and appendable components ignore the
override; the HNSW graph and immutable payload indexes demote fully.
Config reloads follow the new config alone and pass no override.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzbiZFVHooxKVoJX4k6Ais

* Open ReadOnlyEdgeShard under a request-derived load profile

ReadOnlyEdgeShard::open takes an optional LoadProfile, applies it to
every segment open and keeps it so segments discovered by a later
refresh load with the same placement. ScrollRequestInternal and
CoreSearchRequest gain load_profile() constructors, and edge-shard-query
builds the request before the open and passes its profile (opt out with
--no-load-profile).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzbiZFVHooxKVoJX4k6Ais

* Demote pinned quantized vectors and sparse index under a cold profile

Within the immutable layout the quantized RAM and mmap loaders share the
on-disk format — only how the data is brought into memory differs — and
the immutable-RAM sparse index has the same lazy mmap open low-memory
mode already downgrades to. So a cold populate override now demotes the
effective placement itself (Memory::with_populate_override, shared with
the HNSW residency mapping) instead of only skipping cache priming: a
pinned quantized storage opens the mmap kind cold, and a pinned sparse
index opens as Mmap, so neither reads its data on a cold start.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzbiZFVHooxKVoJX4k6Ais

* Add LoadProfile::merge for composite queries

A composite query runs multiple core requests — e.g. a hybrid search
runs one core search per vector, each with its own filter. Its profile
is the union of its parts': merge extends the warm sets and ORs the
payload-storage flag, so a component either part needs warm stays warm.

The union is sound because every placement method is monotone in the
warm sets (growing them only turns "park cold" into "keep configured
placement"), so the merged profile dominates each input; and minimal,
warming nothing no part asked for.

Combine profiles with reduce, not fold: merge's identity element is
the coldest profile (empty warm sets), the opposite of passing no
profile at all — deliberately no empty()/Default constructor exists.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Adapt vanished-segment test to the profile-aware open signature

The test landed on dev (#9777) after the load-profile signature change
was written, so the rebase left its ReadOnlySegment::open calls without
the new load_profile argument.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Defer vector index open entirely under a cold load profile

A cold placement is not enough for the vector index on remote backends:
GraphLinksView requires the whole links file as one contiguous slice,
and the disk cache can only lend a borrowed slice once every block is
locally present — so even a Cold HNSW open mirrors the entire
links_compressed.bin (8.2 MB / 1.1 s per segment in the serverless
cold-start trace), plus the unconditional graph.bin metadata read.
The only way not to fetch the index is not to open it.

LoadProfile::vector_index_placement is replaced by
vector_index_deferred: a vector the request never scores now gets a
DeferredVectorIndex — a new VectorIndexReadEnum variant holding the
open arguments (an owned clone of the segment's raw backend, path,
config, shared component handles) and a OnceLock. Nothing is opened or
prefetched for it at segment open.

Per-method policy of the deferred variant:
- search, fill_idf_statistics and populate open the index on first use
  (with the cold placement the profile chose), so the profile contract
  holds: a request the profile did not predict still works, just pays
  the open then;
- is_index reports true without opening (deferral only ever wraps a
  real HNSW or sparse index; plain opens no files and is never
  deferred);
- telemetry, indexed_vector_count and sizes answer conservative
  defaults rather than trigger a remote fetch for a statistic.

Tests: deleting the vector_index directory before an open under a
scroll profile leaves open, filtered reads and payload reads working —
proof that nothing of the index is read — while a search surfaces the
missing files; and a segment opened under a scroll profile answers
searches identically to an eagerly opened one via the transparent
first-use open.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Make --vector optional in edge-shard-query: random query after open

Omitting --vector on the search sub-command now searches with a random
vector. The request is still built before the shard opens — the load
profile only needs the vector name, not its values — with an empty
placeholder; once the shard is open, fill_random_vector reads the
dimension of the queried vector from the derived shard config and
fills in uniform-random f32s (with a clear error if the named vector
is not in the config).

The vector is generated once, so live-reload iterations re-run the
identical random query and the printed diffs stay meaningful.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Scope index deferral to the HNSW graph via a lazy OnceLock load

Replace the DeferredVectorIndex wrapper (and the VectorIndexReadEnum::
Deferred variant) with deferral inside ReadOnlyHNSWIndex itself: the
graph lives in a OnceLock (same first-wins arbitration as
ReadOnlyRoaringFlags::bitmap) alongside the retained raw backend and
residency, and loads on first use with a cold placement. The config
read stays eager — one tiny, absence-tolerated file — so telemetry,
is_on_disk and indexed_vector_count report real values where the
Deferred arms answered with hard defaults.

The sparse index needs no deferral: its mmap open reads lazily, with
only small JSON metadata eager. A profile that never scores the vector
now passes a cold placement override (LoadProfile::
vector_index_placement) into the eager open_sparse, which demotes
ImmutableRam to the lazy Mmap open like low-memory mode.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Express HNSW graph deferral as a populate override, not a bool param

Replace the `deferred: bool` on the read-only HNSW open/preopen (and
the VectorIndexReadEnum pass-through) with the same
`populate_override: Option<Populate>` every other component takes. A
cold *override* defers the graph load — graph_deferred() mirrors the
cold-override match of open_sparse — while a config-derived cold
placement (or the low-memory clamp) keeps the eager load, since only a
request-specific override carries the "never scored" prediction.

With dense and sparse now consuming the same signal,
LoadProfile::vector_index_deferred is gone: a single
vector_index_placement() serves both index kinds.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Serialize the deferred graph load via once_cell's get_or_try_init

Loading outside the lock (std OnceLock's fallible init is still
unstable) let a search burst on a deferred vector fetch the whole
graph once per thread. Swap the cell for once_cell::sync::OnceCell:
the fallible load runs inside the cell's lock, concurrent first users
block on the one load, and a failed load leaves the cell empty so the
next caller retries.

Addresses https://github.com/qdrant/qdrant/pull/9797#discussion_r3573727836

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 23:03:14 +02:00
f982ae63fc feat: read-only edge grouping/matrix search + object-storage read path (#9691)
* feat: read-only edge grouping/matrix search + object-storage read path

Add query_groups (group by a payload field) and search_matrix (single-shard distance matrix over a random sample) to the read-only edge shard's EdgeShardRead API, in new grouping and matrix modules plus edge test helpers.

Make the object-storage read path available outside tests: drop the #[cfg(test)] gate on the BlobFile UniversalReadExt impl and move io_bridge_object_store/object_store to segment's normal dependencies, so a ReadOnlyEdgeShard can serve segments read from S3.

* Share group-by building blocks between server and edge

Move GroupsAggregator, group candidate query shaping (is-empty filter,
group_by payload selector, prefetch limit scaling) and result-order
derivation into shard::grouping / shard::query, so the collection and
edge grouping implementations cannot silently diverge.

Edge grouping now handles multi-valued group keys, u64 keys, wildcard
group_by paths, prefetch limits and score-ordered groups the same way
as the server.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Drive group-by through a shared sans-IO state machine

Extract the multi-request collect/fill loop into
shard::grouping::GroupByDriver: next_request() yields shaped backend
queries, add_points() advances the state, distill() returns the groups.
Query execution stays with the caller, so the async server path and the
sync edge path drive the same machine, and the request shaping helpers
become private to shard::grouping.

Edge now uses the same request budget (5 collect + 5 fill requests) and
per-request candidates limit (groups * group_size, computed inside the
driver) as the server, replacing its single 4x-oversampled request.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 18:17:58 +02:00
Andrey VasnetsovandClaude Fable 5 f82d7584a2 Resolve vanished segments against the manifest in follower refresh (#9792)
Third step of the ReadOnlyEdgeShard not-found handling (after #9763 and
#9777): shard-level resolution of live-reload failures, per the
live-reload design requirements.

refresh() is now a bounded retry loop over single-manifest-snapshot
attempts. Within an attempt every survivor live-reloads even if one
fails (they are independent); failures split by classification:

- Not-found: resolved against a re-read manifest. Gone from the
  manifest means the leader removed the segment mid-reload — drop it
  and re-run the attempt to pick up its replacements immediately.
  Still listed means essential files are genuinely missing — escalate.
- Anything else: escalate after reloading the other survivors,
  replacing the previous warn-and-swallow that could hide a corrupted
  segment forever. Safe: a failed segment keeps serving its
  pre-refresh state and pending_reload replays the unapplied delta on
  the next refresh.

If attempts run out (leader churning segments continuously) refresh
logs a warning and returns Ok: every swap was atomic, so the shard is
consistent, just possibly not the newest; the next refresh continues.

open_with_enumerator now reuses the refresh machinery: empty holder +
refresh() + the open-only merge_follower_config overlay, removing the
duplicated load/derive logic. At open there are no survivors, so open
behavior is unchanged (#9762 already made it tolerant of unloadable
segments via the superset-biased manifest contract).

Load-side failures stay as settled by #9762: unloadable listed
segments are skipped and retried on every refresh.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 14:50:29 +02:00
Andrey VasnetsovandClaude Fable 5 bb0e5d7544 Fix flaky disk exhaustion in edge-test CI job (#9807)
The publish/ directory is a separate Cargo workspace, so it does not
inherit the root workspace's [profile.dev] debug = "line-tables-only"
override. Example binaries were built with full debug info, each
statically linking qdrant-edge at ~700 MB per binary. Linking several
of them concurrently ran the runner out of disk, surfacing as
"mold: failed to write to an output file. Disk full?" + SIGBUS.

Mirror the debug info override in the publish workspace and free
~20 GB of unused preinstalled software on Linux runners as headroom.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 11:12:07 +02:00
Andrey VasnetsovandClaude Fable 5 8f54076cbe Add uio-grpc backend and key-triggered live-reload to edge-shard-query (#9795)
* Add uio-grpc backend and key-triggered live-reload to edge-shard-query

--backend uio-grpc opens the shard directly over a running Qdrant
peer's StorageRead gRPC service (public gRPC endpoint, api-key aware),
addressed by --collection/--shard-id — no object storage involved. The
UioGrpcSource backend existed since #9634 but was never wired into the
tool's CLI. --bucket is now per-backend optional (required for
aws/gcs), and the default cache dir is scoped by collection/shard for
uio-grpc, whose mirror has no distinguishing key prefix.

--live-reload-key runs the same watch loop as --live-reload with each
reload triggered by pressing Enter instead of a timer — easier when
stepping through a debug scenario. Timer mode is unchanged; the two
flags are mutually exclusive. Closed stdin ends the loop gracefully,
so the flag cannot busy-loop on piped input (and is documented as
incompatible with @- stdin request arguments).

Verified end-to-end against a live instance started with
QDRANT__FEATURE_FLAGS__WRITE_SEGMENT_MANIFEST=true: scroll over
uio-grpc returns all points with payloads, timer mode picks up an
upsert + delete as +/- diff lines, and key mode fires one reload per
Enter and exits cleanly on EOF.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Log StorageRead not-found gRPC failures at debug level

A read-only follower routinely probes files the writer creates lazily
(e.g. the mutable id tracker's mappings and versions before the first
flush), so every uio-grpc follower poll spammed INFO logs like:

  gRPC /qdrant.StorageRead/FileLength failed with NotFound
  "File not found: .../mutable_id_tracker.versions"

NotFound on /qdrant.StorageRead/* now logs at debug; NotFound on all
other services (missing collection etc.) stays at info. The service is
matched by parsing the path's service component and comparing it to
the tonic-generated storage_read_server::SERVICE_NAME constant rather
than a hard-coded path string.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 16:04:54 +02:00
Andrey VasnetsovandClaude Fable 5 a43abbd3b8 Add --live-reload watch mode to edge-shard-query (#9791)
With --live-reload <SECONDS> the tool keeps running after the first
answer: every interval it refreshes the ReadOnlyEdgeShard from object
storage, re-runs the same request, and prints the difference against
the previous results — `+` id appeared, `-` id disappeared,
`~ old -> new` for changed content (payload/vector/score/version).
Each cycle prints a summary line; pure reordering prints nothing.

The request is parsed once into a PreparedRequest (ScrollRequest /
SearchRequest are Clone), so every iteration re-runs the identical
request and the filter/vector parsing and logging no longer repeat.
The first run prints the full result set in the existing format.

A failed refresh is logged and retried next tick — the shard keeps
serving its previous state, so a transiently unreachable bucket does
not kill the watch loop. Status lines go to stderr via log, diff rows
to stdout, keeping stdout machine-consumable.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 10:47:35 +02:00
qdrant-cloud-botandCursor d773266c55 Fix edge shard exact count double-counting across segments (#9790)
Collect per-segment point ids into a BTreeSet before counting so points
visible in multiple segments (e.g. during proxy/merge optimization) are
counted once, matching the canonical collection count path.

Fixes #9789

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-11 09:59:03 +02:00
Arnaud GourlayandClaude Opus 4.8 35bbf0487a Remove unnecessary clippy allow attributes (#9775)
Remove 8 `#[allow(clippy::...)]` attributes that no longer suppress any
lint. Each was verified redundant by rewriting it to `#[expect(...)]` and
confirming the workspace stays clippy-clean under the CI config
(`cargo clippy --workspace --all-targets --all-features -- -D warnings`).

Attribute-only deletions, no behavior change.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 11:47:59 +02:00
Daniel Boros 66ce17e709 feat: skip unloadable segments in read-only follower open (#9762) 2026-07-09 18:13:25 +02:00
Andrey VasnetsovandClaude Opus 4.8 fb681b7d9d Lazy roaring flags bitmap and bool index counts (#9749)
* [AI] make ReadOnlyRoaringFlags bitmap and bool index counts lazy

Opening a read-only segment scanned every flags file end to end:
`ReadOnlyRoaringFlags::open` materialized the whole RoaringBitmap via
`iter_ones()`. Every payload field carries a null index, so this was paid
per field per segment, for bitmaps most queries never touch.

Make the bitmap a `OnceLock`, filled by a scan on first access. Open now
reads only the tiny status file. `ReadOnlyBoolIndex`'s three eager count
fields collapse into one lazily-derived, cached `BoolCounts`; its
`live_reload` refreshes them in place when present and leaves them unset
otherwise, so reloading an index nothing queries stays scan-free.

Propagate the resulting `OperationResult` through `RoaringFlagsRead`,
`PayloadFieldIndexRead::count_indexed_points`, `FieldIndexRead`,
`PayloadIndexRead::{indexed_points, get_telemetry_data}`, `build_info` /
`build_telemetry` and `SegmentEntry::{info, get_telemetry_data}`, out
into shard, edge and collection.

`ram_usage_bytes` stays infallible: an unmaterialized bitmap holds no
RAM, so it reports 0 via the new `bitmap_if_materialized`.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* [AI] correct `preopen` comment: `open` no longer scans the flags file

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* [AI] fix edge examples for fallible `info()`

`EdgeShardRead::info` now returns `OperationResult<ShardInfo>`. The
examples live in their own workspace (lib/edge/publish), so the main
`cargo check --workspace` never saw them.

Every call site sits in `fn main() -> Result<(), Box<dyn Error>>`, so
propagate with `?`. `bm25-search` compiled either way but would have
printed the `Result` rather than the `ShardInfo`.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-09 17:48:40 +02:00
Arnaud GourlayandClaude Fable 5 bd35dc1189 Replace flush_all sync bool with FlushMode enum (#9750)
* refactor: replace flush_all sync bool with FlushMode enum

SegmentHolder::flush_all took two adjacent bools (sync, force), and call
sites read as bare literal pairs like flush_all(true, false). Swapping
the arguments compiles and silently changes flush semantics: a swapped
pair at the snapshot site would make snapshots skip flushing entirely
when a background flush is running.

Introduce FlushMode { Sync, Background } for the first parameter so the
pair is no longer transposable and the behavior is named at each call
site. The force flag stays a bool since it feeds the
SegmentEntry::flusher(force) trait in lib/segment. No behavior change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: exhaustive match on FlushMode instead of equality check

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 16:34:52 +02:00
Daniel Boros 78688733cb feat: support S3 Express One Zone buckets in the S3 bridge (#9757) 2026-07-09 15:02:28 +02:00
Andrey VasnetsovandClaude Opus 4.8 02cacfe192 debug logging of the s3 connection (#9597)
* debug logging of the s3 connection

* refactor(io_bridge): make S3 latency logs opt-in and low-overhead

Move the blob-backend latency traces onto a dedicated `io_bridge::latency`
log target at `trace` level, so they are silent by default and can be
toggled as one group at runtime without a rebuild (e.g.
`RUST_LOG=io_bridge::latency=trace`). Guard the timing `Instant::now()`
behind `log_enabled!` so there is no overhead when the target is disabled.

Also add `list_files` timing, and switch the shard_query CLI logger to
millisecond timestamp resolution.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 18:16:16 +02:00
Andrey VasnetsovandClaude Opus 4.8 8311ab62a9 feat(edge): add --search-threads to shard_query tool (#9736)
Add a `--search-threads` option to the edge-shard-query tool so the
number of threads in the shard's search thread pool can be specified.
When set, it builds an EdgeConfig with `max_search_threads` and passes
it to `ReadOnlyEdgeShard::open`, overriding the CPU-derived default used
for both parallel segment reads at open and running searches.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 14:24:53 +02:00
ab0d3ecc62 Add unified memory: cold|cached|pinned placement parameter for collection components (#9684)
* Add unified `memory: cold|cached|pinned` placement parameter for collection components

Introduce a single `memory` parameter that controls how each collection
component's data is held in RAM, replacing the inconsistent zoo of
`on_disk` / `always_ram` / `on_disk_payload` flags:

- `cold`: not pre-loaded from disk, cached with usage
- `cached`: pre-populated into page cache on load, evictable under pressure
- `pinned`: materialized on heap, never evicted by cache pressure

The parameter is available on dense vectors, HNSW config, all quantization
configs, the sparse index, all payload field index types, and payload
storage (as a new `payload: { memory }` sub-object on collection params).
When set, it overrides the deprecated legacy flag; when unset, behavior is
unchanged. Legacy flags are marked deprecated (Rust + proto) but keep
working; conflicts are resolved in favor of `memory` with a warning.

New capabilities enabled by the tri-state model:
- HNSW graph links can be pinned (first production caller of the existing
  `GraphLinksResidency::Pinned`)
- sparse mmap index, quantized vectors and on-disk payload field indexes
  gain a `cached` tier (mmap + populate on open)

`pinned` is rejected by API validation for components without a heap
variant (dense vector storage, payload storage). Low-memory mode degrades
placements at load time via `Memory::clamp_to_low_memory`, matching the
existing `prefer_disk`/`skip_populate` behavior. Effective-placement
comparison in the config-mismatch optimizer avoids spurious rebuilds when
the same placement is expressed through the new parameter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix gpu-gated tests for the new `memory` field

CI clippy runs with --all-features, which compiles the gpu-gated tests
that were missed locally: add the `memory` field to config literals and
allow deprecated placement params, same as in the rest of the tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Add OpenAPI tests for memory placement, keep sparse config downgrade-clean

- OpenAPI tests: create/update collections with `memory` on every component,
  assert the parameters are echoed in collection info, assert legacy-only
  collections expose no new fields, and assert `pinned` is rejected (422)
  for dense vector storage and payload storage on both create and update.
- Persist only the explicitly requested `memory` parameter in
  `sparse_index_config.json` instead of the legacy-resolved placement, so
  configurations using only the deprecated `on_disk` flag keep byte-identical
  files that older Qdrant versions load without unknown fields.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Validate collection meta ops at construction, not only in the API layer

The `memory: pinned` rejection for dense vectors and payload storage
lived in `Validate` impls on the internal request types, which only ran
through the REST actix extractor. gRPC validates just the proto message,
so a gRPC client could persist `pinned` where it is not supported and
have it silently treated as `cached`.

Run the derived validation in `CreateCollectionOperation::new` and
`UpdateCollectionOperation::new` instead: the constructors are the
common chokepoint for all API paths, before the operation is proposed
to consensus. This covers every validator on these types, not just the
`memory` checks, and keeps consensus-apply unaffected so mixed-version
clusters never reject already-committed operations.

`UpdateCollectionOperation::new` becomes fallible; `remove_replica` now
uses `new_empty` since it carries no user config. Regression tests drive
the gRPC conversion path and assert `InvalidArgument` for `pinned` on
create and update, with `cold`/`cached` accepted as a control.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 14:32:51 +02:00