* Keep the on-disk text index token total in the postings header
The statistics gather asks every segment for `total_tokens` on each query,
and the on-disk index answered by reading its whole length sidecar and
masking it. The build already knows the total, so it goes into the reserved
bytes of `PostingsHeader`, which `open` reads anyway.
Deletions since the build are not subtracted, the same way `posting_len`
keeps them. With every implementation now a field read, `total_tokens`
drops its `OperationResult` and hardware counter.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Note the placement difference in the text index token total
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Report no text index token total for a header that predates it
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Arnaud Gourlay <arnaud.gourlay@gmail.com>
* Gather text corpus statistics across segments
BM25 needs `df`, `N` and `avgdl` over the collection, not over a segment:
computed per segment, a document's score would change whenever the optimizer
moved it. The pre-pass that already collects those for sparse vectors does
nothing sparse-specific, so text reuses the carrier and the walk and brings
its own map.
- `QueryContext` gains `text_stats` beside `idf_stats`, keyed by payload
field and by term string, since a `TokenId` is local to the segment that
assigned it
- the gather runs through `PayloadIndexRead` and `FieldIndexRead` into
`FullTextIndexRead`, which grows `posting_len` for the local `df`
- `TextQueryContext` hands out `idf(term)` and `avg_doc_len()`, both over
the summed totals, with IDF clamped at zero because posting lists keep
deleted documents while `points_count` does not
- `total_tokens` goes absent as soon as one segment records no lengths,
rather than averaging over the part of the corpus that does
- `IdfScopeStats.idf` renamed to `df`, which is what it holds
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Address review of the text statistics gather
- seeded terms must already be tokenized, since resolution is a bare
vocabulary lookup: an untokenized term misses everywhere and keeps a `df`
of zero, which is the largest IDF the formula produces. Stated on both
entry points and checked by a debug assert
- `posting_len` documents what each backend actually counts. The mutable
index drops deleted points from its postings while the other two keep
them, so `df` can exceed `N` and can change when the optimizer runs
- the statistics are shard-local, not collection-wide: the gather runs per
shard and nothing merges across shards
- `idf()` says which end of the range an unreported term sits at, which is
the top, not the bottom
- skip the whole-sidecar read once the corpus total is already absent
- thread the stop flag through the gather and check it per term
- bill the on-disk posting-length read, one per query term per segment
- resolve terms without cloning them, passing the destination slot as the
callback's user data the way `resolve_tokens` does
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Address the second review round of the text statistics gather
- check the stop flag once at the top of the gather, so an empty term list
and the whole-sidecar total read on disk are both covered
- explicit `From` conversions in the integration test
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Verify the gathered average length through a segment
`TextIndexParams::scoring()` is a const `false`, so nothing built through a
segment recorded lengths and the integration test could only assert that
the average was absent. `override_scoring` is a thread-local `testing`
override returning a guard, so a test can build a recording index through
the constructors production uses. The test now asserts the summed average
over two segments, and a separate one keeps the absent case.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Trim the text statistics gather to its contract
Revert the idf to df rename on IdfScopeStats, which is sparse code and
nothing in this stack depends on it. Cut the doc comments that explained
design rationale rather than behavior, state the tokenization requirement
once, and drop the test that asserted a lookup in an empty map is None.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Cut the text statistics gather down to what it needs
- drop the integration test for absent lengths, already covered by the
unit test's second half
- drop TextFieldStats::seeded, used only by that unit test
- trim doc comments that explained design choices rather than behavior
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Make the uio trace visualizer task-agnostic
Lanes now come from the file tree of the traced request paths instead of
a hard-coded path classifier and colour table. Tree nodes expand and
collapse from a side panel or the lane headers (shift for the whole
subtree); files with reported sections split into one child per section.
Order follows first appearance in the trace, sections keep their layout
order, and colours are generated per node.
Requests can be filtered by path (substring or glob) and by operation;
the concurrency and bandwidth columns follow the filter. Phases, marks
and CPU usage render only when the trace has them, and the CPU bin
follows the trace's own sampling interval.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* uio trace visualizer: manual colours, dimming filter, reveal on click
Tree nodes start uncoloured; clicking a node's swatch toggles a random
colour, right-clicking the row picks one from a palette. Requests take the
colour of the deepest coloured node on their path.
The path filter greys out unmatched requests instead of removing them, so
lanes stay put; op filters still hide. Series, totals and tree counts follow
the path filter.
Clicking a request opens its path in the tree and highlights its node.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Each push to a PR branch started a full CI run while the previous one kept
going. Group PR-triggered workflows by PR number and cancel the in-progress
run. Pushes to dev/master use the run id as the group, so they are never
cancelled.
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A compacted tracker holds every mapping on the heap and only reads its file on
open, so the payload storage's report listed its files but not that RAM.
CompactedTracker::ram_usage_bytes is passed up through TrackerEnum, Logstore
and Blobstore::ram_usage_bytes to PayloadStorageImpl, whose memory reporter now
reports it as extra RAM. Append-only trackers and mutable mode report 0: they
read their mappings from the files.
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* [AI] Refactor metrics rendering into private modules
* Update metrics allow-list path in DEVELOPMENT.md
Point contributors at requests.rs after the metrics module split.
---------
Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com>
Behind the new compact_logstore_tracker feature flag, implied by
serverless_compatible. SegmentBuilder calls PayloadStorageEnum::make_immutable
before its flush of a non-appendable segment's payload storage.
Blobstore::make_immutable switches an append-only storage to the compacted
tracker in memory, from the mappings in RAM plus whatever part was flushed, and
removes the append-only tracker file; the next flush saves the compacted one.
The mutable mode is left as is. Sparse vector storages keep the append-only
tracker: plain nearest queries never read them.
CompactedTracker::from_tracker no longer writes, and the tracker keeps its fs
as an object-safe TrackerFs instead of a boxed closure.
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Logstore holds a TrackerEnum, so an indexed segment whose tracker was
compacted opens through the same writable path as any other storage. New
storages still start with the append-only tracker.
A compacted tracker opened writable keeps a type-erased saver over the fs it
was given and rewrites its whole file on flush. The flusher takes the target
point offset count the Logstore flusher recorded, so mappings of puts made
after the flusher was created are not persisted ahead of their value data.
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Read Logstore through either tracker format with TrackerEnum
TrackerEnum holds an AppendOnlyTracker or a CompactedTracker, picked by which
file is on disk, and implements TrackerRead. LogstoreReader, BlobstoreReader
and BlobstoreView use it; the writable Logstore is unchanged. A compacted
storage is complete, so live reload and preload are no-ops for it.
CompactedTracker gains exists/preopen and always opens populated, so a preopen
fetches the whole file. from_tracker converts any tracker into the compacted
form, used by the tests and by an ignored test measuring a real tracker file.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Adapt TrackerEnum reader to the dev UIO and IO stats changes
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Allow in-place enable_hnsw updates on payload indexes
Refactor schema transitions to Compatible(CompatibleDiff) so on_disk and
enable_hnsw flips (including combinations) reuse existing index files
instead of dropping and rebuilding from payload.
* Format enable_hnsw in-place schema update changes
* Keep live index handles on enable_hnsw-only schema change
Re-opening appendable field index files for a metadata-only change
dropped updates not yet flushed by the live handles. Persist the new
schema in drop_index_if_incompatible instead and report AlreadyBuilt.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Replace compatible! macro in schema transition with flag-clearing helper
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Edge facet always merged per-segment counts, ignoring `exact`. Mirror
the local shard's exact facet: collect unique values across segments,
then count each value with an exact, deduplicated count.
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Add CompactedTracker, a RAM-resident tracker with a compact file
Second step towards an immutable tracker for Logstore, on top of the
unified TrackerRead trait. CompactedTracker keeps every mapping in RAM
and implements TrackerRead from there, so a reader decodes the file once
on open and never touches the disk for lookups.
Flushing rewrites the whole file from a snapshot taken when the flusher is
created: the mappings are varint delta encoded against the end of the
previous value (zero for back-to-back values, so the runs compress well)
and LZ4 compressed behind a small header. The file is replaced with
atomic_save, and a stale flusher whose snapshot holds no more mappings
than the file already has is a no-op, so out-of-order flushes never roll
the file back.
Decoding validates every byte: truncated or trailing data, bad magic or
version, and deltas that do not add up to a valid pointer are rejected.
The type is not wired into Logstore or LogstoreReader yet, that is the
next step in the stack.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Split CompactedTracker into a module
mod.rs keeps the struct with its lifecycle and write side, format.rs
holds the file layout and the encode/decode functions, read.rs holds the
TrackerRead impl, and tests.rs the tests. No functional change.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Track dirtiness in CompactedTracker with a flag
set accepts any point offset and replaces existing mappings, so the
mapping count no longer identifies the state. A dirty flag replaces the
persisted count: set marks the tracker dirty, a flusher swaps the flag
off and takes a copy of the pointers, and a clean tracker gets a flusher
that does nothing. A failed write marks the tracker dirty again so the
next flush retries. The stale-flush guard is gone, flushers of one
storage are serialized by the segment flush lock.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Store CompactedTracker mappings as columns instead of LZ4 varints
Gaps, block-bitpacked lengths (min + width per 128 lengths), and a list of
pointers that do not start where the previous value ends. On a real payload
tracker (105k mappings) this is 1.14 bytes per mapping instead of 2.05, 14x
smaller than the flat tracker file, against an entropy floor of ~1 byte.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Document a worked example of the CompactedTracker format
Six point offsets with a gap and a page rollover, their three columns, the 32
serialized bytes and how one pointer decodes. test_documented_example pins the
documented bytes to the encoder.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* zstd-based format (#10787)
* Fix CompactedTracker after UniversalWriteFileOps rename
The stacked merge brought in UniversalWriteFs; update CompactedTracker
bounds so the crate compiles again.
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: xzfc <5121426+xzfc@users.noreply.github.com>
Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com>
* Expose document lengths from the text index read surface
`doc_len` and `total_tokens` on `FullTextIndexRead`, backed by the
`InvertedIndex` accessors deferred from the two previous steps. Nothing
reads them yet, so the recording gate still keeps every answer absent.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Fix up the document length accessors after review
- drop the bounds check in read_point_to_doc_len: the deletion mask the
caller already consults establishes the bounds, and `len()` is an fstat on
io_uring and a blocking HEAD on object storage, once per scored point
- pin the deletion mask in the on-disk total: the agreement test cannot
catch it, since its deletions predate `create` and leave zeroes on disk
- correct three doc comments. `total_tokens` is not the converse of
`doc_len`, `remove` is exactly the invalidation hook the comments claimed
did not exist, and "live points" is wrong under append-only deletion
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Make the document length accessors agree and bill their reads
- `doc_len` answers the same for the same data on every backend. The
on-disk one now bounds by the point space `open` already computes and
reports `Some(0)` for a point it holds no tokens for, instead of letting
the storage placement decide between `Some(0)` and `None`
- both accessors take a `HardwareCounterCell`, so the per-point reads and
the whole-sidecar total are measured like every other IO on the trait
- `ImmutableInvertedIndex` maintains `total_tokens` rather than summing on
every call: `From<MutableInvertedIndex>` already has the counter, the
on-disk conversion already walks every element to mask it, and `remove`
is the only mutation afterwards
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* State the settled population rule on the total tokens accessor
The comment described the interim disagreement between the in-RAM and
on-disk counts, which #10675 removes. Describe the rule instead: both
counts cover the documents still held that carry at least one indexed
token, so the ratio does not depend on the storage placement.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Read document lengths only in batches
A per-point doc_len invites one read per point, a round trip each when
the on-disk index sits on a slow or remote disk. Both InvertedIndex and
FullTextIndexRead now expose doc_len_batch alone: the in-RAM backends
answer from their vector, and the on-disk one answers out-of-range and
inactive points without IO and sends the remaining reads through a
single read_batch.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Meter remote requests and expose IO statistics to edge shard users
Generalize the disk cache statistics from #10637 into a shared OpStats
primitive and add a second layer counting every remote request issued
through BlobFs and BlobFile, reads and writes alike. One helper at each
call site drives the uio_trace request and the stats guard together, so
the trace also gains the write operations.
CachedBlobFs exposes both halves through stats(), DiskCacheFs exposes
its remote filesystem, and both edge shards expose their filesystem, so
an application can read cumulative totals. shard_update logs the IO
counters moved by each phase; shard_query prints the combined dump.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Drop DiskCacheStats in favor of remote request stats
Every disk-cache fetch is exactly one BlobFile read already metered by
RemoteIoStats, so the cache half duplicated the remote Read/ReadFrom
counters. CachedBlobFs::stats() now returns RemoteIoStats directly.
The remote stats dump now includes per-op latency histograms, which
were previously only printed for cache fetches.
edge-shard-query and edge-shard-update initialize serverless_compatible
feature flags instead of running with uninitialized ones.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Sample uio-trace CPU usage every 10ms instead of 2ms
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Address review: not-found bucket, empty-tail reads, strum Op
- OpStats counts "not found" answers separately from errors, and the
trace records them as a `not_found` outcome, so length and existence
probes of absent objects no longer read as failures.
- A read_from tail at or past EOF is settled after the len probe and
counts as a completed empty read; a not-found error skips the probe,
which could only repeat the answer.
- Op derives strum EnumCount/EnumIter/IntoStaticStr instead of a
hand-maintained ALL array.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
`PlannedQuery` fetches `limit + offset` points and leaves cutting the
offset to the caller, which the server does at the collection level.
Edge never did, so `query()` returned the first `offset + limit` hits.
`query_groups` returned hits with only the `group_by` field as payload,
ignoring the requested `with_payload`/`with_vector`. Hydrate the final
hits once after grouping, as the collection-level group-by does.
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* list files concurrently
Read-only shard load took each segment's LIST snapshot one after another,
so cold start paid N sequential object-store round-trips. Take all
segments' snapshots concurrently via a new `list_files_async` on
`UniversalReadFsAsync`, then stage each open over its ready `CachedFs`.
- `list_files_async` has no default, so no backend silently falls back to
a blocking LIST; `BlobFs` spawns it on the `BridgeRuntime` and shares
the traced, latency-logged future with the sync `list_files`.
- `ReadOnlySegment::build_cached_fs_async` + `schedule_open_with_cached_fs`
split `schedule_open` so the LIST can be awaited separately.
- `CachedFs::schedule_open` polls once with `now_or_never` instead of a
nested `block_on`, which would panic under the outer executor.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* use async LIST in live-reload
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Merge UniversalReadFileOps into UniversalReadFs
Every filesystem implemented both traits, so fold list/exists/from_context
into UniversalReadFs and make UniversalWriteFileOps extend it directly.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* rename `UniversalWriteFileOps` -> `UniversalWriteFs`
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Both tool binaries were single main.rs files of about 1100 and 840
lines. Move their items into per-concern modules without changing
behavior: CLI definitions, argument parsing, backend configs, the
prepared request and reporting for shard_query; CLI, backend configs,
schema reading, point generation, dry run and apply for shard_update.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
With `enforce_internal_auth` enabled, the p2p server reused the public
auth layer. A read-only key or any JWT passed that layer, and the Raft
service does not check per-request access, so those credentials could
add a peer to consensus.
The internal auth layer now runs in a dedicated scope that accepts only
`api_key` and `alt_api_key`. The read-only key, JWTs, and missing or
wrong keys are rejected with Unauthenticated. Warn at startup when
internal auth is enforced with only a read-only key configured, since
peers would have nothing to authenticate with.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Harness failures print their seed, but the seeded tests cannot be re-run with it and the
soak binary needs hand-built flags, so a CI failure was effectively one-shot:
- Honour MODEL_TESTING_SEED so the seed from a failing run replays the same op sequence
in the same test; without it a fresh seed is still drawn.
- Print the equivalent model_testing invocation for every run, so any failure (including
the Linux-only harness tests on a non-Linux machine) becomes a one-line reproduction.
- Bind the harness knobs once in smoke() and feed the same bindings to run() and to the
printed command, so a run and its reproduction cannot disagree about them.
- Cover both helpers with platform-independent unit tests: the helpers now compile and are
tested on every platform, not only inside the Linux-gated harness module.
Refs #10406, #10662, #10667.
* Remove shard initializing flag under the local shard lock
A late snapshot recovery from an aborted transfer and a new incoming
transfer could race on the shard initializing flag: initiate_receiving
removed the flag after releasing the local shard lock, deleting the flag
a concurrent restore_local_replica_from had just created. The restore
then failed with ENOENT on its own flag removal, and a crash in between
would have left half-moved shard data without a dirty marker.
Move the removal into init_empty_local_shard, under the local write lock
and after the empty shard is built, tolerating a missing flag. This also
fixes the try_exists(..).is_ok() check, which was true for absent files.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Install empty local shard only after removing the initializing flag
If flag removal fails or the future is cancelled while it is pending,
local keeps the dummy placeholder instead of a Local shard next to a
leftover dirty marker, so a retried transfer re-initializes the shard
and the documented cancel-safety holds.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
One TrackerRead trait, in lib/blobstore/src/tracker/read.rs, now covers
the Gridstore Tracker and ReadOnlyTracker and the Logstore
AppendOnlyTracker. LogstoreView, LogstoreReader, and validate_consistency
are generic over it, as GridstoreView already was over the Gridstore-only
trait.
The trait gains get_range and an access pattern on get, and its iter
returns impl Iterator instead of the concrete Gridstore Iter, which drops
the storage type parameter from the trait. The Gridstore trackers
implement get_range through a new read_slots helper; nothing in Gridstore
calls it yet. The lifecycle methods of LogstoreReader (open, preopen,
files, live preload, live reload, clear cache) stay in an impl block bound
to AppendOnlyTracker, so a future immutable in-RAM tracker plugs into the
read path without inheriting reload semantics.
No runtime change: Gridstore slot reads keep the Random access pattern,
and PointerItem stays the iterator item type.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* Add a BM25-over-sparse baseline benchmark
The text payload index is meant to score about as fast as BM25 over sparse
vectors, so the number it has to match needs to exist before the scorer
does. Measures a local shard end to end, no HTTP.
- embeds the corpus through `lib/bm25` with its defaults, so the baseline
is the route a user migrates from rather than a reimplementation
- Zipf-like vocabulary. On a uniform one every term is equally selective,
IDF is flat and pruning has nothing to prune, which would flatter any
scorer measured against it
- two shards rather than one shard before and after optimization: a shard
that will optimize starts as soon as the upsert lands, so the first cut
timed a half-converted index and called it fresh. Both states assert
what they hold before anything is timed
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Measure the sparse BM25 baseline where the routes separate
- 200k documents by default (BM25_SPARSE_DOCS overrides): at 20k every shape
of both routes measures the same and half of a shard-level query is the
shard; reachable through the shard since #10682
- a third shard with the sparse index on disk, which the optimized state never
exercised
- recall at 10 against BM25 by definition, printed per state: the default
avg_len of 256 on a corpus averaging 110 tokens misses a quarter of the true
top 10, so the optimized state is also timed with the corpus average
- corpus, queries, reference and recall move to segment::fixtures::bm25_corpus,
to be shared with the text-index bench and the comparison harness
- module doc: the sparse shapes measure within a few percent of each other; the
states exist for the text comparison
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Refuse an empty corpus in the sparse BM25 baseline
BM25_SPARSE_DOCS=0 built empty shards, made the average length NaN and
scored every empty truth as recall 1.0.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Measure what the sparse BM25 baseline states claim
The fresh shard kept the default 10 MB indexing threshold and optimized
itself in the background, so it was timed as a second optimized state.
Disable indexing and re-check it after timing. Keep the shard storage
under CARGO_TARGET_TMPDIR so the on-disk index is not read from a tmpfs.
Fix doc comments that named missing files, a no-op IDF clamp and the
wrong reason for the empty-corpus guard.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* AI: Share reads across different diskcache pipelines
* AI: Unify scheduled reads into single ScheduledRead struct
* AI: Implement follower promotion on fetch abandonment
* AI: self-promote if waiting for an abandoned placeholder
* AI: simplify
* AI: use Weak<Placeholder> instead of manually counting strong refs
* AI: track abandoned waiters
* AI + manual: park with timeout
* AI: fix barrier in test
* fmt with latest nightly
* expect remote pipeline
* Carry document length into the immutable and on-disk text indexes
The mutable index records `doc_len`. The immutable and on-disk backends dropped
it on conversion. They now carry it and persist it as a sidecar, so a length
survives an optimizer run and a restart.
- `ImmutableInvertedIndex` gains `point_to_doc_len`, parallel to
`point_to_tokens_count` and zeroed wherever that vector is, so summing it
never counts a deleted document.
- `OnDiskInvertedIndex` writes `point_to_doc_len.dat`, only when the index
records lengths. `files()` lists it only when it exists, so a snapshot carries
it exactly when there is one.
- Deletions are masked on load, not at build time. The file is written once and
a point deleted later through the id-tracker is zeroed in
`TryFrom<&OnDiskInvertedIndex>`. No segment total is stored, since it could
only be summed after that masking.
- A sidecar shorter than the counts is treated as absent. It can only come from
a partially copied file set, and padding it would give every point past the
truncation a zero that reads like a real length.
- A missing sidecar on an index that should have one makes `new_mmap` report the
index absent, which routes into the existing rebuild from payload. The check
lives there rather than in `OnDiskInvertedIndex::open` because the read-only
stack never builds and would drop the field instead.
- `FullTextMmapIndexBuilder::add_many` reached below `index_str_tokens` and so
declared no length at all. It now tokenizes through the shared helper and
measures like the gridstore path.
Recording is still gated behind `TextIndexParams::scoring()`, so none of this is
written today.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Wipe the text index directory instead of its listed files
`files()` lists the `doc_len` sidecar only when the index loaded it, so a
truncated sidecar is omitted. `wipe` then deleted the listed files, failed to
remove the now non-empty directory, discarded that error and returned success,
leaving the directory and a stale sidecar behind.
Remove the directory itself. Each field owns its own `{field}-text` directory,
so this no longer depends on `files()` being an accurate inventory, and real
removal errors propagate rather than being swallowed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Drop the fully qualified visibility paths in the text index
`pub(in crate::index::field_index::full_text_index)` is tedious to read and to
keep correct, and it grants nothing here: `inverted_index` and
`mutable_text_index` are private modules of `full_text_index`, so their types
are not nameable from outside it whatever the field visibility says. Plain
`pub` is no wider in practice.
Widening `Storage` surfaced two `private_interfaces` warnings, since
`ZerocopyPostingValue` and `PostingListHeader` in `types.rs` were still behind
the long path; those move too.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Harden the document length sidecar after review
- check for the sidecar before opening the index rather than after. The
open populates the whole file set, so the first start after scoring is
enabled would fault in every segment's postings only to discard them
- treat a sidecar that covers more points than the index as untrustworthy,
not just one that covers fewer, and warn with both counts the way
`SortedBlockIndex::open` does. A longer one used to be accepted and then
silently truncated when materialized
- unlink a stale sidecar when a build records no lengths. It was the only
file here that could outlive the build that wrote it, and `open` would
have read it as this build's
- say why an index is being rebuilt from payload instead of leaving a
silent full re-index announced at debug level
- read `phrase_matching` from the config in the mmap builder, like the
other two callers of the same helper pair, so the two halves of the
sentinel rule cannot drift apart
- assert the length invariant the writers actually maintain, and correct
what `files()` and the field comment claim
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Address review on the document length sidecar
- shrink the document length vector when it is handed over to the immutable index, it arrives with the mutable index's doubling capacity
- narrow the immutable index fields to pub(super), with a test-only accessor for the one reader outside the module
- let wipe propagate a missing index directory instead of treating it as success, nothing reaches it with the directory already gone
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Configurable key prefix for object storage snapshots
Snapshot object keys were always derived from the local snapshots path,
so every deployment sharing a bucket wrote under the same `snapshots/`
root. Each cloud config block gains an optional `prefix`, and objects
become `<prefix>/snapshots/...`.
The prefix is applied by wrapping the client in `PrefixStore`, so the
snapshot operations and the names returned by the API are unchanged.
Leading, trailing and repeated slashes are dropped, and an empty prefix
is a no-op.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Test that snapshot operations cannot escape the storage prefix
Hostile targets are handed to the cloud manager directly, past
`validate_snapshot_name`: parent references, absolute paths, encoded
slashes, backslashes and empty paths. Writes must stay under the prefix,
and objects planted outside the prefix must be invisible to list,
download, stream and delete.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Normalize empty prefix components
Refactor prefix handling to remove empty components.
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tim Visée <tim+github@visee.me>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* Add `match: { substring }` filter condition
Unindexed `text` and `text_any` matching became token-aware in #10341 and
#10593. Users who relied on the old raw substring behaviour get it back as
an explicit condition: `match: { "substring": "..." }` selects points with a
string value containing the given string, byte-wise and case-sensitive,
consistent with exact keyword and prefix matching.
Execution: a keyword index (with or without the `prefix` option) serves the
condition by scanning its value dictionary and uniting the postings of the
matching keys; cardinality reuses the prefix estimator, generalised into
`keys_union_cardinality`. The per-point checker goes through the forward
index. Without a keyword index the condition falls back to reading the
payload. Text, bool, integer and uuid indexes decline it.
Strict mode: the condition requires the `KeywordMatch` capability, so with
`unindexed_filtering_retrieve: false` it is rejected on unindexed and on
text-indexed fields and allowed on any keyword index.
API: `MatchSubstring` in the REST `Match` union with regenerated OpenAPI,
gRPC `Match.substring = 12`, edge python `MatchSubstring`, edge ffi
`Match::Substring`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Test substring fallback on a text-indexed field, document estimator params
A text index cannot serve `substring`, so on a field that has only a text
index the condition runs through the payload fallback; only strict mode may
reject it. Pin that in the OpenAPI suite and reword the strict-mode unit
test comment, which read as if the text index itself blocked the query.
Also spell out what `keys` and `postings` mean in `keys_union_cardinality`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* Serve `match: { substring }` from the keyword key dictionary
The condition used to enumerate keys through `MapIndexRead::for_each_value`,
which on the on-disk variant drags the whole `value_to_points` file through
`for_each_entry`, plus one random read per matching key for its postings
count. Query planning paid that scan in full, before deciding whether to use
the clause at all.
Route it through the `prefix_index.bin` key dictionary instead: front-coded
keys with their postings counts inline, no postings. Estimation now reads
keys only and never touches `value_to_points`; filtering takes the matched
key list and resolves postings in one batched read, as prefix matching
already does.
This makes the `prefix` option a requirement: a keyword index without it has
no key dictionary, so it declines the condition and falls back to the payload
scan, the same as a text index. Strict mode follows — substring now infers
`KeywordPrefix`, so `unindexed_filtering_retrieve: false` names
`keyword (with prefix: true)` as the index to create.
`PrefixIndex::for_each_key` reads blocks in ~1 MiB chunks rather than the
whole key section at once: a substring cannot be pruned by the block index,
so the one-shot read would grow with the dictionary.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Sync MatchSubstring OpenAPI description with Rust docs
After routing substring matching through the prefix key dictionary,
the schema docstring required regenerating so OpenAPI stays consistent.
* Scan the map index keys for substring match without a dictionary
Without the keyword dictionary a substring condition was declined by the
field index and left to the per-point condition checker, which reads the
forward index for every candidate point. Enumerate the distinct keys of
`values_to_points` instead: the same one-pass-over-distinct-values shape as
the dictionary scan, only over a structure that interleaves keys with their
postings. Filtering and cardinality estimation are then always served, so
the condition can act as a primary clause on a plain keyword index.
Prefix matching keeps its per-point fallback: an ordered dictionary is what
makes a prefix a bounded range, and enumerating every key to answer one is
not a trade worth making implicitly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Reject substring matching in strict mode
A substring condition is answered by looking at every distinct value of the
field: no index gives it a bounded access path, so there is no index a user
could create to make it affordable. Reject it under strict mode instead,
wherever a filter reaches verification — read and write filters, nested
sub-filters, and prefetch filters.
Filter limits are now checked before the unindexed-field check, so the
rejection is not reported as "create an index for this key", advice that
would lead to the same rejection afterwards.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Update the strict mode substring test to the new rejection
The test asserted that substring filtering under strict mode asks for a
keyword index with the `prefix` option. It is now rejected whatever index
the field carries, so every case in the test gets the same answer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Estimate a substring condition without scanning
Counting the keys a substring matches costs the same scan as answering the
condition, and `filter` then repeats it to collect those keys. Report the
uninformed estimate instead — the one an unindexed condition has always
reported — and keep the primary clause, so the scan happens once, in
`filter`, and only when the planner picks the condition to drive iteration.
With no counts to collect, `substring_scan` collapses into `substring_keys`:
the in-RAM variants no longer look up a posting count per matched key.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Don't to parse everything as UTF-8
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com>
Co-authored-by: timvisee <tim@visee.me>
Snapshot storage was limited to S3 although the object_store dependency
already ships the GCS and Azure backends. `snapshots_storage` now accepts
`gcs` (alias `gcp`) and `azure`, configured through `gcs_config` and
`azure_config` blocks next to the existing `s3_config`. The legacy S3
shape is unchanged.
Client construction is split into one builder per backend, all sharing
the Qdrant user agent and the plain-HTTP rule for `http://` endpoints.
Startup warns when a config block for an unselected backend is present.
The e2e snapshot recovery test is parameterized over the cloud backends.
The GCS case is skipped because fake-gcs-server does not implement the
XML multipart upload API that object_store uses for GCS uploads.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
A value that tokenizes to nothing is indexed and matches nothing. The mmap
build already treats it as no document, folding it into the "no tokens"
mask, while the mutable index counted it, so `points_count` for the same
data depended on whether the segment had been optimized yet.
- `index_tokens` counts a transition rather than incrementing, so a point
becomes a document when it gains its first token and stops being one when
rewritten to nothing. Re-indexing the same point no longer counts it twice
- `remove` decrements only when the removed token set had tokens
- the builder counts the same way
`points_count` feeds `count_indexed_points` for cardinality estimation, and
is `N` for the IDF of a text score.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(segment): replace timing-based building cancellation test
`test_building_cancellation` compared wall-clock times of independent
builds cancelled after fractions of a measured baseline. Its tolerances
had to be widened repeatedly (#2039, #8243, #10346) and it still failed on
Windows CI, e.g. when an early stop landed in a slow non-cancellable setup
step (time_early: 969, time_later: 631, baseline 2488).
Replace it with tests that target build phases explicitly and measure
work instead of time:
- test_building_cancelled_before_start: a build cancelled up front returns
`Cancelled` without starting any vector index work.
- test_building_cancelled_during_main_graph: a watcher cancels the build
once the main HNSW graph (observed via the progress tracker) reaches
1000 of 10000 points. The build must return `Cancelled`, leave the graph
unfinished, and insert at most 2 points per build thread after the flag
is set (only in-flight insertions may complete).
Both also check that a cancelled build leaves nothing behind in the
segments or temp directory. Use random vectors with a fixed seed instead
of identical zero vectors.
Verified the main-graph test fails when the per-point stop check is
removed, or only done every 64 points (102 points after cancel, limit 16).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(segment): cover cancellation during HNSW graph healing
it: an uncancellable heal phase is what blocked consensus when removing a
replica during an optimization.
Track progress of the `migrate` phase (healed `(point, level)` pairs out of
the total), like `main_graph` already does. Besides showing healing
progress in optimization telemetry, this lets a test tell "stopped right
away" apart from "healed everything, then noticed the flag"; both return
`Cancelled`, because the flag is checked again right after healing.
Generalize the main-graph cancellation test into a helper that cancels
any phase at 10% of its work and checks that at most 2 items per build
thread complete afterwards, and add test_building_cancelled_during_heal:
it builds an HNSW segment, deletes a quarter of its points (below the
default healing_threshold of 0.3), and cancels the rebuild while healing.
Verified the test fails with the heal stop check removed ("migrate was
completed despite cancellation"). Measured: 8 items after cancel with 8
build threads, out of 3000-5700 to heal.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>