* QuerySimd::dotprod_batch: score a contiguous run of vectors in one call
The entry point for scanning a contiguous run of encoded vectors at a
stride: `out[v]` ← score of the vector at `data[v * stride..]`. It
scores vector by vector for now; the SIMD batch kernels that share the
query loads across vectors follow.
Bench: `query{4,2,1}bit_dotprod_scan` — a hot query against runs of 512
consecutive vectors streaming from DRAM at the TurboQuant stride, per
vector and through `dotprod_batch`.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* QuerySimd: interleaved AVX-512 batch kernel with a fused reduction
Vectors up to four cache lines are scored in groups of four that share
every query block load and tail mask; the group's independent
accumulators keep `VPDPBUSD` saturated while one vector's reduction
overlaps with the next group's loads. Longer vectors keep the
per-vector walk, since the hardware prefetcher streams four interleaved
byte streams far worse than one (measured at the 4-bit width: +10 % at
dim 512, 2× slower at dim 1024).
The per-vector reduction fuses the query bytes before the horizontal
sum — `low + K · high` in i32 lanes, then one tree that widens to i64
at the end — for vectors within a per-width lane bound derived from the
encoding (2040 bytes at 4 bits, 1020 at 2, 255 for the wide 1-bit
query; unbounded for a one-byte query). A test pins the derivation to
the hand-computed 4-bit value and drives every width to its bound with
the heaviest possible inputs.
Bench: `batch_avx512_vnni` rows in the scan groups.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* QuerySimd: AVX2 batch kernel; one fused reduction for AVX2 and AVX-512
The AVX2 batch kernel scores vectors one at a time: its loop-carried
chain is a single `vpaddd` per accumulator (the `maddubs → madd`
products hang off the loads), so interleaving vectors only adds
register pressure on the 16 YMM registers — measured 10–15 % slower
with groups of two or four at the 4-bit width.
The AVX2 per-vector reduction now uses the same fused tree as the
AVX-512 one, within the same per-width lane bound; the bound test
drives both kernels.
Bench: `batch_avx2` rows in the scan groups.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* QuerySimd: copy the tail block in constant-size pieces
The SSE, AVX2 and NEON kernels run their last partial block on a
zero-padded copy of the remaining bytes. A `len`-byte copy compiles to
a `memcpy` call plus a `memset` for the padding — and the call forces
the accumulators out of their registers around it. Copy in power-of-
two pieces of constant size instead: `len` is the same for every vector
of a query, so the piece branches predict perfectly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* QuerySimd: interleaved NEON batch kernels
The SDOT and plain NEON block loops take `N` vectors at a stride, and
the batch entry points score vectors up to four cache lines in groups
of four — the same policy as the AVX-512 kernel, with the group
threshold carried over from the AVX-512 measurement rather than tuned
on ARM hardware.
Bench: `batch_neon` and `batch_neon_sdot` rows in the scan groups.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* Split async IO into extension traits; only async-capable backends implement them
Move `read_bytes_async` / `open_async` off the universal `UniversalRead` /
`UniversalReadFs` traits into dedicated extension traits, `UniversalReadAsync`
and `UniversalReadFsAsync` (traits/async_io.rs). Only backends with a genuine
async story implement them — the blob family, the disk caches layered over it,
and a trivial ready-impl for mmap (tests and the mmap lookup path) — each in a
dedicated async_io.rs next to its sync impl.
`CachedFs` now requires its inner filesystem to be `UniversalReadFsAsync`; the
requirement reaches segment code through one supertrait bound on
`UniversalReadExt`. io_uring implements no async surface anymore: the
tokio_uring bridge thread, its tests, the musl-gated tokio-uring dependency,
and the `IoUringFile` read-only-segment wiring (`UniversalReadExt` impl and
the *RoIoUring condition-checker variants) are deleted — io_uring is not a
read-only-segment backend.
The payoff for live reload: `CachedFs::resolve_prefetched` awaits every parked
prefetch, and the edge refresh flow now runs preload -> resolve -> reload, so
the per-segment write locks never wait on IO.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Decouple UniversalReadExt from the async filesystem requirement
UniversalReadExt is condition-checker dispatch; it never consumed the async
surface itself. Drop its `Fs: UniversalReadFsAsync` supertrait bound and relax
CachedFs's struct-level bound back to `UniversalReadFs` — the async requirement
now lives on the one impl that consumes it, `CachedReadFs for CachedFs`
(schedule_open parks the inner filesystem's `open_async` futures).
The bound then surfaces only on the lifecycle/preload impl blocks that go
through CachedReadFs (segment open, live-preload/reload, config reload, edge
load/refresh); the search path carries no async bounds at all.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* existing segments: wait for IO outside of search pool
* new segments: wait for IO outside of search pool
* extract reload into separate function
* Update lib/edge/Cargo.toml
---------
Co-authored-by: Tim Visée <tim+github@visee.me>
* rename `reopen`->`live_reload` and `schedule_reopen`->`live_preload`
* `UniversalRead::live_preload` returns a shared future
* assert snapshot-miss eagerly on `live_preload`
`live_reload` cannot see the failed preload: its blocking fallback
re-resolves the length from the remote and succeeds. The error
surfaces at preload time, as callers (`ok_not_found`) expect.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* [CachedFs] new `schedule` and `wait_all` primitives
* [AppendableIdTracker] don't reopen if just opened
* eager NotFound in `schedule_open`
* add traces for async reads
* finish `preopen`/`preload` with `wait_all`
* lock all segments in parallel for `live_reload`
* LIST before everything
to do: we don't have whole-fetch in async mode. to prevent sequential
`len`, we won't overlap static files with LIST.
* `wait_all` returns nothing
* Add repro test for Gridstore stale-gaps allocation panic
The region gaps (gaps.dat) are an acceleration structure derived from
the bitmask (bitmask.dat), persisted to a separate file without
ordering guarantees. After an unclean shutdown (power loss, kernel
crash) the gaps can claim free space where the bitmask has the blocks
marked used. An allocation in that state panics with "New page has
just been created", seen in production during WAL replay on startup.
This test simulates the torn state and expects it to be recovered; it
fails with that panic until the next commit.
* Rebuild Gridstore region gaps once on detected inconsistency
Offsets returned by the block search always come from scanning the
bitmask itself; the region gaps only steer where to look. Stale gaps
can therefore only cause missed allocations, never a wrong allocation:
every torn state funnels into the allocation failure that used to
panic with "New page has just been created".
Instead of paying for gaps validation on every open, detect the
inconsistency at that failure point, log a warning, rebuild the gaps
from the bitmask (repairing content and length), and retry. This is
allowed at most once per instance: after a rebuild the gaps are kept
consistent in memory, so a second failure would be a logic bug and
still panics. Also clamp proposed search windows to the bitmask length
so a length-diverged gaps file reaches the recoverable path instead of
an out-of-bounds panic.
* Fix typo
* Add repro test for gaps length divergence breaking page creation
BitmaskGaps::extend grows the file with zeroes before writing the new
all-free entries through the mmap. After an unclean shutdown the growth
can be persisted while the entry contents are lost, leaving phantom
all-zero entries beyond the bitmask, each claiming a full region.
Phantom full entries are invisible to the gap search, but they force
trailing_free_blocks to report zero, so the next allocation always
tries to create a new page and cover_new_page panics on its "Bitmask
length mismatch" assertion — before the lazy gaps rebuild from the
previous commit can detect anything.
The test expects opening the storage to repair the divergence; it
fails with that panic until the next commit.
* Repair gaps-to-bitmask length divergence when opening Gridstore
The number of regions the gaps file covers must match the bitmask, but
an unclean shutdown can break that: a lost extend writeback leaves
phantom all-zero entries beyond the bitmask, and a lost file growth
leaves the gaps file short. Phantom full entries force page creation
(they zero out trailing_free_blocks) and cover_new_page then panics on
its length assertion — before the lazy content rebuild can detect
anything, so that path cannot recover from this state.
Comparing the lengths is cheap, so do it on every open: on divergence,
log a warning, rebuild the gaps from the bitmask right away, and
consume the once-per-instance rebuild allowance. Allocation behavior
is unchanged on consistent storages.
* Reference to pull request
* Make gaps rebuild safe on Windows
Windows refuses to resize a file with a live user mapping, so the gaps
reset that recreated the file under its own mapping failed there with
OS error 1224 (ERROR_USER_MAPPED_FILE).
Split the rebuild along that constraint. The lazy content rebuild
keeps the mapping and overwrites the entries in place: it never needs
to resize, because a length divergence is repaired when the storage is
opened, and refuses with an error if it encounters one anyway. The
open-time length repair consumes the Bitmask by value so it can drop
the gaps mapping, atomically replace the file with the rebuilt
entries, and map it again — no resize of a mapped file on any
platform.
* Simplify gaps rebuild code
Cleanups from a review pass, no behavior change:
- compute_gaps: one read_all pass over region chunks instead of a
read_bit_range call per region, which also removes the loop body
duplicated from update_region_gaps
- BitmaskGaps::overwrite: take a slice instead of collecting an
iterator the only caller already holds as a Vec
- find_available_blocks: gate the divergence clamp on the O(1)
bit_len instead of hoisting read_all above it
- Gridstore::open: flatten the match-to-tuple into an if let, and
shorten the rebuild warning to match the runtime one
- tests: shared bitmask setup and value read-back helpers; drop the
length-divergence scenario from test_rebuild_gaps that
test_gaps_length_mismatch already covers (its search assertion
moved there)
- fix garbled log and comment wording
* TQ SIMD: one backend ladder, resolved once per query
The 2- and 4-bit kernels share the same preference order (AVX-512 VNNI
→ AVX2 → SSE → NEON + SDOT → NEON → scalar), spelled out six times as
chains of `is_x86_feature_detected!` — and `Query{2,4}bitSimd::dotprod`
re-ran its chain for every vector scored.
Move the ladder into one `simd::SimdBackend` enum with a single `detect()`.
The query types resolve it in `new()` and dispatch on the stored value;
the symmetric `score_{2,4}bit_internal*` entry points dispatch on
`SimdBackend::detect()`.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* QuerySimd: one query layout and scalar reference for every packing width
`Query{1,2,4}bitSimd` are three copies of the same idea — quantize the
query into i8 halves, multiply them against an integer codebook — each
with its own query layout and its own set of SIMD kernels. Introduce
`simd::query::QuerySimd<PLANES>`, generic over the number of codes per
packed byte (2, 4 or 8), which the three widths will share.
The query halves are stored as planes, one per code position within a
byte: plane `k` entry `j` is the half of query dim `PLANES · j + k`.
That is the order the codes come out of raw data bytes with a shift and
a mask, so a kernel never has to unpack them into dim order. Planes are
zero-padded to the widest SIMD block, so a partial last block on the
data side multiplies against zeros.
The widths contribute only their integer encoding (`Encoding`: codebook
table, offset, scale and query range); the 1-bit width gets one here —
`{0, 128}` with offset 64 on x86_64, `∓127` on aarch64 — chosen so the
query keeps full i8 halves. Only the scalar reference exists yet; the
SIMD kernels follow, and the width types switch over once they're in.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* QuerySimd: AVX-512 VNNI kernel on the query planes
One ZMM of packed codes per block; for each plane the codes are shifted
down by the code width, masked and looked up in the codebook with one
`vpshufb`, then `VPDPBUSD` folds them into the plane's low and high
accumulators. Two accumulator pairs per vector keep the VNNI latency
off the critical path at every width. The last partial block is a
masked load whose dead lanes multiply against the planes' zero padding.
The shift count is an immediate, so the shift-by-width helper spells
out the three widths in a `match` — the only place the kernel is not
literally generic over `PLANES`.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* QuerySimd: NEON SDOT kernel on the query planes
The AVX-512 kernel's shape on 128-bit registers: one `TBL` codebook
lookup per plane, `SDOT` (inline asm — `vdotq_s32` is still unstable)
into two accumulator pairs. The last partial block runs on a
zero-padded copy of the remaining bytes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* QuerySimd: AVX2 kernel on the query planes
One YMM of packed codes per block, one `vpshufb` lookup per plane and
`maddubs → madd` against ones into the same two accumulator pairs as
the VNNI kernel. The `maddubs` pair sums stay inside i16 by the
per-width query bounds.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* QuerySimd: SSE and plain NEON kernels on the query planes
The 128-bit forms of the AVX2 and SDOT kernels: `maddubs → madd` on
XMM, `vmull_s8 → vpadalq_s16` on NEON without `dotprod`. Every backend
of the shared query type now has its kernel.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Query4bitSimd: score through QuerySimd<2>
`Query4bitSimd` becomes an alias of the shared query type; its own
chunk-and-tail query layout and the per-backend kernels built on it go
away, along with the accuracy tests the shared module now runs for
every width. What stays in `query4bit` is the 4-bit encoding and the
symmetric `score_4bit_internal*` paths.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Query2bitSimd: score through QuerySimd<4>
The 2-bit asymmetric kernels unpacked every 4 packed bytes into 16
centroid bytes through two `pshufb` / `TBL` pair-table lookups and a
zip before a single multiply-accumulate step — on AVX-512 that was four
128-bit unpacks and six lane inserts per pair of `VPDPBUSD`. On the
query planes the same 16 codes cost one shift, one mask and one lookup
per plane, straight from a full-width load.
`Query2bitSimd` becomes an alias of the shared query type; its chunk
layout and per-backend kernels go away. The pair-table unpack stays
for the symmetric `score_2bit_internal*` paths.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Query1bitSimd: score through QuerySimd<8>
The 1-bit asymmetric kernels bit-plane-transposed the query and scored
`Σ_b 2^b · popcount(data AND plane_b)` per 16-byte block — eight
AND + popcount + add steps on XMM (even with AVX-512, through the VL
forms) and a `BITS`-deep accumulator array. On the query planes a
sign bit is just a one-bit code: shift, mask, a two-entry codebook
lookup and the same multiply-accumulate as the wider widths, on full
256-/512-bit registers.
`Query1bitSimd` becomes an alias of the shared query type. Its query
width was a const parameter (8 bits by default, 16 for TQ+ through the
`Bits1Wide` variant); the shared encoding always carries 16-bit halves,
so the variant and the TQ+ special case go away. The popcount kernels
stay for the symmetric `score_1bit_internal`, where XOR + popcount is
the right tool.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Bench: cold per-vector rows for every width
`query{4,2,1}bit_dotprod_cold` from one generic body — scalar reference,
public `dotprod` and each backend — so the widths can be compared on one
host. `TURBO_SIMD_DIMS` narrows or widens the dims of a run and
`TURBO_SIMD_POOL_KB` shrinks the pool to L1 for hot-kernel numbers.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* QuerySimd: query bytes as a parameter; 8-bit queries for the 1-bit width
The shared kernels always carried two query bytes (a ~16-bit query),
which cost the 1-bit width its 8-bit-query speed: the old bit-plane
kernel scored a 16-byte vector in 4.9 ns hot (34.7 ns cold) against
9.1 ns (54.7 ns) through the planes, the difference being the second
byte's multiply-accumulates on a vector that fills a quarter of one
block. Above 512 dims the planes win either way.
Make the number of query bytes a parameter: `QuerySimd<PLANES,
QUERY_BYTES>` with one plane per query byte and code position, and one
accumulator pair per query byte. A one-byte query is scaled to the
range of a single byte, `RADIX / 2 − 1`. `Query1bitSimd` is the
one-byte instance — at parity with the old kernel at small dims (cold
36.9 / 38.3 / 38.4 ns at d = 128 / 256 / 512) and 1.8× faster at 1536
(68 vs 121 ns) — and `Query1bitWideSimd` the two-byte one, which TQ+
selects through the `Bits1Wide` variant as before. The 2- and 4-bit
widths keep two bytes.
Bench: `query1bit_wide_dotprod_cold` and a `query1bit_wide` row next
to BQ.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
The debug-tools workflow cross-compiles for x86_64-unknown-linux-musl,
but tokio-uring 0.5.0 requires libc::statx which is unavailable on musl.
Skip the tokio-uring dependency on musl and fall back to sync io_uring reads.
* `UniversalReadFs::open_async`
* `schedule_open` polls once
Scheduled opens must start eagerly: sync backends complete their
`open_async` on the first poll, preserving the prefetch contract
(handles outlive later file deletions/replacements). Moved down from
the integration branch so this PR stays green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* Read vector runs that straddle a chunk boundary
Resolve a run into per-chunk parts instead of a single range, borrowing
when it lands in one chunk and copying when it spans two. The read
pipeline schedules one range per read, so a straddling run is read
outside it.
No writer produces such a run yet, so this changes nothing on its own.
It is what a reader needs before one does — including edge and
live-reload readers, which read files a different version wrote.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Place multivector runs without regard to chunk boundaries
Writers appended a multivector's inner vectors at the end of the row
space unless the run would cross a chunk boundary, in which case they
skipped the chunk tail — the batch writers padding the skipped rows with
explicit zero rows. That made chunk geometry part of the interface every
multivector storage had to reuse.
Runs now go at the end unconditionally and the chunked storage splits
the write across chunks, as it already did for a batch of single
vectors.
What is left of the geometry is a size cap: a multivector may not exceed
one chunk. It is fill-independent, so it constrains nothing about
placement, and it is what the volatile storage needs anyway — that one
returns a plain slice and so cannot serve a straddling run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Split a run at chunk boundaries in one place
Reading, writing in place and appending each derived the split from
`remaining_chunk_capacity`, so every one of them had to know that a run
does not necessarily fit where it starts.
`split_run` hands out the parts instead: one per chunk the run covers,
each carrying where it goes and how much of the run it takes. Nothing
asks how much room is left any more, and `get_chunk_offset` goes with it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Keep straddling runs on the read pipeline
Reading a straddling run outside the pipeline blocked the scheduling
loop on one read, which costs a round trip on a backend that fetches
remotely and drops the batch back to sequential.
A run is now scheduled as one read per chunk it covers. Parts complete
in any order, so each run holds what has landed until the last part
does, then hands the callback the stitched vectors. Runs taking a single
read carry the caller's data in the tag and never touch that table.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Stop capping a multivector at one chunk
The cap outlived its reason on disk, but the volatile storage still
needed it: its `get_many` handed out a slice of one chunk, so a run that
crossed a boundary had nowhere to come from. And since a volatile
storage is a target of the batched copy that builds a segment, dropping
the cap only on disk would have turned a rejected write into a failed
merge.
So the volatile storage splits and stitches too. Both are a few lines
each, and placing a run no longer skips a chunk tail, so `extend` is now
`insert_many` at the end of the storage.
Nothing user-facing moves: `MAX_MULTIVECTOR_FLATTENED_LEN` caps a
multivector at 1M elements, far inside a 32 MiB chunk, so the storages
only ever rejected what reached them unvalidated.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Schedule a single-read run without the queue
The scheduling loop resolved every run into the queue and then took it
straight back out, so the overwhelmingly common run — one that fits a
chunk — paid a push and a pop for nothing. It now goes to the pipeline
directly, and the queue holds only what a straddling run leaves behind.
Worth ~10% on the multivector read benchmark, and it collapses the
"top up, then take" pair into one decision. Extracting that bookkeeping
into helpers instead was measured and is much worse: the mmap pipeline
alternates one schedule with one wait, so the loop body is a few dozen
nanoseconds, and a helper carrying the cold map and stitching paths is
too big for the compiler to inline back into it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Repoint the multivector WAL-replay test at a live rejection
The test upserted a multivector too large for a storage chunk, which no
longer fails: the storages stopped capping one at a chunk. Nothing else
covered a multivector operation that only the apply path rejects.
A raw blob that is not a whole number of quantized records still does,
so the test now uses that, alongside its dense and sparse siblings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Test reading multivectors with legacy chunk-tail padding
Locks the compatibility contract that pre-straddle files — runs that
skip a chunk's leftover slots — still reopen as single-chunk borrows.
* chore: retrigger CI after flaky test-consensus-compose
* Move ReadTag into for_each_vector, its only user
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AN4Hgbd65gDhesthJk5bUY
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com>
* Add SetFlushInterval op to the model tester
Changes the collection's flush_interval_sec mid-run through the same path
update_collection takes (persist the optimizer-config diff, then recreate
the optimizers in the background). The model is untouched: what it perturbs
is the flush cadence, so how much of the workload is still WAL-only when a
restart hits, plus the worker stop/start race in on_optimizer_config_update.
Kept in FORCE_OFF for now: with the optimizer on it makes stale point state
visible within a few ops of the config change. Narrowed to
recreate_optimizers_background, see the comment on Swarm::BASE.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011AjmS5GFeGztfP3JqutnXj
* Keep SetFlushInterval enabled in the swarm
Drops it from FORCE_OFF so the divergence it surfaces is reachable without
--enable-force-off (which would also enable the broken vector-name ops).
The evidence moves from the FORCE_OFF comment onto the op's own doc.
The two optimizer-on harness gates now fail whenever the swarm draws the op.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011AjmS5GFeGztfP3JqutnXj
* Propagate proxied changes when unwrapping proxies on optimization failure
unwrap_proxy puts the wrapped segments back into the segment holder, so the
changes recorded on the proxy while the optimization ran (deleted points,
index and vector-name changes) have to reach the wrapped segment first. They
did not, so every point deleted or overwritten during the optimization kept
its pre-optimization copy live next to the new copy in the write segment, and
reads saw both: counts too high, scroll and search returning the stale copy.
The snapshot unproxy path already does this; the optimizer failure path was
the only place putting a wrapped segment back without it. It is reachable
whenever the shard outlives the cancellation, in particular an update_collection
that recreates the optimizers while an optimization is in flight.
Lock order is holder-then-updates, matching try_unproxy_segment: updates-then-
holder-write deadlocks against the snapshot path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011AjmS5GFeGztfP3JqutnXj
* Drop the stale failure note from the SetFlushInterval doc
The divergence it described is fixed in this branch.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011AjmS5GFeGztfP3JqutnXj
* test as well with 0s as flushing interval
* Fail optimization unwrapping when proxy propagation fails
Losing proxied deletes and index changes is data corruption, so return the
error instead of logging it: no proxy is unwrapped and the changes stay
served by the proxies. The cancelled-segment cleanup moves ahead of
unwrap_proxy so the orphan is still removed when that error fires.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CpoaAtGbAQAEuxEi8ScHc5
* Drop the model tester --flush-interval-sec flag
SetFlushInterval covers the interval now, so the run starts at the shipped
5s default (fixture::INITIAL_FLUSH_INTERVAL_SEC, still traced in the header)
and the ops move it from there. Also documents what 0 does now that it is a
generated value.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CpoaAtGbAQAEuxEi8ScHc5
* Fail snapshot unproxying when proxy propagation fails
Both paths logged the error and unwrapped anyway, dropping the deletes and
index changes that never reached the wrapped segment. Same reasoning as
unwrap_proxy in the optimizer.
try_unproxy_segment hands the lock back and leaves the proxy installed, the
failure mode its doc already describes: the caller keeps it in `proxies` and
unproxy_all_segments retries the propagation right after. unproxy_all_segments
returns before touching the holder, so the temp segment the surviving proxies
write into stays in place (remove_segment_if_not_needed only checks whether it
is empty and appendable, not whether a proxy still references it).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CpoaAtGbAQAEuxEi8ScHc5
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
ast-grep 0.45 no longer parses a leading `::` fragment as a pattern, so
`pattern: ::$MOD` stopped matching and every `::common::` / `::wal::` path
survived into the generated qdrant-edge crate, failing `just rs-check`.
Matching the node text instead keeps the rule working on 0.44 and 0.45: the
amalgamation output is byte identical to what 0.44 produced before.
Claude-Session: https://claude.ai/code/session_01CpoaAtGbAQAEuxEi8ScHc5
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
`quantized_size_for` pads its `dim` argument, so passing the already
padded `self.padded_dim` applied the x1.5 expansion of `Bits1_5` a
second time. Nothing on disk depends on the value for that width: the
quantization path sizes its records with `quantized_size_for` from the
raw dim, and the Turbo datatype storages, which do use
`quantized_size()` as their record size, are fixed at Bits4, where the
padding is idempotent. The wrong value only over-reserved the
`quantize` output buffer.
Compute the packed size from `padded_dim` directly and cover `Bits1_5`
in the byte-length test.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(strict-mode): enforce max_query_limit on scroll requests when limit is omitted (fixes#10373)
* style: rustfmt scroll query_limit for CI lint
---------
Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com>
* fixup! Implement `ChangeAliases` operation
* fixup! Add `ChangeAliases` to replay-safety tests
* fixup! Add `ChangeAliases` tests
* De-slop ⛷️
* De-slop 🏂
* Add `TestSlowDown` and `TestTransientError` actions
These would have to be implemented on `TableOfContent` when switching
to `ConsensusStateMachine` as main consensus impl
* Handle more stupid corner-cases for `ChangeAliases` prop tests
* Add `AliasMapping::remove` and `AliasMapping::rename` methods
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Implement `ChangeAliases` operation
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `ChangeAliases` to replay-safety tests
Multi-action operations that rename an alias are skipped in the convergence
property: the current implementation does not replay them convergently, and
the machine reproduces that.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `ChangeAliases` tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add test-only `PeerMetadata::new` constructor
Test state needs peers at a version other than this build. The `version` field
is crate-private and `current()` is the only constructor, so gate the new one
on the `testing` feature and enable it for the `storage` test build.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Implement `UpdatePeerMetadata` operation
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `UpdatePeerMetadata` to replay-safety tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `UpdatePeerMetadata` tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Implement `UpdateClusterMetadata` operation
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `UpdateClusterMetadata` to replay-safety tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `UpdateClusterMetadata` tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Implement `SetQuotaConfig` operation
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `SetQuotaConfig` to replay-safety tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `SetQuotaConfig` tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Implement `TestSlowDown` and `TestTransientError` operations
Both are node-local: one sleeps, the other fails at random. They plan no
actions, like `Nop`, so the replay-safety properties have nothing to add.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add `TestSlowDown` and `TestTransientError` tests
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fixup! Add `UpdatePeerMetadata` to replay-safety tests
* fixup! Add `UpdateClusterMetadata` tests
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Unindexed Match::Text and Match::Phrase previously shared a String::contains
arm, so phrase order was ignored and queries matched across token boundaries.
Use the default Word tokenizer for best-effort parity with indexed fields.
Fixes#10182
* Keep consensus operation awaiters alive for concurrent waiters
Callers proposing an identical consensus operation deduplicate onto one
broadcast channel, and the map holds its only sender. Removing the entry on
timeout therefore closed the channel for every other waiter, failing their
still in-flight operation with "Channel sender dropped".
Only remove the entry once no receiver is left, dropping our own receiver
first so the last caller out cleans up. Apply the same to
await_for_multiple_operations, which registered awaiters but never
deregistered them when it timed out.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add debug assert to ensure we clean up consensus operation waiters
* Deregister consensus operation awaiters when the waiter is dropped
Dispatcher::submit_collection_meta_op registers the expected operations before
proposing, then drops that future unpolled whenever the proposal itself fails.
The awaiters stayed in the map with no receiver left, so the next identical
request deduplicated onto a dead entry and never heard back. This is what
tripped the new debug assert in CI: a rejected create-collection left a
SetShardReplicaState awaiter behind, and the next run of the same test hit it.
Move registration into an OperationAwaiters guard that deregisters on drop, so
timeout, drop-before-poll and request cancellation are all covered by one path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Close the race the awaiter debug assert trips on
The assert is sound only if no one can observe an entry whose receivers are all
gone. Both cleanup sites dropped their receiver before taking the map lock, so a
concurrent register could see exactly that and panic. Drop the receiver while
holding the lock instead, and take that lock once per batch rather than once per
operation: creating a collection registers an awaiter per replica, on the mutex
the consensus thread needs for every entry it applies.
Collect the awaiters into the guard as we go, so giving up part way still
deregisters the ones already registered, and only build the broadcast channel
when the operation is not already in-flight.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: timvisee <tim@visee.me>
Add stop checks during the pre-HNSW build setup phase so cancellation
is observed promptly, and widen timing tolerance for noisy Windows debug
builds where post-stop delays can exceed 1s.
* docs(schema): declare enforced 1..=65536 bound on VectorParams.size
The REST layer enforces an upper bound of 65536 on VectorParams.size
via a custom validator (validate_nonzerou64_range_min_1_max_65536),
but custom validators contribute no bounds to the generated JSON
schema - so the published OpenAPI document only declared minimum: 1,
while DenseVectorConfig.size already documented both bounds.
Add an explicit #[schemars(range(min = 1, max = 65536))] attribute so
clients validating requests against the schema see the same contract
the server enforces, update docs/redoc/master/openapi.json
accordingly, and pin the bound with a unit test asserting the
generated schema.
Fixes#9942
* test(openapi): accept documented size bound as a rejection path
test_vector_dimension_limit asserted that an oversized VectorParams.size
reaches the server and returns the exact runtime 422 message. Now that the
enforced 1..=65536 bound is documented in the served OpenAPI schema (#9942),
request_with_validation rejects such payloads client-side before sending.
Accept either layer: a client-side jsonschema.ValidationError or the
server-side validation error.
* test(openapi): handle both rejection layers in dimension limit
pytest.raises only covered the client-side jsonschema rejection; if the
request reached the server instead, the test would fail on an unhandled
response. Use try/except around request_with_validation and assert the
server-side status and exact error message in the else branch.
* test(openapi): assert exact HTTP 422 on server-side rejection
A broad not-ok check would pass on any error status carrying the same
error text; pin the documented contract to 422.
* refactor(tests): address review feedback
Remove the unit test asserting the generated VectorParams schema shape -
it only restates the schemars attribute and adds maintenance cost.
Reduce test_vector_dimension_limit to its actual contract: an oversized
dimension is rejected by the documented OpenAPI schema before the
request is sent.
* Drop obsolete clippy large-error-threshold override
The 256 threshold was pinned for clippy 1.87 while tonic's `Status` was a
large error type. Upstream boxed its contents in `5de7bad` (hyperium/tonic#2253),
which is in the pinned 0.14.6 fork, so `Status` is now a single `Box` and the
default threshold of 128 passes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Remove stale clippy allows
These 11 allows no longer suppress anything under any of the three CI clippy
configurations (default, --all-targets, --all-targets --all-features).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Drops 6 crates from the release build and 7 from the workspace test
build, with no source changes.
- geo: no triangulation, only Contains/Intersects/Haversine (spade, earcut)
- jsonwebtoken: HS256 from_secret only, no PEM keys (pem, simple_asn1)
- tar: nothing sets unpack_xattrs, which defaults to false (xattr)
- duplicate: every duplicate_item names its module (proc-macro2-diagnostics)
- pprof: no C++ frames to demangle (cpp_demangle)
Also promotes duplicate to a workspace dependency.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* perf: skip external-id resolution in search post-processing
process_search_result hands its scored points to retrieve(), which resolves
every external id back into an internal offset — although the offsets are
already known there: they come straight from the vector index as
ScoredPointOffset. Every request therefore performs top_k x segments
redundant external->internal lookups.
Whether that costs anything depends on the id tracker being mutable: a
freshly optimized segment looks ids up through a BTree, while a restarted
node maps an immutable tracker and resolves in constant time. So the effect
shows up right after ingestion or optimization and disappears after a
restart, which is what makes it easy to miss — a benchmark that starts from
a freshly booted node never sees it.
Split retrieve() into resolution + retrieve_resolved() and pass the offsets
from process_search_result directly, applying the deferred cutoff by offset
instead of resolving ids just to filter them. retrieve() behaviour is
unchanged for all other callers; retrieve_resolved() is private and states
that deferred filtering is the caller's responsibility.
Measured on glove-100-angular (1.18M points, 21 segments, top_k=10) at a
fixed request rate, on a node that had just finished ingesting: median
latency ~15-24% lower, ~8% less CPU per request. Both figures come from the
same node before and after the change.
* review fix
* review fix
---------
Co-authored-by: Ivan Dashchinskiy <iadashchinskiy@sbertech.ru>
* Optimization: Skipping items before pushing into PriorityQueue.
* Apply suggestion from top-k-update branch
* Stabilize order of equal scored items in tests
* Steer writes away from appendable segments at max_segment_size
* pick a write target that stays under the configured size cap, instead of
growing an appendable segment past it
* clamp the deferred points threshold to max_segment_size, treating a zero cap
as uncapped
* apply the same cap when replaying the WAL, so recovery matches live updates
* plumb the cap through the update worker and cover the silent-failure gaps
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Extract the per-segment capacity check into a helper
Lets has_appendable_segment_with_capacity short-circuit on the first segment
below the cap instead of collecting every eligible ID.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Replace cgroups-rs with direct cgroup memory file reads
We used cgroups-rs in exactly one place, to read the memory limit and
usage of our own cgroup, so read those files directly instead. Drops 34
crates from the lockfile, including the zbus stack that carries
RUSTSEC-2026-0221.
Also fixes two latent cgroup v1 bugs (the LONG_MAX unlimited sentinel
reported ~9 EB of total memory, an unreadable limit file reported 0
bytes) and the hierarchy mix-up on hybrid hosts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Decline cgroup memory reporting when the usage read fails
Reporting a usage of 0 made available_memory_bytes claim the whole cgroup
limit as free. Fall back to sysinfo when the usage file cannot be read at
init, and keep the last known value on a failed refresh.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Keep the last known memory limit when its read fails
A transient read failure cleared the cached limit and silently fell back
to host memory while the process was still capped, the same direction of
over-reporting as the usage read. Both now keep their last known value,
and a limit lifted at runtime still clears.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Treat a malformed memory limit as an error, not as unlimited
Parse failures returned Ok(None), so garbage in the limit file cleared a
valid cached limit on refresh and read as unlimited at init. Reserve
Ok(None) for "max" and the v1 sentinel, and report anything else as
InvalidData so the last known limit survives.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-up to #10287, which added `max`. Expressing a minimum still
required spelling out `(a + b - |a - b|) / 2`, the sign flip of the max
identity — drop the `neg` and you silently get a maximum instead. It also
only works for two operands and mentions each one twice, so the scorer
walks every sub-tree twice per candidate point.
The pair is what makes clamping expressible:
{"max": [0.0, {"min": [1.0, "$score"]}]}
`min` mirrors `max` throughout, and both guard helpers introduced in
#10287 already took an `operator: &str`, so they are reused unchanged: an
empty operand list is rejected at parse time rather than folding to
+infinity, and the Edge FFI rejects it at construction time. The result
needs no `is_finite` check, since `min` cannot produce a non-finite value
from finite inputs.
The unindexed-field walker shares one arm for `Max | Min` as the bodies
are identical, with a test pinning `min` separately so a later split
cannot silently drop it.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
11 names across 7 files. Renames that did not reach the comment above them
(further_searches for further_results, query_context for segment_query_context,
block_ranges for local_block_ranges, op for operation twice, request for
requests, max_threads for max_kmeans_threads), and 3 arguments that were
removed from a signature and left documented (is_on_disk, collection_params,
search_runtime_handle with timeout).
Documentation only, no behaviour change.
* feat: add a dedicated max operator to score formulas
Expressing a maximum in a score formula required spelling out the
arithmetic identity `(a + b + |a - b|) / 2`. That is easy to get wrong
(the `/ 2` is load-bearing), only works for two operands, and mentions
each operand twice, so the scorer evaluates every sub-tree twice per
candidate point.
`max` is variadic, mirroring `sum` and `mult`:
{"max": ["$score", {"mult": [0.5, "popularity"]}]}
Unlike `sum` and `mult`, `max` has no identity element for the empty
case, so an empty operand list is rejected at parse time rather than
folding to -infinity and scoring every point with a non-finite value.
The check lives in `ExpressionInternal::parse_and_convert`, which every
entry point passes through, and the Edge FFI additionally rejects it at
construction time to match how that crate validates elsewhere.
The result needs no `is_finite` check: unlike `log10`, `exp`, `div`,
`sqrt` and `pow`, `max` cannot produce a non-finite value from finite
inputs, so it follows the existing `sum` convention.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: cover max error propagation and datetime operands
An operand that fails must fail the whole expression rather than being
passed over in favour of a finite sibling. Covered with the failure both
before and after the finite operand: `mult` short-circuits on zero and
so can skip evaluating later operands, and this pins down that `max`
must not grow a similar shortcut that would swallow an error.
Also covers `max` over datetime operands, which reach the scorer through
a separate conversion to seconds, so that "score by whichever timestamp
is newer" is verified rather than assumed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Use a warmup baseline and adaptive timeout so the search is cancelled
mid-flight on both fast macos ARM runners and overloaded hosts, instead
of relying on a fixed 350ms cutoff.
* impl live_preload for payload indexes
enable live_preload for bool and null indexes
* (not) impl live_preload for `VectorIndexReadEnum`
* impl live_preload for `ReadOnlyPayloadStorage`