mirror of
https://github.com/qdrant/qdrant.git
synced 2026-09-29 01:17:56 -05:00
* Read vector runs that straddle a chunk boundary Resolve a run into per-chunk parts instead of a single range, borrowing when it lands in one chunk and copying when it spans two. The read pipeline schedules one range per read, so a straddling run is read outside it. No writer produces such a run yet, so this changes nothing on its own. It is what a reader needs before one does — including edge and live-reload readers, which read files a different version wrote. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Place multivector runs without regard to chunk boundaries Writers appended a multivector's inner vectors at the end of the row space unless the run would cross a chunk boundary, in which case they skipped the chunk tail — the batch writers padding the skipped rows with explicit zero rows. That made chunk geometry part of the interface every multivector storage had to reuse. Runs now go at the end unconditionally and the chunked storage splits the write across chunks, as it already did for a batch of single vectors. What is left of the geometry is a size cap: a multivector may not exceed one chunk. It is fill-independent, so it constrains nothing about placement, and it is what the volatile storage needs anyway — that one returns a plain slice and so cannot serve a straddling run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Split a run at chunk boundaries in one place Reading, writing in place and appending each derived the split from `remaining_chunk_capacity`, so every one of them had to know that a run does not necessarily fit where it starts. `split_run` hands out the parts instead: one per chunk the run covers, each carrying where it goes and how much of the run it takes. Nothing asks how much room is left any more, and `get_chunk_offset` goes with it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Keep straddling runs on the read pipeline Reading a straddling run outside the pipeline blocked the scheduling loop on one read, which costs a round trip on a backend that fetches remotely and drops the batch back to sequential. A run is now scheduled as one read per chunk it covers. Parts complete in any order, so each run holds what has landed until the last part does, then hands the callback the stitched vectors. Runs taking a single read carry the caller's data in the tag and never touch that table. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Stop capping a multivector at one chunk The cap outlived its reason on disk, but the volatile storage still needed it: its `get_many` handed out a slice of one chunk, so a run that crossed a boundary had nowhere to come from. And since a volatile storage is a target of the batched copy that builds a segment, dropping the cap only on disk would have turned a rejected write into a failed merge. So the volatile storage splits and stitches too. Both are a few lines each, and placing a run no longer skips a chunk tail, so `extend` is now `insert_many` at the end of the storage. Nothing user-facing moves: `MAX_MULTIVECTOR_FLATTENED_LEN` caps a multivector at 1M elements, far inside a 32 MiB chunk, so the storages only ever rejected what reached them unvalidated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Schedule a single-read run without the queue The scheduling loop resolved every run into the queue and then took it straight back out, so the overwhelmingly common run — one that fits a chunk — paid a push and a pop for nothing. It now goes to the pipeline directly, and the queue holds only what a straddling run leaves behind. Worth ~10% on the multivector read benchmark, and it collapses the "top up, then take" pair into one decision. Extracting that bookkeeping into helpers instead was measured and is much worse: the mmap pipeline alternates one schedule with one wait, so the loop body is a few dozen nanoseconds, and a helper carrying the cold map and stitching paths is too big for the compiler to inline back into it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Repoint the multivector WAL-replay test at a live rejection The test upserted a multivector too large for a storage chunk, which no longer fails: the storages stopped capping one at a chunk. Nothing else covered a multivector operation that only the apply path rejects. A raw blob that is not a whole number of quantized records still does, so the test now uses that, alongside its dense and sparse siblings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Test reading multivectors with legacy chunk-tail padding Locks the compatibility contract that pre-straddle files — runs that skip a chunk's leftover slots — still reopen as single-chunk borrows. * chore: retrigger CI after flaky test-consensus-compose * Move ReadTag into for_each_vector, its only user Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AN4Hgbd65gDhesthJk5bUY --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com>
Collection
Crate, which implements all functions required for operations with a single collection of points. Points within a collection should share the same payload schema and have same vector size. So that search requests could be performed over all points of a single collection.

