27 Commits
Author SHA1 Message Date
Andrey VasnetsovandClaude Opus 5.5 aef8cb6dd9 Report the RAM of compacted payload trackers in memory usage (#10749)
A compacted tracker holds every mapping on the heap and only reads its file on
open, so the payload storage's report listed its files but not that RAM.
CompactedTracker::ram_usage_bytes is passed up through TrackerEnum, Logstore
and Blobstore::ram_usage_bytes to PayloadStorageImpl, whose memory reporter now
reports it as extra RAM. Append-only trackers and mutable mode report 0: they
read their mappings from the files.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-27 22:24:11 +02:00
Andrey VasnetsovandClaude Opus 5.5 56a7b8a84f Compact the payload tracker of non-appendable segments the optimizer builds (#10748)
Behind the new compact_logstore_tracker feature flag, implied by
serverless_compatible. SegmentBuilder calls PayloadStorageEnum::make_immutable
before its flush of a non-appendable segment's payload storage.

Blobstore::make_immutable switches an append-only storage to the compacted
tracker in memory, from the mappings in RAM plus whatever part was flushed, and
removes the append-only tracker file; the next flush saves the compacted one.
The mutable mode is left as is. Sparse vector storages keep the append-only
tracker: plain nearest queries never read them.

CompactedTracker::from_tracker no longer writes, and the tracker keeps its fs
as an object-safe TrackerFs instead of a boxed closure.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-27 21:53:12 +02:00
Andrey VasnetsovandClaude Opus 5.5 3ec084c066 Open writable Logstore over either tracker format with TrackerEnum (#10746)
Logstore holds a TrackerEnum, so an indexed segment whose tracker was
compacted opens through the same writable path as any other storage. New
storages still start with the append-only tracker.

A compacted tracker opened writable keeps a type-erased saver over the fs it
was given and rewrites its whole file on flush. The flusher takes the target
point offset count the Logstore flusher recorded, so mappings of puts made
after the flusher was created are not persisted ahead of their value data.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-27 21:21:13 +02:00
Andrey VasnetsovandClaude Opus 5.5 934c22441b Read Logstore through either tracker format with TrackerEnum (#10745)
* Read Logstore through either tracker format with TrackerEnum

TrackerEnum holds an AppendOnlyTracker or a CompactedTracker, picked by which
file is on disk, and implements TrackerRead. LogstoreReader, BlobstoreReader
and BlobstoreView use it; the writable Logstore is unchanged. A compacted
storage is complete, so live reload and preload are no-ops for it.

CompactedTracker gains exists/preopen and always opens populated, so a preopen
fetches the whole file. from_tracker converts any tracker into the compacted
form, used by the tests and by an ignored test measuring a real tracker file.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Adapt TrackerEnum reader to the dev UIO and IO stats changes

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-27 20:43:33 +02:00
ed1ce7dc37 Add CompactedTracker: RAM-resident Logstore tracker with a compact file (#10737)
* Add CompactedTracker, a RAM-resident tracker with a compact file

Second step towards an immutable tracker for Logstore, on top of the
unified TrackerRead trait. CompactedTracker keeps every mapping in RAM
and implements TrackerRead from there, so a reader decodes the file once
on open and never touches the disk for lookups.

Flushing rewrites the whole file from a snapshot taken when the flusher is
created: the mappings are varint delta encoded against the end of the
previous value (zero for back-to-back values, so the runs compress well)
and LZ4 compressed behind a small header. The file is replaced with
atomic_save, and a stale flusher whose snapshot holds no more mappings
than the file already has is a no-op, so out-of-order flushes never roll
the file back.

Decoding validates every byte: truncated or trailing data, bad magic or
version, and deltas that do not add up to a valid pointer are rejected.

The type is not wired into Logstore or LogstoreReader yet, that is the
next step in the stack.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Split CompactedTracker into a module

mod.rs keeps the struct with its lifecycle and write side, format.rs
holds the file layout and the encode/decode functions, read.rs holds the
TrackerRead impl, and tests.rs the tests. No functional change.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Track dirtiness in CompactedTracker with a flag

set accepts any point offset and replaces existing mappings, so the
mapping count no longer identifies the state. A dirty flag replaces the
persisted count: set marks the tracker dirty, a flusher swaps the flag
off and takes a copy of the pointers, and a clean tracker gets a flusher
that does nothing. A failed write marks the tracker dirty again so the
next flush retries. The stale-flush guard is gone, flushers of one
storage are serialized by the segment flush lock.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Store CompactedTracker mappings as columns instead of LZ4 varints

Gaps, block-bitpacked lengths (min + width per 128 lengths), and a list of
pointers that do not start where the previous value ends. On a real payload
tracker (105k mappings) this is 1.14 bytes per mapping instead of 2.05, 14x
smaller than the flat tracker file, against an entropy floor of ~1 byte.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Document a worked example of the CompactedTracker format

Six point offsets with a gap and a page rollover, their three columns, the 32
serialized bytes and how one pointer decodes. test_documented_example pins the
documented bytes to the encoder.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* zstd-based format (#10787)

* Fix CompactedTracker after UniversalWriteFileOps rename

The stacked merge brought in UniversalWriteFs; update CompactedTracker
bounds so the crate compiles again.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: xzfc <5121426+xzfc@users.noreply.github.com>
Co-authored-by: qdrant-cloud-bot <111755117+qdrant-cloud-bot@users.noreply.github.com>
2026-09-25 21:18:11 +02:00
Luis CossíoandClaude Opus 5.5 1f0b3de20c [UIO] Fold *FileOps* into *Fs* traits (#10768)
* Merge UniversalReadFileOps into UniversalReadFs

Every filesystem implemented both traits, so fold list/exists/from_context
into UniversalReadFs and make UniversalWriteFileOps extend it directly.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* rename `UniversalWriteFileOps` -> `UniversalWriteFs`

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-24 15:30:28 -03:00
Andrey VasnetsovandClaude Fable 5.1 878843e6e0 Unify the blobstore tracker read interface across storage modes (#10735)
One TrackerRead trait, in lib/blobstore/src/tracker/read.rs, now covers
the Gridstore Tracker and ReadOnlyTracker and the Logstore
AppendOnlyTracker. LogstoreView, LogstoreReader, and validate_consistency
are generic over it, as GridstoreView already was over the Gridstore-only
trait.

The trait gains get_range and an access pattern on get, and its iter
returns impl Iterator instead of the concrete Gridstore Iter, which drops
the storage type parameter from the trait. The Gridstore trackers
implement get_range through a new read_slots helper; nothing in Gridstore
calls it yet. The lifecycle methods of LogstoreReader (open, preopen,
files, live preload, live reload, clear cache) stay in an impl block bound
to AppendOnlyTracker, so a future immutable in-RAM tracker plugs into the
read path without inheriting reload semantics.

No runtime change: Gridstore slot reads keep the Random access pattern,
and PointerItem stays the iterator item type.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-23 15:40:51 +02:00
Luis CossíoandClaude Fable 5 690d92e751 [updater] genericize fs to use UniversalAppendFs (#10451)
* introduce UniversalAppendFs helper

* AI: migrate to UniversalAppendFs bound

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* don't use it in Gridstore

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-08 16:56:34 -03:00
Luis CossíoandClaude Fable 5 fda819d45a [updater] strip fs from components (#10450)
* strip fs from id tracker

* strip fs from Gridstore and Logstore

* strip fs from UpdateOnlyBlobstore

* strip fs out of null and bool indexes

* strip fs out of chunked vectors

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-07 23:45:02 -03:00
Luis Cossío 0798695099 [UIO] Segment live_preload waits for all IO before returning (#10357)
* `LiveReload::live_preload` returns futures

* await reopens and reloads concurrently
2026-09-01 10:06:29 -04:00
Luis CossíoandClaude Fable 5 515e8ace69 [UIO] make UniversalRead::live_preload async (#10356)
* rename `reopen`->`live_reload` and `schedule_reopen`->`live_preload`

* `UniversalRead::live_preload` returns a shared future

* assert snapshot-miss eagerly on `live_preload`

`live_reload` cannot see the failed preload: its blocking fallback
re-resolves the length from the remote and succeeds. The error
surfaces at preload time, as callers (`ok_not_found`) expect.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 10:06:28 -04:00
Luis Cossío 83311bc243 [UIO] CachedFs waits for scheduled files to resolve + misc (#10353)
* [CachedFs] new `schedule` and `wait_all` primitives

* [AppendableIdTracker] don't reopen if just opened

* eager NotFound in `schedule_open`

* add traces for async reads

* finish `preopen`/`preload` with `wait_all`

* lock all segments in parallel for `live_reload`

* LIST before everything

to do: we don't have whole-fetch in async mode. to prevent sequential
`len`, we won't overlap static files with LIST.

* `wait_all` returns nothing
2026-09-01 10:06:28 -04:00
Tim Visée 02106ff176 Fix gridstore new page panic (#10399)
* Add repro test for Gridstore stale-gaps allocation panic

The region gaps (gaps.dat) are an acceleration structure derived from
the bitmask (bitmask.dat), persisted to a separate file without
ordering guarantees. After an unclean shutdown (power loss, kernel
crash) the gaps can claim free space where the bitmask has the blocks
marked used. An allocation in that state panics with "New page has
just been created", seen in production during WAL replay on startup.

This test simulates the torn state and expects it to be recovered; it
fails with that panic until the next commit.

* Rebuild Gridstore region gaps once on detected inconsistency

Offsets returned by the block search always come from scanning the
bitmask itself; the region gaps only steer where to look. Stale gaps
can therefore only cause missed allocations, never a wrong allocation:
every torn state funnels into the allocation failure that used to
panic with "New page has just been created".

Instead of paying for gaps validation on every open, detect the
inconsistency at that failure point, log a warning, rebuild the gaps
from the bitmask (repairing content and length), and retry. This is
allowed at most once per instance: after a rebuild the gaps are kept
consistent in memory, so a second failure would be a logic bug and
still panics. Also clamp proposed search windows to the bitmask length
so a length-diverged gaps file reaches the recoverable path instead of
an out-of-bounds panic.

* Fix typo

* Add repro test for gaps length divergence breaking page creation

BitmaskGaps::extend grows the file with zeroes before writing the new
all-free entries through the mmap. After an unclean shutdown the growth
can be persisted while the entry contents are lost, leaving phantom
all-zero entries beyond the bitmask, each claiming a full region.

Phantom full entries are invisible to the gap search, but they force
trailing_free_blocks to report zero, so the next allocation always
tries to create a new page and cover_new_page panics on its "Bitmask
length mismatch" assertion — before the lazy gaps rebuild from the
previous commit can detect anything.

The test expects opening the storage to repair the divergence; it
fails with that panic until the next commit.

* Repair gaps-to-bitmask length divergence when opening Gridstore

The number of regions the gaps file covers must match the bitmask, but
an unclean shutdown can break that: a lost extend writeback leaves
phantom all-zero entries beyond the bitmask, and a lost file growth
leaves the gaps file short. Phantom full entries force page creation
(they zero out trailing_free_blocks) and cover_new_page then panics on
its length assertion — before the lazy content rebuild can detect
anything, so that path cannot recover from this state.

Comparing the lengths is cheap, so do it on every open: on divergence,
log a warning, rebuild the gaps from the bitmask right away, and
consume the once-per-instance rebuild allowance. Allocation behavior
is unchanged on consistent storages.

* Reference to pull request

* Make gaps rebuild safe on Windows

Windows refuses to resize a file with a live user mapping, so the gaps
reset that recreated the file under its own mapping failed there with
OS error 1224 (ERROR_USER_MAPPED_FILE).

Split the rebuild along that constraint. The lazy content rebuild
keeps the mapping and overwrites the entries in place: it never needs
to resize, because a length divergence is repaired when the storage is
opened, and refuses with an error if it encounters one anyway. The
open-time length repair consumes the Bitmask by value so it can drop
the gaps mapping, atomically replace the file with the rebuilt
entries, and map it again — no resize of a mapped file on any
platform.

* Simplify gaps rebuild code

Cleanups from a review pass, no behavior change:

- compute_gaps: one read_all pass over region chunks instead of a
  read_bit_range call per region, which also removes the loop body
  duplicated from update_region_gaps
- BitmaskGaps::overwrite: take a slice instead of collecting an
  iterator the only caller already holds as a Vec
- find_available_blocks: gate the divergence clamp on the O(1)
  bit_len instead of hoisting read_all above it
- Gridstore::open: flatten the match-to-tuple into an if let, and
  shorten the rebuild warning to match the runtime one
- tests: shared bitmask setup and value read-back helpers; drop the
  length-divergence scenario from test_rebuild_gaps that
  test_gaps_length_mismatch already covers (its search assertion
  moved there)
- fix garbled log and comment wording
2026-09-01 15:40:03 +02:00
Luis Cossío e9cc3b8673 schedule_open returns nothing (#10355) 2026-08-31 12:35:51 -04:00
Luis CossíoandClaude Fable 5 2279c4a79b [UIO] UniversalReadFs::open_async (#10352)
* `UniversalReadFs::open_async`

* `schedule_open` polls once

Scheduled opens must start eagerly: sync backends complete their
`open_async` on the first poll, preserving the prefetch contract
(handles outlive later file deletions/replacements). Moved down from
the integration branch so this PR stays green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-31 12:35:50 -04:00
Luis Cossío 4e4aca8893 [UIO] renames + enforce LiveReload::live_preload (#10351)
* make LiveReload::live_preload required

* rename `schedule_prefetch`->`schedule_open`

* rename `reschedule_prefetch`->`reschedule_open`
2026-08-31 12:35:50 -04:00
Arnaud GourlayandClaude Opus 5 34fae7221f Remove stale clippy allows and the obsolete large-error-threshold override (#10337)
* Drop obsolete clippy large-error-threshold override

The 256 threshold was pinned for clippy 1.87 while tonic's `Status` was a
large error type. Upstream boxed its contents in `5de7bad` (hyperium/tonic#2253),
which is in the pinned 0.14.6 fork, so `Status` is now a single `Box` and the
default threshold of 128 passes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Remove stale clippy allows

These 11 allows no longer suppress anything under any of the three CI clippy
configurations (default, --all-targets, --all-targets --all-features).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 10:52:54 +02:00
Anton Karpov b88becd3b5 docs: parameter names in doc comments that the signatures do not have (#10290)
11 names across 7 files. Renames that did not reach the comment above them
(further_searches for further_results, query_context for segment_query_context,
block_ranges for local_block_ranges, op for operation twice, request for
requests, max_threads for max_kmeans_threads), and 3 arguments that were
removed from a signature and left documented (is_on_disk, collection_params,
search_runtime_handle with timeout).

Documentation only, no behaviour change.
2026-08-21 15:36:27 +02:00
Yash Singh a5a8427a86 fix(gridstore): correct operator precedence in Bitmask::create page-size assertion (#9963)
The debug_assert wrote 'page_size_bytes % block_size_bytes * region_size_blocks == 0', which parses as ((page % block) * region) == 0 and is therefore always true. Use is_multiple_of(block * region), matching the sibling assertion in open(). Rebased onto dev after the gridstore crate moved under lib/blobstore.
2026-08-12 14:29:18 +02:00
Luis Cossío fb1d00ecc1 [LiveReload] impl LiveReload::live_preload for Blobstore (#10039)
* impl for Gridstore

* impl for Logstore

* fix path in `Tracker::preopen`

* take non-mut `&self`

* support UnchangedOpen in `live_reload`
2026-08-11 19:14:44 -04:00
Luis Cossío bdd4faccf7 [LiveReload] Add live_preload (#10036)
* LiveReload: change associated type, add `live_preload`

* adjust existing trait implementations
2026-08-11 17:41:45 -04:00
Andrey VasnetsovandClaude Fable 5 6633205704 [UpdateOnly] create Blobstore-backed storages append-only under a feature flag (#10154)
* [UpdateOnly] create Blobstore-backed storages append-only under a feature flag

A new `append_only_storages` feature flag, enabled by `serverless_compatible`,
switches every Blobstore creation site — the payload storage, the appendable
field indexes (numeric, map, geo, full-text) and the sparse vector storage —
to the append-only Logstore mode. One shared helper maps each site's Gridstore
layout to its Logstore counterpart, carrying the page size and compression
over; blocks and regions have no append-only equivalent. Only creation
consults the flag: an existing storage keeps its persisted mode, both modes
are always readable, so flipping the flag never strands data.

Two changes make the flag usable rather than booby-trapped:

`Logstore::delete_value` now succeeds trivially where nothing is stored, as
mutable mode does, and errors only for a stored value. The ordinary write
paths delete defensively — an index clears a slot before filling it, an empty
value is stored as a deletion — and only ever hit occupied slots when
something is genuinely mutated in place.

A segment derives `append_only_storages` from the persisted payload storage
mode when it opens — not from the flags, which may have changed since it was
created — and it forces append-only mutation semantics on itself: every
mutation clones to a fresh slot, and the same-operation slot-reuse shortcut
is disabled, since the second step of a multi-step write would rewrite a
payload row those storages cannot rewrite.

The end-to-end test runs as its own binary (feature flags are process-global)
with `serverless_compatible` on: the segment comes out holding Logstore
storages, and upserts, updates of existing points, multi-step same-operation
writes, deletes, an index build over existing points, a flush and a reload all
run against them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [UpdateOnly] assume the flag pairing instead of deriving it, trim the docs

Per review: `append_only_storages` without `append_only_mutations` is not a
state to defend against — `init_feature_flags` forces the pairing, and the
same-operation slot-reuse check reads the flag directly. That deletes the
segment-side derivation: the `append_only_storages` segment field, the
persisted-mode read at open, and the `is_append_only` accessor chain through
`Blobstore`, `PayloadStorageImpl` and `PayloadStorageEnum`.

Docstrings and comments trimmed to the guarantees; how the write paths use
them is their own business.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [UpdateOnly] restructure the creation config around mode-neutral options

`CreateOptions` in the blobstore crate holds what a caller actually decides —
page size, block size, compression — and `into_config(append_only)` turns them
into the config of either mode, each taking the fields it can express. The
segment-side `storage_config` supplies only the mode, from the feature flag.

That removes the misnomer chain the previous cut left behind: nothing named
gridstore returns a config that might not be one, and no call site builds a
`GridstoreConfig` just to have its fields repacked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [UpdateOnly] keep Logstore strict; fix the callers that deleted nothing

Per review, `Logstore::delete_value` goes back to an unconditional error: a
delete reaching an append-only storage is a caller bug to fix, not a case to
absorb. The callers that issued vacuous deletes are fixed instead:

- The numeric and geo indexes only delete from the storage when their
  in-memory index actually held values at the slot — the two are written in
  lockstep, so an empty slot has nothing stored either. The map and text
  indexes already worked this way.
- The sparse storage skips the delete for keys at or past its end, where
  nothing was ever stored.

Each removed call was wasted work in mutable mode too. The e2e test now also
drives a numeric index and the sparse storage against append-only mode, and
asserts that deleting a stored sparse vector fails.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* [UpdateOnly] drop the same-op slot reuse; upserts write the whole point at once

The append-only path never needs a multi-step point write: the one real
multi-stepper was the shard's upsert — `upsert_point` followed by a payload
step under one operation number — and it now goes through
`upsert_moved_point`, which writes vectors and payload as one operation and
one slot. With that, the same-operation slot-reuse carve-out in
`handle_point_mutate` has nothing to carry: on an append-only segment every
mutating step clones to a fresh slot, unconditionally, and the
`append_only_storages` special case disappears with it. The version gate
skips only on strictly newer versions, so a caller that still multi-steps
stays correct — it pays a slot per step.

`PointToUpsert` now exposes the point's parts — raw vectors, decoded vectors,
payload — and both write paths are provided from them: `upsert_into` hands
the parts to `upsert_moved_point`, `write_moved` adapts them to the
copy-on-write move callback. The two hand-written `upsert_into` bodies and
the follow-up payload helper are gone.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Regenerate OpenAPI for the `append_only_storages` feature flag

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* one extra debug assertion

* fmt

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 15:35:55 +02:00
Andrey VasnetsovandClaude Opus 5 c8261beaec [UpdateOnly] implement UpdateOnlyPayloadStorage (#10146)
* [UpdateOnly] implement `UpdateOnlyPayloadStorage`

The payload write half for the update-only segment writer: a short-lived
storage opened for one batch and dropped with it, over a backend that only
appends.

Backed by a `Logstore` — the append-only mode of the same storage the writable
`PayloadStorageImpl` uses — so a slot's payload is written once and never
rewritten. `append_many` takes one payload per point at the slot the ID tracker
claimed for it and flushes, so a batch is durable when the call returns and
nothing is buffered across calls. Puts only buffer, so the flush is what
touches the files: one append per touched page file plus one to the tracker,
regardless of how many points the batch holds. A point with an empty payload is
skipped, since an unwritten slot already reads back as an empty payload, and so
is any gap between slots, which the tracker materializes as unmapped entries.

`Logstore` had to leave the `Blobstore` facade for this: `Blobstore`'s type is
bound at `UniversalWrite + UniversalAppend` for the sake of its `Gridstore`
variant, so it cannot be named on a backend that only appends. Its cross-crate
surface is `open_or_create`, `put_value` and `flusher`, nothing more; the new
`open_or_create` mirrors `Blobstore`'s and rejects a storage created in mutable
mode rather than opening it.

Not wired into `AppendableSegment` yet — `store_points` stays `todo!()` until
the vector storages and field indexes exist, as with `UpdateOnlyChunkedVectors`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [UpdateOnly] trim the doc comments on the payload storage writer

Keep the guarantees and the non-obvious rationale, drop the restatements — the
merged-baseline style of `UpdateOnlyChunkedVectors` and `AppendableSegment`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 10:19:49 +02:00
Srimon Danguria f4ad4f4c25 chore(deps): drop dead dependencies in edge-path crates (#10109)
- blobstore: move `dataset` to dev-dependencies (test/bench only)
- shard: remove unused `fs4`
- segment: move `tap` to dev-dependencies (test/bench only)
- sparse: move `tempfile` to dev-dependencies (test only)

Removes the `dataset -> reqwest -> hyper/tower/h2` root from the
`edge` dependency graph.
2026-08-06 17:41:57 +02:00
df177a31e2 blobstore: batched raw bytes read (#10033)
Add `read_values_bytes`, the byte-blob analogue of `read_values`, so callers
that want the stored blob rather than the parsed value keep the batched
(io_uring-friendly) read path instead of looping over `get_value_bytes`.
`read_values` becomes a thin adapter over it.

Co-authored-by: Ivan Pleshkov <ivan.pleshkov@qdrant.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 08:59:24 +02:00
Ivan Pleshkov af4e9e5d77 blobstore with raw bytes api (#10024) 2026-07-30 11:39:37 +02:00
be543561e5 Add Logstore and Blobstore wrapper (#9673)
* Gridstore: introduce storage operating mode in config

Add a mode field to the gridstore config, selecting between the dynamic
mode (current behavior, the default) and the upcoming serverless mode.
The mode is specified through StorageOptions on creation, persisted in
config.json, and read back first when opening so the correct variant can
be selected automatically. Configs written before this field existed
deserialize as dynamic.

For now, selecting the serverless mode returns an error; the variant
itself is added in follow-up commits.

* Gridstore: move dynamic implementation into dedicated module

Mechanical move of the current Gridstore implementation into
gridstore/dynamic.rs as DynamicGridstore. The public Gridstore struct
becomes a thin wrapper holding a mode variant enum, propagating every
call into the selected variant. For now the enum only has the dynamic
variant; the serverless variant is added in follow-up commits.

No logic changes to the dynamic implementation itself: only visibility,
the config parameter now passed into open (the wrapper reads it first to
select the mode), and open_or_create staying on the wrapper.

* Gridstore: add serverless tracker

Add the append-only mapping tracker for the serverless storage mode.

The tracker file is a plain array of 16-byte mapping entries without any
header: the number of mappings is defined by the exact file length, and
the entry index is the point offset. The file starts empty and only ever
grows by appending, existing bytes are never rewritten. Mappings must be
set in monotonically increasing point offset order; skipped offsets are
backfilled as zeroed entries which decode as None.

New mappings are buffered in memory and appended with a single write per
flush. A flush with a stale target is a no-op so bytes are never written
twice. A torn trailing entry (file length not a multiple of the entry
size) is ignored when reading and truncated away when opening writable.

Unlike the dynamic tracker, the file is read and written directly with
positional file IO instead of memory mapping, as serverless environments
do not handle memory mapped files well.

* Gridstore: add serverless storage variant

Add the append-only gridstore variant for serverless deployments, which
restrict IO to appending to files: existing bytes can never be
rewritten, and IO is expensive so as few files as possible are used.

The variant stores all value data in a single page file next to the
serverless tracker and the storage config, three files in total. Both
data files start empty and only ever grow by appending; there is no
preallocation, no used-block bitmask and no gap/region bookkeeping.
Values are appended at put time at the next block aligned offset, with
the zero padding included in the write so it lands exactly at the end
of the file. Mappings are buffered and appended to the tracker with a
single write per flush, after the page file is synced, so a mapping on
disk never points at data that is not durable.

Values cannot be updated or deleted, and must be put at monotonically
increasing point offsets; violations are rejected before any data is
written. Files are read and written directly, never memory mapped.

The mode is selected through StorageOptions on creation and picked up
automatically from the persisted config when opening.

* Gridstore: serverless support in reader and view

Extend the read-only GridstoreReader and the GridstoreView with the
serverless mode, keeping both public types unchanged: like the writable
Gridstore they now hold a mode variant internally, selected
automatically from the persisted config when opening.

The serverless reader holds the tracker and page directly and reads the
files positionally, without memory mapping. A live reload re-reads the
mapping count from the exact tracker file length (there is no size
header), ignoring a torn trailing entry, and never truncates as it is
read-only. Value reads always go directly to the file, so newly
appended data is readable without remapping anything.

* Gridstore: document storage operating modes

* Gridstore: review fixes for the serverless mode

Hardening and cleanup from a review pass over the new serverless
storage variant:

- Batch the reader side iteration like the writer already did, instead
  of materializing tracker mappings for the full range in one go, which
  could transiently allocate gigabytes on large storages.
- Recover the append cursors when a positional write fails partway:
  truncate the file back to the tracked length so a retried append or
  flush never rewrites bytes that already landed in the file.
- Validate page addressability before appending value data, a rejected
  put must not grow the page file.
- Cross-check tracker and page consistency when opening: mappings that
  reference value data past the end of the page file (e.g. after a
  partial copy or restore) now fail fast instead of surfacing as
  opaque read errors per point.
- Reject value pointers into any page other than page 0 on the
  serverless read path with PageNotFound, matching the dynamic mode
  contract, instead of silently reading from a wrong location.
- Refresh the reported storage size on reader live reload even when no
  new mappings were flushed, unflushed value data may have grown the
  page file already.
- Validate configs read from disk: a corrupt config with zero sized
  blocks, pages or regions is now rejected when opening instead of
  panicking on a division by zero later.
- Classify rejected serverless puts as UnsupportedOperation, consistent
  with rejected deletes, so they don't surface as user-facing
  validation errors at the segment level.
- Deduplicate the compression dispatch into Compression::compress and
  Compression::decompress, and the serverless file create/open patterns
  into shared direct IO helpers, so the two modes and files can't
  silently drift apart.

* Gridstore: cover both operating modes in mode-agnostic tests

Parameterize the gridstore tests that exercise mode-agnostic behavior
over both the dynamic and serverless mode with rstest, using a
single and bulk put/get roundtrips, storage files, basic persistence,
corrupt config rejection, batched read congruence, reader live reload,
and the different block sizes.

Mode specific expectations branch inside the tests: expected file
names, storage size semantics (whole blocks vs exactly packed bytes),
value pointer layout (page spill over vs a single packed page), and
gaps (created by deletes in dynamic mode, by skipped puts in serverless
mode). Dynamic-only internals assertions are kept behind a mode check.

Tests around updates, deletes, page spanning, block reuse and other
dynamic-only behavior intentionally stay dynamic; the serverless
specific format invariants remain covered by the dedicated serverless
tests.

* Gridstore: port serverless specific tests from sibling branch

Source the serverless specific test cases that the
serverless-gridstore-updates branch added, adapted to the dedicated
variant implemented here (distinct file names, headerless tracker
with 16 byte entries, a single packed page without trailing padding,
and rejected re-puts):

- writes only ever append: tracker and page files only grow and
  previously written bytes stay byte-for-byte untouched
- new mappings land exactly at the end of the tracker file, which
  always covers the exact number of mappings
- mapping gaps are zero-padded on disk and survive reopening
- values are packed back to back at block aligned offsets, the page
  file ends exactly at the last value
- serverless mode never creates nor reports block flag files
- a flusher persists exactly the mappings that existed at its
  creation, later puts stay pending
- a config claiming the wrong mode fails loudly in both directions
  instead of loading the incompatible file format of the other mode

Tests around their mode switching, page spanning and tolerated deletes
don't apply to this design and are intentionally not ported.

* Gridstore: test serverless production risk scenarios

Add tests for the operational aspects that matter before serverless
mode goes to production, each covering a scenario that wasn't
evaluated yet:

- Replayed puts of already persisted offsets (a WAL redo after a
  crash where the flush completed but was never acknowledged) are
  rejected without appending anything, and max_point_offset is the
  exact offset a replay must resume at.
- The accepted crash case of a tracker file extended with zeroed
  bytes: the entries count as permanent None mappings, can never be
  put again, and the storage stays consistent and writable past them.
- The read-only reader never modifies the files: opening over a torn
  tracker tail, reading, iterating and live reloading leave both
  files byte-for-byte untouched.
- A multi-round put/flush/reopen cycle always exposes exactly the
  flushed prefix, with the mapping count matching the exact tracker
  file length and unflushed offsets reusable.
- An append beyond the maximum addressable block offset is rejected
  before writing anything, keeping retried puts from growing the page
  file unboundedly.

* Gridstore: rename serverless mode to append-only, split into module

Rename the mode after its defining characteristic instead of its
deployment target: files only ever grow, existing bytes are never
rewritten. Renames Mode::Serverless to Mode::AppendOnly (persisted as
"mode": "append_only") and the on-disk file names to
append_only_tracker.dat and append_only_page_0.dat. The serverless
deployment motivation stays in the documentation.

Also split the single 2300 line serverless.rs into an append_only
module with dedicated files for the storage, page, view, reader and
tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Use universal IO in Gridstore

* Include upstream preopen logic in new Gridstore variant

* Gridstore: buffer append-only value writes until flush

In append-only mode, put previously wrote the value data to the page
file right away, one write operation per put, while mappings were
already buffered and batch persisted on flush. Buffer value writes the
same way: both the value and its mapping now only land on disk once a
flush cycle executes.

This batches all new value data into a single write operation per
flush, which is significantly more efficient on S3 based storage where
every write is a costly operation. A flush now performs exactly two
writes: one appending all buffered value data to the page file, one
appending all pending mappings to the tracker file, in that order, so
a mapping on disk never points at value data that is not durable.

The page mirrors the tracker's pending mechanism: an in-memory buffer
that is byte for byte the next append (zero padding between block
aligned values included), a watermark captured at flusher creation so
puts made during a flush stay buffered, a stale-flush no-op guard so
appended bytes are never written twice, and truncate-back recovery on
failed writes. Reads transparently serve buffered values from memory.

As a side effect, a crash between flushes now leaves nothing on disk
at all, where the write-through approach left orphaned value bytes in
the page file. The buffered data is held in memory until the next
flush, bounded by the flush cadence.

Universal IO filesystem handles are now required to be Send + Sync, so
the flusher closure can carry one to grow the page file at flush time;
all existing backends already satisfied this.

* Gridstore: rename inner DynamicGridstore to Gridstore

The dynamic variant keeps the Gridstore name; the outer dispatching type
will be renamed to Blobstore in a follow-up. Until then the inner type is
referred to as dynamic::Gridstore to distinguish it from the outer type.

* Gridstore: rename append-only variant to Arenastore

The append-only variant stores all value data in a single ever-growing
page, allocating space by appending, hence: arena store.

* Gridstore: rename outer storage type to Blobstore

The outer type dispatching between the two storage variants is now called
Blobstore, being more generic than Gridstore. This frees up the Gridstore
name, which now exclusively refers to the dynamic mode variant, next to
Arenastore for the append-only variant. Storage components keep using the
outer type, so they now use Blobstore.

The gridstore crate name, GridstoreError, and the persisted names
(config.json mode, payload config storage_type) are unchanged.

* Gridstore: split Gridstore and Arenastore into dedicated modules

The outer module is now blobstore, matching the Blobstore type it
defines. The two storage variants each get their own submodule: the
dynamic Gridstore moves from dynamic.rs into gridstore/ with its reader
and view extracted from the shared files, mirroring the arenastore/
module (previously append_only/) which already had this layout.

* Rename gridstore crate to blobstore

The crate is named after the outer Blobstore storage type it provides.
The gridstore name lives on in the dynamic mode variant. GridstoreError
and the persisted names (config.json mode, payload config storage_type)
are unchanged.

* Arenastore: pack values back to back across multiple pages

Drop the block alignment from the append-only mode: values are packed
byte to byte, without blocks, and the tracker offset is now a plain byte
offset within the page. Blocks and regions are dynamic mode concepts;
their page size constraints no longer apply to append-only configs.

Bring back support for multiple pages. Once appending a value would
grow the current page beyond the configured page size, a new page is
started, bounding the size of and the number of appends to each file:
object stores like S3 Express limit the number of appends per object.
A value larger than the page size gets a page of its own; values never
span pages.

A rollover creates the new, empty page file at put time; the value data
itself stays buffered until the next flush, which appends to each
touched page with a single write, using per-page watermarks captured at
flusher creation. The reader scans for consecutively numbered page
files when opening, validates the most recent mappings against them,
and adopts pages created since on a live reload.

* Blobstore: rename dynamic mode to mutable

Rename Mode::Dynamic to Mode::Mutable, and the persisted config value
with it: config.json now writes "mode": "mutable". There is no
compatibility alias for "dynamic", released versions never wrote the
mode field (a missing field still defaults to mutable), only unreleased
storages did.

The Gridstore type and module names for the mutable variant are
unchanged.

* Fix Edge compilation due to package rename

* Review remarks

* Extract Gridstore preopen into module

* Rename Arenastore files

* Use universal IO for append operations

* Rename GridstoreError to BlobstoreError

The error type belongs to the Blobstore crate and is shared by both the
Gridstore and Arenastore variants, so it follows the crate naming. Also
update the user-facing error messages that referred to the old name.

* Split config into per-variant types

* Rename Arenastore to Logstore

Rename the Arenastore type to Logstore, including the reader, view,
config, module and variant names. The storage file names follow:
log_page_{n}.dat and log_tracker.dat. The persisted mode tag stays
"append_only".

* Move bitmask module into the Gridstore variant

The bitmask tracks free blocks, which only exists in the mutable mode.
Move the module from the crate root into the Gridstore variant that
owns it. It stays re-exported at the crate root because the bitmask
benchmark needs a public path.

* Move pages module into the Gridstore variant

Like the bitmask, the block based pages module is only used by the
mutable mode. Move it from the crate root into the Gridstore variant
that owns it. The Logstore variant has its own page implementation.

* Use universal IO for every Logstore operation

Replace the direct_io module with universal IO in the append-only
tracker, making the whole Logstore go through a universal IO backend
bounded by UniversalRead and UniversalAppend:

- The tracker is generic over the backend now. Reads go through
  UniversalRead with the caller's access pattern, flushes land as one
  atomic append with the same offset compare-and-swap recovery as the
  pages: a retried append after a lost acknowledgement is adopted
  instead of appended twice. A torn trailing entry is still truncated
  away on writable open, through a fresh handle since shrinking is not
  supported through an open one.
- The reader now schedules a prefetch for the tracker file too, it no
  longer bypasses the backend.
- The config write, clear and wipe use the backend file operations
  instead of local filesystem calls, matching the Gridstore variant.

* Batch reads in Logstore read_values

Apply the same batching logic as the Gridstore variant: resolve all
mappings first, then fetch the value data, both through the backend's
read pipeline so async backends can serve the reads in parallel.

The tracker gains a batched lookup mirroring the mutable tracker's
iter, serving pending mappings and out of range point offsets directly
from memory. The pages gain a batched value read; unflushed values are
served from the in-memory buffers, and since values never span pages
each value is a single read without reassembly.

Like in the Gridstore variant, the callback may now be invoked in a
different order than the requested point offsets.

* Better describe logstore live reload ordering

* use enum for options, swap `*Options`<->`*Config` naming

* don't wrap enum in struct

* ditch unused `StorageConfig`, make deserialization more ergonomic

* rename `*Options`->`*Config`

* make `preopen` non-blocking

* fixup! ditch unused `StorageConfig`, make deserialization more ergonomic

* fixup! use enum for options, swap `*Options`<->`*Config` naming

* fixup! don't wrap enum in struct

* fix rebase

* use `populate` param in Logstore

* test: failing repro of stale page after live reload across rollover

A reader that live-reloads between a page rollover and the following
flush adopts the new, still empty page. The previous page is then no
longer the last one and is never reloaded again, so the tail that the
next flush appends to it stays invisible to the reader forever:

    value pointer at byte 100 with length 100 is out of range

AppendOnlyPages::live_reload only reloads the last held page, assuming
earlier pages never change once a newer page exists. But the rollover
creates the new page file eagerly at put time, while the previous
page's buffered tail only lands at the next flush (see
test_rollover_writes_no_value_data_before_flush), so a page can keep
growing on disk after its successor exists.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: reload all pages that grew

* use Fs in `open_or_create`

* fix: publish tracker mappings only after the pages reload

`AppendOnlyTracker::live_reload` observed the mapping count and made it
visible in one step, before `LogstoreReader::live_reload` reloaded the
pages. Every failure path in the page reload -- `list_files`, reopening a
grown page, opening an adopted one, the truncation check -- therefore left
the reader with mappings referencing value data it never loaded, so reads
in the new offset range fail until a later reload happens to succeed. The
edge refresh loop keeps a segment whose reload failed, expecting it to keep
serving its pre-refresh state, which it then does not.

Split observing from publishing: `reload_count` refreshes the handle and
returns the count as a `PendingReload` token, `commit_reload` publishes it.
The reader still observes the tracker first, as the writer persists pages
before the mappings referencing them, but only commits once the pages are
loaded. Reopening without committing is harmless: reads stay bounded by the
unchanged count, and the bytes below it never change.

A partial failure inside the page reload needs no unwinding, pages running
ahead of the tracker is the safe direction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* perf: batch the value reads in Logstore iteration

`LogstoreView::iter_range`, the path behind `Logstore::iter` and
`LogstoreReader::iter`, fetched the mappings for the whole range with a
single read but then read the values themselves one at a time, serially.
Gridstore routes its `iter` through `read_values` and pipelines both stages,
so a full scan of an append-only storage was the one read path without
batching -- one blocking round trip per value on the object store backends
this variant exists for. It is reached by payload storage iteration and by
the payload index build, which scans every payload.

Feed the pointers into `read_batch_values` instead, keeping the single
contiguous tracker read, which is better than the per-offset pipeline
scheduling Gridstore does on that side.

Values are now delivered through the read pipeline, so the callback may be
invoked out of order, as it already could be for Gridstore's `iter` and for
`read_values` in both variants. Both segment callers are order independent.
Tests that happened to rely on the mmap backend completing reads in
scheduling order now sort before comparing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test: don't run the failed-page-reload test on Windows

The test shrinks a page file out of band to make the page reload fail, but
Windows refuses to resize a file while the reader holds it mapped, which it
does by construction here: "the requested operation cannot be performed on a
file with a user-mapped section open". The panic is on the injection itself,
the code under test never runs.

There is no portable injection. Truncating a page the reader holds is what
the check under test detects, so the mapping cannot be avoided; failing the
adopted page open instead needs a listed but unopenable file, and
`local_list_files` descends into matching directories rather than listing
them; failing the directory listing needs the storage directory removed,
which Windows also refuses while pages are mapped.

The storage itself is fine on Windows, its append path grows mapped pages
there and every other Logstore test passes. The logic under test is platform
independent and stays covered elsewhere, with the tracker half of the
guarantee pinned by `test_live_reload`, which runs on every target.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Luis Cossío <luis.cossio@outlook.com>
2026-07-26 20:23:28 +02:00