* feat: create appendable segments when the write target is optimizing or the shard is empty
* review: SegmentManifestState::is_writable, caller-supplied temp dir, uuid from token
- `SegmentManifestState::is_writable` with a full match replaces the ad-hoc
`matches!` in the manifest enumerator.
- `ListedSegment` is destructured in `open` so every field is accounted for.
- `create_appendable_from` is test-only; `create_appendable` is the API.
- `create_appendable` builds the scratch segment in a caller-supplied local
`temp_path` (conventionally `<shard>/temp_segments`) instead of the system
temp dir, and takes the uuid from the build token instead of parsing the path.
Upload speed of `copy_dir_via` is tracked in #10433.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* feat: pin per-shard search pool to a core
* fix: plumb search_pool_core through bindings
* feat: expose search_pool_core in python bindings
* fix: validate search pool core before pinning
* chore: trim comments
* feat: segment manifest optimizing state with lease
* fix: clippy
* fix: exhaustive manifest state matching
* fix: merge manifest rebuilds under the write lock
* refactor: named state predicates, drop redundant enumerator test
Review follow-up: move the enumerator's filter into
SegmentManifestState::is_usable, name the preserving predicate
is_optimizer_mark, delete the enumerator test that re-tested serde.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat: read-only edge grouping/matrix search + object-storage read path
Add query_groups (group by a payload field) and search_matrix (single-shard distance matrix over a random sample) to the read-only edge shard's EdgeShardRead API, in new grouping and matrix modules plus edge test helpers.
Make the object-storage read path available outside tests: drop the #[cfg(test)] gate on the BlobFile UniversalReadExt impl and move io_bridge_object_store/object_store to segment's normal dependencies, so a ReadOnlyEdgeShard can serve segments read from S3.
* Share group-by building blocks between server and edge
Move GroupsAggregator, group candidate query shaping (is-empty filter,
group_by payload selector, prefetch limit scaling) and result-order
derivation into shard::grouping / shard::query, so the collection and
edge grouping implementations cannot silently diverge.
Edge grouping now handles multi-valued group keys, u64 keys, wildcard
group_by paths, prefetch limits and score-ordered groups the same way
as the server.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Drive group-by through a shared sans-IO state machine
Extract the multi-request collect/fill loop into
shard::grouping::GroupByDriver: next_request() yields shaped backend
queries, add_points() advances the state, distill() returns the groups.
Query execution stays with the caller, so the async server path and the
sync edge path drive the same machine, and the request shaping helpers
become private to shard::grouping.
Edge now uses the same request budget (5 collect + 5 fill requests) and
per-request candidates limit (groups * group_size, computed inside the
driver) as the server, replacing its single 4x-oversampled request.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* chore: update object_store to 0.14.0
* fix: resolve AWS S3 default credentials from the environment
The default path built the client with AmazonS3Builder::new(), which reads nothing, so object_store's resolver always fell through to EC2 IMDS (the node role) — AWS_WEB_IDENTITY_TOKEN_FILE / AWS_ROLE_ARN (EKS IRSA), ECS and EKS Pod Identity were silently ignored.
Seed the builder with from_env() for the default chain so web identity (IRSA), container credentials and AWS_* keys are honoured; explicit bucket/region/endpoint from config still override anything from_env picks up. Static credentials keep using new() with the provided keys.
* perf: read whole objects in a single GET instead of HEAD+GET
* feat: read object tail from offset to EOF in a single GET
Generalize the whole-object single-GET read so a tail read from a known
offset also avoids the separate len()/HEAD round-trip. The backend
primitive becomes `read_from(path, from)`, issuing an open-ended
`GetRange::Offset(from)` GET (or a plain GET for from == 0) and reporting
the object's total size from the response; `read_whole` is now a thin
wrapper over it.
The owned blob pipeline's `schedule_whole` no longer HEADs the remote to
size a tail read. An offset at/past EOF is an unsatisfiable range (HTTP
416) rather than an empty body, so the buffer builder disambiguates with
a single len() only on the error path, yielding an empty read when the
tail is genuinely empty. The disk cache's reopen prefiller is made
tolerant of that empty tail so it no longer truncates the local mirror.
Also split the now-hard-to-follow simple_disk_cache `file.rs` into a
`file/` module (type/state, init state machine, reopen, read surface)
and document the FromScratch vs Prefiller init sources.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: immutable field index live_reload
* feat: read-only segment live_reload
* fix: linter
* fix: linter
* fix: review comments
* fix: linter
* feat: ReadOnlyVectorData::live_reload covering all components
Extract the per-vector reload into a dedicated method that destructures
`ReadOnlyVectorData` so storage, index and quantized vectors are all
covered. Adding a field without reloading it won't compile, which guards
against silently skipping a component. The segment orchestrator now just
delegates per named vector.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: keep id-tracker delta pending until all reloads succeed
ReadOnlySegment::live_reload drained the id-tracker delta (which advances
tracker state and cannot be replayed) and only then ran the fallible
payload/vector reloads. On error the delta was lost, so un-updated
components drifted out of sync permanently.
Accumulate the delta into a new `pending_reload` field and clear it only
once every component has reloaded successfully. On a later reload the
tracker's fresh delta is folded in via `LiveReloadResult::merge` and the
union is replayed, so a partial failure self-heals. Offsets are monotonic,
so the only merge conflict is an inserted-but-unapplied offset later
deleted: it is dropped from `inserted` and kept in `deleted` so a
partially-applied component drops it on replay.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: add open for read-only sparse vector index + enum sparse dispatcher
* fix: universal-IO loads for read-only sparse index open
* refactor: rename load_via/open_via to load_universal/open_universal
* refactor: drop StorageVersion::load in favor of load_universal
The regular-IO `load` duplicated `load_universal` over plain `File` IO.
Remove it and route all callers through `load_universal(&MmapFs, ..)`.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(sparse): decouple inverted index from concrete storage; split segment_constructor_base (#9461)
Make the read-only index enum generic over storage `S` (no concrete MmapFile),
and remove construction callbacks from the sparse index open paths.
- InvertedIndex is now a pure read/search trait: `open`/`from_ram_index`/
`type Fs` moved off the trait to inherent methods on each concrete index
type, so construction no longer requires `S::Fs: Default`.
- SparseVectorIndex open split into a generic `plan` (load-vs-build decision +
RAM-index build) and generic `finish` (assembly); callers do the concrete
per-type construction, so no construction callbacks are needed.
- ReadOnlySparseVectorIndex::open takes the already-constructed inverted index
and caller-loaded config instead of a `load_inverted_index` callback.
- VectorIndexReadEnum is generic over `S: UniversalRead`; sparse mmap variants
hold `InvertedIndexCompressedMmap<_, S>` rather than a concrete `MmapFile`.
- Split the 1182-line segment_constructor_base.rs into a module (paths,
vector_storage, payload_storage, id_tracker, vector_index,
sparse_vector_index, create_segment, segment, legacy_state); the sparse
dispatcher's match arms collapse into three per-family helpers.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat: LiveReload (no-op) dispatch for read-only vector index enum (#9436)
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: live_reload for read-only immutable id tracker + enum dispatch
* deleted should be already sorted
---------
Co-authored-by: Luis Cossío <luis.cossio@outlook.com>
* feat: add open to read-only HNSW index
* fix: dedicated universal-IO load for read-only graph
* fix: report actual graph residency in read-only is_on_disk
* refactor: rename load_via to load_universal
* refactor: unify graph links loading on universal IO
Replace the dual GraphLinks load paths (regular mmap + universal IO) with a
single universal-IO loader, and the `Mmap` GraphLinksEnum variant with a
`Universal` variant that keeps a type-erased `UniversalRead` handle alive
behind `Box<dyn GraphLinksStorage>` (UniversalRead is not object-safe).
- GraphLinksEnum::from_storage picks the variant from UniversalKind:
borrowable (mmap-backed) kinds stay `Universal`, others are materialized
into `Ram`, so the borrowability invariant in GraphLinksStorage::bytes
holds by construction.
- Loading takes a `Populate` parameter, derived from `hnsw_config.on_disk`:
on_disk -> Populate::No (lazy), otherwise Populate::Blocking.
- is_on_disk is reported from config again instead of the enum variant.
- Split graph_links.rs into a graph_links/ module (format, vectors, storage,
links, tests) with a relationship diagram in mod.rs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor: drop LoadOption in favor of a Populate parameter
After unifying the link load paths on universal IO, LoadOption only encoded a
populate choice with a fs that was always MmapFs. Replace it with a `Populate`
argument to GraphLayers::load, removing the enum, its constructors, the
load_links helper, and the unused generic backend.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: live_reload for read-only quantized vectors
* Propagate live_reload params through quantized chunked storage
Thread fs, deleted_points, new_points and hw_counter through the full
quantized live_reload chain instead of synthesizing empty deltas and a
disposable hardware counter at the leaf storages.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: use LiveReload trait in chunked mmap
* fix: use LiveReload trait
* fix: compiler errors
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: live_reload for read-only sparse vector storage
* fix: pass multi_vector_config in multi read-only live_reload test
* feat: live_reload dispatch for VectorStorageReadEnum
* feat: live_reload for ChunkedVectorsRead
* feat: live_reload for read-only chunked dense/multi storage
* fix: full reread
* fix: linter
* refactor: chunked vectors live_reload as trait impl
* refactor: route chunked-vectors reads through the Fs backend
Use the `fs: &S::Fs` argument for filesystem access in the read-only
chunked-vectors path instead of going through `fs_err` (direct local FS),
so non-local backends are addressed correctly:
- `read_chunks_from`: list chunk files via `Fs::list_files` with the
`chunk_` prefix instead of `fs_err::read_dir` over the whole directory,
avoiding enumeration of unrelated files.
- `read_status_len`: read the status file via `read_whole_via(fs, ...)`
instead of `fs_err::read`; thread `fs` through its `open`/`live_reload`
call sites.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fmt
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: read-only open for ReadOnlyChunkedDenseVectorStorage
Add ReadOnlyChunkedDenseVectorStorage::open, the read-only counterpart of
open_appendable_memmap_vector_storage_impl: open the chunked vectors/
directory via ChunkedVectorsRead::open and materialize the deleted/ flags
into an owned bitvec via DynamicStoredFlags::load_bitvec, threading every
file open through the UniversalRead backend without writing anything.
Share the vectors/ and deleted/ directory names with the writable storage
by exposing them pub(crate).
Covered by a round-trip test: write and delete through the writable
storage, reopen read-only, and assert vectors, deletions and counts match.
Part of the read-only vector-storage open constructors (#9241).
* feat: read-only open for ReadOnlyChunkedMultiDenseVectorStorage
Add ReadOnlyChunkedMultiDenseVectorStorage::open, the read-only counterpart
of open_appendable_memmap_multi_vector_storage_impl: open the chunked
vectors/ and offsets/ directories via ChunkedVectorsRead::open and
materialize the deleted/ flags into an owned bitvec via
DynamicStoredFlags::load_bitvec, threading every file open through the
UniversalRead backend without writing anything.
Share the vectors/, offsets/ and deleted/ directory names with the writable
storage by exposing them pub(crate).
Covered by a round-trip test: write and delete multivectors through the
writable storage, reopen read-only, and assert per-point multivector
contents, deletions and counts match.
Part of the read-only vector-storage open constructors (#9241).
* feat: read-only open for ReadOnlySparseVectorStorage
Add ReadOnlySparseVectorStorage::open, the read-only counterpart of
MmapSparseVectorStorage::open: open the Gridstore store/ directory via
GridstoreReader::open and materialize the deleted/ flags into an owned
bitvec via DynamicStoredFlags::load_bitvec, threading every file open
through the UniversalRead backend without writing anything.
next_point_offset is reconstructed the same way the writable storage does
on reopen (highest deleted id or the Gridstore pointer count).
Share the store/ and deleted/ directory names with the writable storage by
exposing them pub(crate).
Covered by a round-trip test: write and delete sparse vectors through the
writable storage, reopen read-only, and assert live contents, deletions and
the reconstructed point count match.
Part of the read-only vector-storage open constructors (#9241).
* feat: open dispatcher for VectorStorageReadEnum
Add VectorStorageReadEnum::open, the read-only counterpart of
open_vector_storage: route a VectorDataConfig to the matching read-only
variant. The storage type selects the on-disk layout (mmap -> immutable
DenseVectorStorageImpl, chunked-mmap -> appendable chunked storage), the
datatype selects the element type, and a multivector config routes to the
chunked multi-dense storage (mmap multivectors are appendable-only).
advice/populate are derived from the storage type exactly as the writable
path does.
Expose open_dense_vector_storage_impl as pub(crate) so the dispatcher can
build the immutable Dense* variants. Sparse storage keeps its own
ReadOnlySparseVectorStorage::open and is wrapped at the call site.
Covered by tests asserting each config routes to the expected variant and
round-trips.
Part of the read-only vector-storage open constructors (#9241).
* fix: false values update
* fix: linter
* fix: ci/cd
* chore: add comment
* fix: reopen and update trues_count and falses_count
* feat: implement LiveReload for null index
Reload the null index incrementally through the roaring-flags LiveReload trait: reopen the has_values/is_null bitslices and re-read only the changed points, then grow total_point_count, instead of re-materializing both bitmaps via a full open.
* chore: trigger ci
* feat: read-only open for ReadOnlyChunkedDenseVectorStorage
Add ReadOnlyChunkedDenseVectorStorage::open, the read-only counterpart of
open_appendable_memmap_vector_storage_impl: open the chunked vectors/
directory via ChunkedVectorsRead::open and materialize the deleted/ flags
into an owned bitvec via DynamicStoredFlags::load_bitvec, threading every
file open through the UniversalRead backend without writing anything.
Share the vectors/ and deleted/ directory names with the writable storage
by exposing them pub(crate).
Covered by a round-trip test: write and delete through the writable
storage, reopen read-only, and assert vectors, deletions and counts match.
Part of the read-only vector-storage open constructors (#9241).
* feat: read-only open for ReadOnlyChunkedMultiDenseVectorStorage
Add ReadOnlyChunkedMultiDenseVectorStorage::open, the read-only counterpart
of open_appendable_memmap_multi_vector_storage_impl: open the chunked
vectors/ and offsets/ directories via ChunkedVectorsRead::open and
materialize the deleted/ flags into an owned bitvec via
DynamicStoredFlags::load_bitvec, threading every file open through the
UniversalRead backend without writing anything.
Share the vectors/, offsets/ and deleted/ directory names with the writable
storage by exposing them pub(crate).
Covered by a round-trip test: write and delete multivectors through the
writable storage, reopen read-only, and assert per-point multivector
contents, deletions and counts match.
Part of the read-only vector-storage open constructors (#9241).
* feat: read-only open for ReadOnlySparseVectorStorage
Add ReadOnlySparseVectorStorage::open, the read-only counterpart of
MmapSparseVectorStorage::open: open the Gridstore store/ directory via
GridstoreReader::open and materialize the deleted/ flags into an owned
bitvec via DynamicStoredFlags::load_bitvec, threading every file open
through the UniversalRead backend without writing anything.
next_point_offset is reconstructed the same way the writable storage does
on reopen (highest deleted id or the Gridstore pointer count).
Share the store/ and deleted/ directory names with the writable storage by
exposing them pub(crate).
Covered by a round-trip test: write and delete sparse vectors through the
writable storage, reopen read-only, and assert live contents, deletions and
the reconstructed point count match.
Part of the read-only vector-storage open constructors (#9241).
* feat: read-only open for ReadOnlyChunkedDenseVectorStorage
Add ReadOnlyChunkedDenseVectorStorage::open, the read-only counterpart of
open_appendable_memmap_vector_storage_impl: open the chunked vectors/
directory via ChunkedVectorsRead::open and materialize the deleted/ flags
into an owned bitvec via DynamicStoredFlags::load_bitvec, threading every
file open through the UniversalRead backend without writing anything.
Share the vectors/ and deleted/ directory names with the writable storage
by exposing them pub(crate).
Covered by a round-trip test: write and delete through the writable
storage, reopen read-only, and assert vectors, deletions and counts match.
Part of the read-only vector-storage open constructors (#9241).
* feat: read-only open for ReadOnlyChunkedMultiDenseVectorStorage
Add ReadOnlyChunkedMultiDenseVectorStorage::open, the read-only counterpart
of open_appendable_memmap_multi_vector_storage_impl: open the chunked
vectors/ and offsets/ directories via ChunkedVectorsRead::open and
materialize the deleted/ flags into an owned bitvec via
DynamicStoredFlags::load_bitvec, threading every file open through the
UniversalRead backend without writing anything.
Share the vectors/, offsets/ and deleted/ directory names with the writable
storage by exposing them pub(crate).
Covered by a round-trip test: write and delete multivectors through the
writable storage, reopen read-only, and assert per-point multivector
contents, deletions and counts match.
Part of the read-only vector-storage open constructors (#9241).
Add ReadOnlyChunkedDenseVectorStorage::open, the read-only counterpart of
open_appendable_memmap_vector_storage_impl: open the chunked vectors/
directory via ChunkedVectorsRead::open and materialize the deleted/ flags
into an owned bitvec via DynamicStoredFlags::load_bitvec, threading every
file open through the UniversalRead backend without writing anything.
Share the vectors/ and deleted/ directory names with the writable storage
by exposing them pub(crate).
Covered by a round-trip test: write and delete through the writable
storage, reopen read-only, and assert vectors, deletions and counts match.
Part of the read-only vector-storage open constructors (#9241).
* feat: implement LiveReload for read-only roaring flags
Make ReadOnlyRoaringFlags<S> implement the shared LiveReload trait (open retains the StoredBitSlice; reopen + read only the changed positions on reload), threading the backend S through the read-only bool/null index structs, the field-index enum and the facet enum. The per-index reload impls live in the bool/null branches stacked on top.
* fix: review commits
* feat: add remark comment
* feat: implement LiveReload for numeric index
* fix: missing remove_point method
* fix: remove comments
* fix: bad var usage
* chore: trigger ci
* fix: dev rebase
* feat: wire up read only field indexes
* feat: add missing indexes
* feat: add try_from impl for TextIndexParams
* refactor: unify read-only field index open and drop RocksDB storage
Merge ReadOnlyFieldIndex::open_gridstore/open_mmap into a single `open`
that picks the appendable vs immutable path from the stored
FullPayloadIndexType::storage_type. The choice is modeled as a ReadMode
(Appendable/Immutable) rather than a concrete backend, since the
read-only stack is generic over UniversalRead and mmap is now just one
implementation of it.
Also remove the unsupported RocksDb variant from payload_config's
StorageType and its now-dead match arms.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: linter
* fix: linter
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: add open for read only numeric index
* feat: add parent open_*, get_mutability_type for ReadOnlyNumericIndex
* fix: naming and generic fs
* feat: detect not-found via error in numeric index open
* fix: hw counter write -> read
* feat: detect not-found via error in bool, null and geo index open
* fix: error on inconsistent storage in bool and null index open
* fix: hw counter write -> read
* feat: add readonly null index open
* feat: add open method
* refactor: split variants into modules
* chore: remove duplicated impl
* feat: remove generic in Roaring flag move into open method
* feat: add open methods for map
* fix: review comments
* feat: detect not-found via error instead of path.exists in map index open
Replace the path.exists() pre-check in the read-only appendable map index
open path with error-driven detection through ok_not_found(). Generalize
OkNotFound over an IsNotFound trait, implemented for io::Error, mmap::Error,
UniversalIoError, and GridstoreError so a NotFound surfacing through any
layer (including mmap's inner io::Error) maps to Ok(None).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* review: revert internal error conversion
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: add io_bridge
* fix: cached dispatcher
* feat: add open_with_handle
* fix: naming
* fix: wording
* fix: rebase io_bridge onto split BorrowedReadPipeline/OwnedReadPipeline
* fix: linter
* feat: add S3 backend
* chore: remove io_bridge
* fix: linter
* fix: linter
* fix: read handle
* chore: simplify S3Source
* fix: s3 test
* feat: support multi runtime
* fix: clippy errors
* fix: review comments
* feat: add io design
* feat: add S3 backend
* chore: fix docs
* fix: dev changes
* chore: add some docs
* chore: remove explicit type
* feat: add new methods
* [WIP] review refactor
* fmt
* fix: bytes alignment
* fix: linter
* feat: remove Bytes
* fix: ci/cd
* fix: tests
* fix: tests
* refactor: simplify io_bridge pipeline to Handle-based dispatch
Replace the BridgeRuntime worker thread + request channel + boxed
BridgeRequest with direct tokio Handle usage:
- BridgeRuntime is now just an Arc<Runtime>; schedule() spawns the read
future via Handle::spawn instead of routing it through a dispatcher
thread. Removes BridgeRequest and the now-unreachable S3RuntimeShutDown
error variant.
- Guard against a panicking read task hanging wait() forever: the spawned
task catches unwinds and converts them into a TaskPanicked error reply,
so every scheduled slot is always answered.
- Encapsulate slot bookkeeping in PendingSlots, exposing only the needed
operations instead of a public map + counter.
- Split the grown pipeline.rs into a pipeline/ module (slots / inner /
borrowed / owned), de-duplicating the shared read-future construction.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor: move pipeline buffer ownership into the read future
Instead of the pipeline owning the destination Vec<T> in a slot map and
the future writing through a SendBytePtr raw pointer, let the future
allocate the buffer itself and return it through BridgeResponse. The
buffer crosses the worker-thread boundary as a normal move via the reply
channel, wrapped in a SendableVec<T> newtype that asserts Send for
T: bytemuck::Pod only.
This removes the entire unsafe SendBytePtr apparatus from the pipeline:
no raw pointer, no unsafe fn, no per-call-site unsafe blocks, no
heap-stability invariants. The only remaining unsafe in the crate is one
bounded `unsafe impl<T: Pod> Send for SendableVec<T>` with a trivially
true invariant (Pod types are plain bytes).
PendingSlots collapses back to PendingSlots<U>: slots no longer carry
buffers. AlignedBufWriter::from_raw_bytes (used only by the SendBytePtr
path) and its test are removed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor: drop SendableVec wrapper now that Item: Send
UniversalRead's element type is now bound to `Item` (`Pod + Send`),
so the io_bridge pipeline no longer needs a hand-rolled `Send` wrapper
around `Vec<T>` to ship buffers through the reply channel. Replace
`SendableVec<T>` with `Vec<T>` end-to-end and tighten the local impls
to `T: Item`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat: implement UniversalRead::reopen for BlobFile
BlobFile has no cached file metadata or mapping — `len()` queries the
object store fresh on each call — so reopen is a no-op, matching the
io_uring impl.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat: add new fs impl for Blob
* fix: is_in_ram_or_mmap for S3
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor: read only full text
* fix: linter
* refactor: follow new file structure
* chore: remove storage enums
* fix: linter
* fix: FullTextReadIndex methods
* chore: remove duplicated impls
* fix: format
* refactor: unify FullTextIndex telemetry via trait default method
Add a `get_telemetry_data` default method to `FullTextIndexRead` built
from the existing telemetry methods, so the `ReadOnlyFieldIndex` match
arm is consistent with every other variant.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor: move full-text lifecycle methods off FullTextIndexRead trait
`populate`, `clear_cache`, `files` and `immutable_files` are lifecycle
concerns, not part of the read surface. Move them from the
`FullTextIndexRead` trait to inherent methods alongside the rest of each
index's lifecycle code (open / wipe / flusher). Also split the
`mutable_text_index` tests into their own `tests.rs`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor: restructure full_text_index module layout
Move the `FullTextIndex` enum into `full_text_index/mod.rs` and split its
impls following the inner-module convention: lifecycle (constructors,
builders, ValueIndexer, PayloadFieldIndex) into `lifecycle.rs`, read-path
impls and shared free functions into `read_ops.rs`. Extract the
`FullTextIndexRead` trait into its own `full_text_index_read.rs`.
Drops the inherent `FullTextIndex::get_telemetry_data`, which duplicated
the `FullTextIndexRead::get_telemetry_data` default method.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat: read-only counterpart of MutableFullTextIndex
Extract the shared in-memory state of `MutableFullTextIndex` into
`MutableFullTextIndexInner` (inverted index, config, tokenizer), which
implements `FullTextIndexRead`. `MutableFullTextIndex` now wraps it
alongside a writable `Gridstore`.
Add `ReadOnlyAppendableFullTextIndex<S>`, wrapping the same inner state
with a `GridstoreReader` over generic `UniversalRead`. Mirrors the
`MutableMapIndex` / `ReadOnlyAppendableMapIndex` split; loading and
lifecycle land in a follow-up.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor: turn ReadOnlyFullTextIndex into a two-variant enum
Rename the `read_only_text_index` module to `read_only` and reshape
`ReadOnlyFullTextIndex` from a thin `MmapFullTextIndex` wrapper into an
`Appendable` / `Immutable` enum, mirroring `ReadOnlyMapIndex`. Its
`FullTextIndexRead` impl is now a dispatcher forwarding to the active
variant.
Drop the no-op inherent `populate` (mutable, immutable) and
`immutable_files` (mutable) methods that tripped `clippy::unused_self`
and `clippy::unnecessary_wraps` after the earlier trait-to-inherent
move; `FullTextIndex::populate` inlines the in-RAM no-op directly.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: generall <andrey@vasnetsov.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor: read only geo index
* fix: linter
* fix: drop open temp
* fix: read_ops
* fix: address review fixes and use same file structure like map index
* fix: review comments
* fix: linter
* feat: forward api keys header
* fix: linter
* feat: add api_keys to grpc
* fix: imports
* fix: extract api keys from grpc headers
* refactor: make apikeys extracor more compact
* fix: linter issue
* feat: update header key logic
* fix: clippy issues
* make external keys a part of InferenceParams::new argument, so we wont forget using it
* avoid cloning metadata on header extraction
* combine inference token with ext api key for simpler usage
* clippy
---------
Co-authored-by: generall <andrey@vasnetsov.com>