Commit Graph
1429 Commits
Author SHA1 Message Date
Arnaud Gourlay 0d9db0ea05 Optimize process search results (#8163) 2026-02-17 17:42:11 +01:00
xzfc 66d0c36009 Merge io and memory into common (#8155)
* Unify parking_lot/arc_lock feature

* Move lib/common/{io,memory}/* -> lib/common/common/*

- Mmap-related items are grouped into `common::mmap` sub-module:
  - `memory/src/chunked_utils.rs`      -> `common/src/mmap/chunked.rs`
  - `memory/src/madvise.rs`            -> `common/src/mmap/advice.rs`
  - `memory/src/mmap_ops.rs`           -> `common/src/mmap/ops.rs`
  - `memory/src/mmap_type_readonly.rs` -> `common/src/mmap/mmap_readonly.rs`
  - `memory/src/mmap_type.rs`          -> `common/src/mmap/mmap_rw.rs`
- Filesystem-related items are grouped into `common::fs` sub-module:
  - `common/src/fs.rs`          -> `common/src/fs/sync.rs`
  - `io/src/file_operations.rs` -> `common/src/fs/ops.rs`
  - `io/src/move_files.rs`      -> `common/src/fs/move.rs`
  - `io/src/safe_delete.rs`     -> `common/src/fs/safe_delete.rs`
  - `memory/src/checkfs.rs`     -> `common/src/fs/check.rs`
  - `memory/src/fadvise.rs`     -> `common/src/fs/fadvise.rs`
- Rest is moved straight into `common`:
  - `io/src/storage_version.rs` -> `common/src/storage_version.rs`

The old `io` and `memory` are now hollow crates that re-export items
from `common`. These hollow crates will be removed in next commits.

* Replace uses of `io` and `memory` with new paths in `common`

Since `io` and `memory` are just re-exports of `common`, these
replacements are no-op.

* Remove `io` and `memory` crates
2026-02-17 11:01:14 +01:00
Arnaud Gourlay acbfbd3edf Optimize MmapPointToValues from_iter (#8158)
* Optimize MmapPointToValues from_iter

* fix const typo
2026-02-17 11:01:14 +01:00
Andrey Vasnetsov 2f79e8fd15 [manual] skip extra links indexing if vectors are deleted (#8125) 2026-02-16 10:17:29 +01:00
Luis CossíoandArnaud Gourlay bc6246df5a [relevance feedback] add rest and grpc interfaces (#7399)
* add rest and grpc interfaces

Also handle inference for this query

* follow refactor from base branch

* renaming from Ms Cooper

* get started on validation testing

* more validations

* test equivalence with query when less than 2 feedback elements

* rename feedback query to relevance_feedback query

* fix extraction of context pairs

* rename feedback vector to example

* make coderabbit happier

* upd test

---------

Co-authored-by: Arnaud Gourlay <arnaud.gourlay@gmail.com>
2026-02-13 09:55:37 +01:00
Andrey Vasnetsovandtimvisee 2110d456b5 parse round floats as integers on payload indexing (#8100)
* parse round floats as integers on payload indexing

* use backward conversion to validate if float is actualy integerable

* fmt

* Use then_some

---------

Co-authored-by: timvisee <tim@visee.me>
2026-02-13 09:55:36 +01:00
Luis Cossío b88a1d76ce fix: handle score_threshold in Formula queries (#8097)
* handle `score_threshold` in Formula queries

* AI: add openapi test

prompt: Upload 4 points with a numeric payload, use the payload as the
score, and set a score threshold. We can assert which ids should be in
the result and which shouldn't
2026-02-13 09:55:18 +01:00
Andrey Vasnetsov 037600e8af Weighted rrf (#8063)
* weighted rrf implementation

* test

* fmt

* fix edge

* validate number of sources and number of weights

* do not partial match

* upd schema

* review fixes

* update formula

* remove calcualtions from tests

* update comment, because AI have OCD

* fmt
2026-02-10 00:04:33 +01:00
Jojiiandgenerall 65d6b7d016 Don't lock SegmentHolder for the entire duration of read operations (#8056)
* Don't lock SegmentHolder during read operations

* Add comment clarifying eager segment allocation

* review fixes

* Remove TODO

* Shorter locking of segment holder in calculate_local_shard_stats

---------

Co-authored-by: generall <andrey@vasnetsov.com>
2026-02-10 00:03:52 +01:00
Ivan BoldyrevandAndrey Vasnetsov 47da1f6661 Split SegmentEntry into appendable and non-appendable part (#8047)
* Split `SegmentEntry`

Add `ImmutableSegmentEntry` for operations that can be applied to
immutable segments, and make the `SegmentEntry` as it subtrait.

* Rename to `NonAppendableSegmentEntry`

It differs semantically from `ImmutableSegmentEntry` by allowing point
deletion.  Move point deletion to the trait too.

* Fix docstring

* use NonAppendableSegmentEntry where possible

---------

Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
2026-02-10 00:02:32 +01:00
Luis Cossío 876609c4c1 cfg import behind linux flag (#8058) 2026-02-09 23:58:49 +01:00
xzfc 5d72481af3 Callback-based DenseVectorStorage::for_each_in_dense_batch (#8033)
* Generalized vector offset

* Replace get_batch-like methods with for_each-like methods
2026-02-09 23:58:14 +01:00
Daniel Boros 7422d20f19 feat/edge-facet-api (#8045)
* feat: add sync facet api

* feat: add facet-tests

* fix: formatting

* chore: remove exact typing

* feat: reduce code duplication

* fix: add missing type def

* fix: serde default value
2026-02-09 23:57:47 +01:00
Jojii 985885d0b1 Fix test consistency in random memory id tracker (#8043) 2026-02-09 23:56:02 +01:00
d91e8db5f9 Streaming snapshot unpacking (#8025)
* download tar

* compute sha256 for stream download

* wip: propagate unpacking into down to the logic, todo: validation

* unpacked snapshot validation

* Minor tweaks

* Fix typo

* validation during unpack

* cancellation token

* update docstring

* remove redundant dep

* Rearrange unpack functions

- Rename `safe_unpack.rs` into `tar_unpack.rs` so it would be listed
  near `tar_ext.rs` in IDEs.
- Replace calls like `ar = open_snapshot_archive(…); safe_unpack(ar, …);`
  with a single call to `tar_unpack_file(…)`.
- Put calls to `Archive::new(); Archive::set_overwrite(false);` inside
  `tar_unpack_reader` (was `safe_unpack`). So, now it is the only place
  that does `set_overwrite`.

* Let clippy complain if tar::Archive::unpack used

* Mock snapshot download URL

Instead of downloading from storage.googleapis.com every time the test
runs, put small snapshot file to the repo.

The snapshot file is created using this command:

    curl -s \
      https://storage.googleapis.com/qdrant-benchmark-snapshots/test-shard.snapshot \
    | tar \
      --delete segments/4ea958d8-0b64-4312-9a53-0cd857e93535.tar \
      --delete segments/65ac6276-8cca-4f5c-b767-9722190cee8b.tar \
      > lib/storage/src/content_manager/snapshots/test-shard.snapshot

File contents:

    $ tar tf lib/storage/src/content_manager/snapshots/test-shard.snapshot
    wal/
    wal/closed-255
    newest_clocks.json
    replica_state.json
    shard_config.json
    
    $ du -sh lib/storage/src/content_manager/snapshots/test-shard.snapshot
    12K	lib/storage/src/content_manager/snapshots/test-shard.snapshot
    
    $ sha256sum < lib/storage/src/content_manager/snapshots/test-shard.snapshot       
    5d94eac5c1ede3994a28bc406120046c37370d5d45b489a0d2252531b4e3e1f2  -

---------

Co-authored-by: timvisee <tim@visee.me>
Co-authored-by: xzfc <xzfcpw@gmail.com>
2026-02-09 23:53:16 +01:00
Ivan Boldyrev d8ffb098c7 Fix gridstore Option unsoundness (#8011)
* Implement `Optional<T>` type

Define a type with the same presumed layout as `Option<T>`, but with defined behavior.

* Make `transmute_*` functions unsafe

The functions `memory::mmap_ops::transmute_*` are inherently unsafe, but
are not marked as are.  Their usage is documented, but it is not always clear
if the code is correct.

* Add `CsrHeader` to resolve another unsoundness

Tuples have no defined layout.
2026-02-09 23:53:02 +01:00
Tim Visée 0ce9667ab7 Apply clippy suggestions (#8029) 2026-02-09 23:46:02 +01:00
Andrey Vasnetsov f052b3e774 gridstore iter livelock (#7983)
* granular lock for writing and flushing pending updates in gridstore tracker

* avoid grigstore tracker lock in iterator
2026-02-09 23:27:03 +01:00
Ivan Pleshkov 411011477d Remove ChunkedVectorStorage trait (#7977)
* Remove ChunkedVectorStorage trait

* fix rocksdb feature

* fix multidense rocksdb
2026-02-09 23:26:00 +01:00
b2ad969b11 in ram single mmap file (#7971)
* WIP: introduce new vector store type

* handling of InRamMmap

* fmt

* feature-flag

* fmt

* Use if else

Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>

* Update lib/common/common/src/flags.rs

Co-authored-by: Tim Visée <tim+github@visee.me>

* also choose madvise for single-file in-ram-mmap

* simplify generics

* gpu fix

* fix bug

---------

Co-authored-by: Tim Visée <tim+github@visee.me>
Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>
2026-02-09 23:23:45 +01:00
Roman Titov 096f68fb68 Improve SnapshotManifest/SegmentManifest construction (#7961)
* Improve `SnapshotManifest`/`SegmentManifest` construction

* fixup! Improve `SnapshotManifest`/`SegmentManifest` construction

Fix stupidity 😬
2026-02-09 23:22:45 +01:00
krapcys1-maker 6ab5e39f20 fix: improve datetime_range RFC3339 parse error (#7919)
* fix: improve REST datetime_range RFC3339 error message

* docs: add docstrings for datetime_range parsing

* fix: avoid serde_json intermediary for non-human-readable deserialization

* fix: avoid serde_json intermediary for non-human-readable range deserialization

* refactor(segment): reduce Condition JSON routing boilerplate

* fix(segment): preserve binary datetime deserialization compatibility

* fix(segment): preserve binary datetime deserialization

* test(segment): add datetime deserialization regression tests

* chore: rerun CI

* chore: rerun CI (2)

* refactor(segment): simplify datetime deserialization
2026-02-09 23:20:55 +01:00
xzfc 6abd5a1776 Enforce segment UUIDs (#7958)
* Swap docstrings for segment_ids/segment_uuids

Terse internal docs, detailed user-facing docs, not vice versa.

* Enforce segment UUIDs

* Export upcoming segment UUID to telemetry

* Add `optimize_for_test` wrapper

* Tune log message

Reason: the UUID is now guaranteed.

* Improve argument naming/docs
2026-02-09 23:20:44 +01:00
Luis CossíoandTim Visée 1e23057ab8 More graceful handle of flush cancellation (#7798)
* Reapply "return an error when cancelling a flush (#7781)" (#7799)

This reverts commit 4a1103acb7.

* log::trace cancellation of components

* gracefully handle flush cancellation inside segment flusher

* Update lib/segment/src/segment/entry.rs

Co-authored-by: Tim Visée <tim+github@visee.me>

* don't wrap cancellation errors

---------

Co-authored-by: Tim Visée <tim+github@visee.me>
2026-02-09 23:19:37 +01:00
Roman Titov 3b2cdd6b9b Always Copy-on-Write when updating payload in immutable segments (#7952) 2026-02-09 23:17:56 +01:00
Luis Cossío b12abffce9 move cfg statements (#7954) 2026-02-09 23:17:23 +01:00
Kumar Shivenduandtimvisee 14dbd5e4cc Expose segment UUIDs in telemetry response (#7735)
* Add segment ID in telemetry response

* Migrate optimization log segment IDS to use UUIDs

* fmt

* clippy

* Bring back proxy segment SegmentEntry impl

* update openapi spec

* Bring back segment ID

* update comment

* Precompute segment uuid from filesyste while creating segment

* Update return types

* Update OpenAPI spec

* remove segment_ prefix from uuid field and use inline block

* Update OpenAPI spec

* Trigger CI

* clippy

* Improve comments to differentiate between segment id and uuid

* Remove duplicate function that shadowed segment entry function

---------

Co-authored-by: timvisee <tim@visee.me>
2026-02-09 23:16:07 +01:00
Andrey Vasnetsov 7b9f094b13 rollback experimental fix (#7948) 2026-02-09 23:15:37 +01:00
Andrey Vasnetsovandtimvisee 2de3e774f0 async scroll (#7928)
* implementation of async batch vectors reading

* EXPERIMENT: async-io for reading on scroll

* disable on non-linux

* make retrieve sequential for test

* wip: implement vector reading via callback

* simplify operations, remove duplicates

* use batch retrieve also for post-processing search results

* clippy

* fix tests

* review fixes

* Replace big match statement with simple option filter and equal check

* Inline format arguments

---------

Co-authored-by: timvisee <tim@visee.me>
2026-02-09 23:12:40 +01:00
Arnaud Gourlay eb48738cda Optimize SegmentBuilder update with proper Vec sizing (#7921) 2026-02-09 23:11:36 +01:00
Arnaud Gourlay fe3316e265 Optimize allocation in max_available_vectors_size_in_bytes (#7924) 2026-02-09 23:10:01 +01:00
Ivan Pleshkov 8dc1544916 Use p square to find ranges (#7733)
* use p square to find ranges

* use sample size

* add stopper checks

* review remarks

* review remarks
2026-02-09 23:04:25 +01:00
Roman TitovandAndrey Vasnetsov 4c7ed0de44 Implement info request for Qdrant Edge (#7890)
* Implement `edge::Shard::info`

* Add `ReprStr` marker-trait

* Add Python bindings for payload index types

* Implement `PyShard::info`

* fixup! Add Python bindings for payload index types

Add `enable_hnsw` fields

* add __repr__ and exted tests

---------

Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
2026-02-09 23:02:57 +01:00
Andrey Vasnetsov 4ef24e5ee6 log mutable id tracker mapping updates (#7894)
* add log for persisting id-tracker changes

* trace log flush topology

* fmt

* make trace
2026-02-09 22:56:00 +01:00
Roman Titovandgenerall fd18177766 Implement scroll and count requests for Qdrant Edge (#7880)
* Cleanup `shard` crate module declarations

* Move `ScrollRequestInternal` into `shard` crate

* fixup! Move `ScrollRequestInternal` into `shard` crate

Fix imports

* fixup! Move `ScrollRequestInternal` into `shard` crate

`const fn default_*`

* Implement `edge::Shard::scroll`

* fixup! Implement `edge::Shard::scroll`

Re-export `OrderByInterface`

* Cleanup `edge` module declarations

* Cleanup `qdrant-edge-py` module declarations

* Move `PyWithPayload` and `PyWithVector` into `types::query`

* Add `PyScrollRequest` type

* Implement `PyShard::scroll`

* Move `CountRequestInternal` into `shard` crate

* fixup! Move `CountRequestInternal` into `shard` crate

Fix imports

* fixup! Move `CountRequestInternal` into `shard` crate

Rename `default_exact_count` into `CountRequestInternal::default_exact`

* Implement `edge::Shard::count`

* Implement `PyShard::count`

* review: offset for scroll, default values, examples

* ai review

---------

Co-authored-by: generall <andrey@vasnetsov.com>
2026-02-09 22:55:21 +01:00
Arnaud Gourlay abbcab9be6 Improve logs regarding cleanup on startup (#7885) 2026-02-09 22:55:13 +01:00
SapphireandEC2 Default User 37de1f219d Add enable_hnsw option for payload field schema (#7887)
* feat: Add enable_hnsw option for payload field indexes

Add optional enable_hnsw parameter to all payload index types to control
whether additional HNSW graph links are built for each indexed field.

- Add enable_hnsw field to all 8 payload index param types
- Update gRPC proto definitions and conversions
- Update OpenAPI schema
- Modify HNSW graph builder to respect enable_hnsw flag
- Add enable_hnsw() helper methods to PayloadSchemaParams and PayloadFieldSchema
- Update all tests to include new field (default: None)

When enable_hnsw is true and payload_M > 0, additional HNSW links will
be built for the payload field. Default value is true for backward compatibility.

* Fix Some format problems

* fix: address comment problem

---------

Co-authored-by: EC2 Default User <ec2-user@ip-10-78-171-148.ec2.internal>
2026-02-09 22:53:47 +01:00
Ivan Pleshkovandtimvisee a6909aefc1 id tracker persisted mappings offset (#7877)
* id tracker persisted mappings offset

* remove obsolete check

* review remarks

* review remarks

* When persisting mappings, truncate file that is larger than we expect

* If persisting mappings fails, truncate file to what we had before

This isn't necessary, but it is nice to clean up partial mappings.

* Fix typos

* Fix typos

* Just truncate the file, we don't have to seek anymore

* Rename length variable to mappings_expected_len

---------

Co-authored-by: timvisee <tim@visee.me>
2026-02-09 22:53:35 +01:00
xzfc b08c19ed88 Use syncfs (#7883)
* Remove call to Archive::set_sync

Also, this method was the last remaining part of our `tar-rs` fork,
so we can switch to the upstream version now.

* Do syncfs
2026-02-09 22:48:16 +01:00
d685cab358 CoW defines flush topology (#7850)
* CoW defines flush topology

* add explicit logging for staging builds

* clippy

* partial cleanup of topology after flush

* Complete comment

* Format remaining segments in debug assertion

* fix typo

* clarify comment

* Fix comment

* Minor tweaks

---------

Co-authored-by: Tim Visée <tim+github@visee.me>
Co-authored-by: Arnaud Gourlay <arnaud.gourlay@gmail.com>
Co-authored-by: timvisee <tim@visee.me>
2026-02-09 22:46:52 +01:00
Arnaud Gourlay 45bc203a99 Always capture segments flusher first for consistency (#7882) 2026-02-09 22:46:19 +01:00
Andrey Vasnetsov 85d4d1d23c remove rocksdb from creating snapshots path (#7854)
* remove rocksdb from creating snapshots path

* disable rocksdb in tests

* disable rocksdb in tests

* fix clippy

* fmt

* fix clippy again

* less default payload storage types

* fix another test, which assumed rocksdb
2026-02-09 22:39:34 +01:00
Andrey Vasnetsov 906136346b restore snapshot in edge (#7852)
* Restore shard snapshot in Edge python bindings

* method to request snapshot manifest

* move snapshot manifest into Shard crate

* move shapshot manifest reading

* implement inplace update of the shard from snapshot

* fmt

* move shapshot-related functions again, into a dedicated struct

* fmt

* implement partial snapshot recovery for edge

* test for partial recoverying snapshot on edge
2026-02-09 22:36:13 +01:00
xzfc 62f18d5f94 Safe delete (#7830)
* Replace `Option<Segment>` with `enum LoadSegmentOutcome`

* Replace some Path/PathBuf with str/String

* Rename field Segment::{current_path -> segment_path}

* safe_delete
2026-02-09 22:35:57 +01:00
Tim Visée feff43f477 Use RwLock in flushers, drop is alive lock early (#7811)
* Use RwLock for pending changes in MmapBitSliceBufferedUpdateWrapper

* Use RwLock for pending operations in DatabaseColumnScheduledDeleteWrapper

* Use RwLock for pending updates in DatabaseColumnScheduledUpdateWrapper

* Drop alive guard before reconciliation, we don't touch files after

* Remove redundant clone

* Update comments
2026-02-09 22:26:33 +01:00
Roman Titov df20fdbb74 Replace lazy_static with std::sync::LazyLock (#7808) 2026-02-09 22:25:33 +01:00
Tim Visée 1287d74026 Fix reconciliation in more flushers (#7805)
* Reconcile BufferedDynamicFlags flusher

* Reconcile DatabaseColumnScheduledDeleteWrapper flusher

* Rename function for consistency
2025-12-19 12:16:17 +01:00
Tim Visée 66d7022b25 Switch to drain, it is simpler and recommended by docs (#7803) 2025-12-19 12:16:17 +01:00
Luis Cossío 6c8c9179ef Clone pending updates in buffered storages (#7801)
* in MmapSliceBufferedUpdateWrapper

* in MmapBitsliceBufferedUpdateWrapper

* in MutableIdTracker's versions updates

* in MutableIdTracker's mapping updates

* clone updates only when non-empty

* only lock for reconciling pending changes

* simpler reconciling

* use Mutex as argument to ensure we only lock within reconciliation
2025-12-19 12:16:16 +01:00
Tim Visée 66bc5524dc Revert "return an error when cancelling a flush (#7781)" (#7799)
This reverts commit 2f443ec055.
2025-12-18 17:29:24 +01:00