in grpc:
- restructure `ContextInput`, so that it becomes easy to handle both pairs and positive/negative lists in the future
in rest:
- restructure context queries so that they don't repeat the label in order to be used . E.g. turn `{ query: { context: { context: [ ... ] } } }` into `{ query: { context: [ ... ] } }`
- fixes fusion so that `{query: { fusion: "rrf" } }` actually works
both:
- renames `RecommendInput` positives and negatives to singular (same as in reco api)
* add `query` to shard trait
* add missing conversions for query
* update grpc docs
* Query response has intermediate results
* add ShardQueryResponse description
* move pub use to the top, keep only one way of reaching reexports
* introduce QueryContext, which accumulates runtime info needed for executing search
* fmt
* propagate query context into segment internals
* [WIP] prepare idf stats for search query context
* Split SparseVector and RemmapedSparseVector to guarantee we will not mix them up on the type level
* implement filling of the query context with IDF statistics
* implement re-weighting of the sparse query with idf
* fmt
* update idf param only if explicitly specified (more consistent with diff param update
* replace idf bool with modifier enum, improve further extensibility
* test and fixes
* Update lib/collection/src/operations/types.rs
Co-authored-by: Arnaud Gourlay <arnaud.gourlay@gmail.com>
* review fixes
* fmt
---------
Co-authored-by: Arnaud Gourlay <arnaud.gourlay@gmail.com>
* Add grey color for collection info status
* Restructure locks in local shard info method
* Set collection status to grey if we have pending optimizations
* Update OpenAPI specification and gRPC documentation
* Set optimizer status instead of color for compatibility reasons
* Only set and check for grey status if there are no other statuses
* Change `ClockMap` to always reject tick 0
* Enable rejecting operations in `LocalShard::update`...
...and propagate updated `clock_tick` back to the caller
* WIP: Retry operation on *all* nodes (with new tick)...
...if *any* node rejected the operation
* fixup! WIP: Retry operation on *all* nodes (with new tick)...
Keep the same `clock` for the whole duration of `update`
* Add update response status for clock rejection (#3725)
* Add clock rejected update status
* Don't hide internal items in docs
* Propagate clock rejected status if operation clock is too old
* Fix test in which we manually define clock ticks
* Expect forward proxy rejection, node already has that operation
* Add constant for first valid clock tick, init clocks in recovery point
* In testing, advance all zero clocks to start ticking from 1
* Attempt to fix clock_set_clock_map_workflow
* Rename WalError::Rejected to ClockRejected
* Limit retry attempts when operation is rejected for an old clock tick
* Don't log rejected clock tag if tick was 0
* Cleanup `ClockMap`/`RecoveryPoint`
* Cleanup `ClockSet` and `ShardReplicaSet::update`
* Fix `clock_set_clock_map_workflow` test
* add debug assertions
* fmt
---------
Co-authored-by: Tim Visée <tim+github@visee.me>
Co-authored-by: timvisee <tim@visee.me>
Co-authored-by: generall <andrey@vasnetsov.com>
* When restarting shard transfer, use sync if old or new one had sync set
* refactor restart transfer models
---------
Co-authored-by: generall <andrey@vasnetsov.com>
* Add consensus operation to restart shard transfer with different config
* Require shard transfer restart to have a changed configuration
* implement api
---------
Co-authored-by: generall <andrey@vasnetsov.com>
* Add first stubs for WAL delta shard transfer method
* Repurpose queue proxy, use it for transferring WAL diff as well
* Integrate WAL delta transfer is transfer selection logic
* Add WalDelta shard transfer type which is not exposed in public API
* Basic implementation of falling back to stream records transfer
* Share await_consensus_sync function
* Ask remote shard for recovery point
* During WAL delta transfer, resolve shard diff locally for recovery point
* Rebase on latest dev, support empty WAL diff
* Rebase on latest dev, support empty WAL diff
* Use partial snapshot state for WAL delta transfer
* Set cutoff point on remote shard after shard WAL delta transfer
* Set cutoff point on remote shard after stream records transfer
* During WAL delta transfer, set s tate from partial snapshot to partial
* Describe WAL delta transfer in a comment
* Allow updating cutoff point in stream records transfer to fail
* Do not set cutoff point on remote shard on WAL delta transfer
* Make await consensus sync logic easier to read and reason about
* Fix fallback to other shard transfer method on WAL delta transfer fail
* Various minor improvements
* Add TODO for just ignoring API unimplemented errors
* Only allow stream records cutoff point error if remote is older version
* Allow switching to partial to fail when falling back
* Only change shard state to partial if not in partial state already
* Change default shard transfer method back to stream records
* Add important TODO back
* Add WAL delta shard transfer method in gRPC
* Prefer configured shard transfer method as default
* Extract shard transfer fallback logic into separate function
* Add recovery shard replica set state
* Accept forced operations in partial snapshot state
* Fix switching into wrong state
* Add some helpful comments
* Deduplication in match statement
* Support set by key in low level.
* Rename key field.
* Format.
* Pass key.
* Format.
* Test.
* Clippy.
* Fix ci lint.
* Check grpc consistency.
* Update openapi.
* Fix empty key test case.
* Support array index.
* Format.
* Add test for non exists key.
* Clippy fix.
* Add idempotence test.
* Update index by updated payload.
* Add ut for utils.
* Add ut for 1 level key.
* Fix ut.
* Support no exits key.
* Fix test result.
* Fix after rebase
* handle wildcart insertion into non-existing array
* avoid double read of payload during update
* fix missing removing data from index in case if set_payload removes indexed field
---------
Co-authored-by: Shylock Hg <shylock@DESKTOP-40I855A>
Co-authored-by: Albert Safin <xzfcpw@gmail.com>
Co-authored-by: generall <andrey@vasnetsov.com>
* Add min_should field in Filter struct
* min_should clause checks whether at least given number (min_count) of conditions are met
* modify test cases due to change in Filter struct (set min_should: None)
* add simple condition check unit test
* docs, cardinality estimation, grpc not implemented yet
* Add min_should field in Filter struct
* min_should clause checks whether at least given number (min_count) of conditions are met
* modify test cases due to change in Filter struct (set min_should: None)
* add simple condition check unit test
* Impl min_should clause in REST API
* perform cardinality estimation by estimating cardinalities of intersection and combining as union
* add openapi spec with docs update
* add integration test
* Impl min_should clause in gRPC
* Cargo fmt & clippy
* Fix minor comments
* add equivalence test between min_should and must
* shortcut at min_count matches
* use `Filter::new_*` whenever possible
* Add missing min_should field
* Fix gRPC field ordering & remove deny_unknown_fields
* Empty commit
---------
Co-authored-by: Luis Cossío <luis.cossio@outlook.com>
* support \`start_from\`: DateTime
* add order by datetime test
* generate openapi models and grpc docs
* fixup after rebase
* allow string representation of datetime in grpc
* add TODO
* fix `.start_from()`
* use custom deserialization on datetime
* first PR implementation (#2865)
- fetch offset id
- restructure tests
- only let order_by with numeric
- introduce order_by interface
cargo fmt
update openapi
calculate range to fetch using offset + limit, do some cleanup
enable index validation, fix test
Fix pagination
add e2e tests
make test a little more strict
select numeric index on read_ordered_filtered
add filtering test 🫨
fix filtering on order-by
fix pip requirements
add grpc interface, make read_ordered_filtered fallible
fmt
small optimization of `with_payload` and `with_vector`
refactor common logic of point_ops and local_shard_operations
Make filtering test harder and fix limit for worst case
update openapi
small clarity refactor
avoid extra allocation when sorting with offset
stream from numeric index btree instead of calculating range
use payload to store order-by value, instead of modifying Record interface
various fixes:
- fix ordering at collection level, when merging shard results
- fix offset at segment level, to take into account also value offset
- make rust tests pass
remove unused histogram changes
fix error messages and make has_range_index exhaustive
remove unused From impl
Move OrderBy and Direction to segment::data_types::order_by
Refactor normal scroll_by in local_shard_operations.rs
More cleanup + rename OrderableRead to StreamWithValue
empty commit
optimization for merging results from shards and segments
fix case of multi-valued fields
fix IntegerIndexParams name after rebase
precompute offset key
use extracted `read_by_id_stream`
Expose value_offset to user
- rename offset -> value_offset
- extract offset value fetching logic
* remove offset functionality when using order_by
* include order_by in ForwardProxyShard
* extra nits
* remove histogram changes
* more nits
* self review
* resolve conflicts after rebase, not enable order-by with datetime index schema
* make grpc start_from value extendable
* gen grpc docs
---------
Co-authored-by: kwkr <kawka.maciej.93@gmail.com>
Co-authored-by: generall <andrey@vasnetsov.com>
* Merge serde attributes
* Remove obsolete conversion
* Add integer type with parameters
* Make integer lookup and range parameters non-optional
* Add parameterized integer index types test
Co-authored-by: Di Zhao <diz@twitter.com>
* Cleanup
---------
Co-authored-by: Di Zhao <diz@twitter.com>
* add checksum to SnapshotDescription
* implement storing snapshot checksums in a file
* Don't serialize checksum if it's None for backwards compatibility
* Remove hex dependency, use Rust std formatter for this
* Do not error if we cannot remove checksum file for snapshot
Some snapshots may not have a corresponding checksum file. Maybe it was
created in an older Qdrant version that didn't have support for this, or
a user hasn't provided any.
* Add debug message when hashing snapshot, can be expensive on large files
* Inline debug messages
* Add checksum to shard snapshots
* If creating snapshot fails, delete snapshot target and checksum file
* Use Rust idiomatic ok() and improve debug messages
* Use correct snapshot checksum paths, clean up after shard snapshot
* Use better path type in get_checksum_path
---------
Co-authored-by: timvisee <tim@visee.me>
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
* feat: Expose git commit id in the health check endpoint
* fix: CI errors
* test: Add test for health check api
* feat: Add / endpoint to openapi schema
* Make git commit hash optional
* ci: Enable debugging setup-protoc action
* Install later protobuf compiler through GitHub Action
* Disable debug mode for setup-protoc job
* refactor: Use commit instead of commit_id
* fix: Use commit instead of commit_id gRPC docs
* test: Update ping API test
* refactor: Rename ping api to root api
---------
Co-authored-by: timvisee <tim@visee.me>
* Make point, indexed point and vector counts optional in collection info
* Update OpenAPI specification and gRPC docs
* Use functional entry API
* Mark point/vector count fields as deprecated
* Use internal collection info structure to remove internal optionals
* Don't deprecate, but mark point counts as approximate
* WIP: collection-level storage for payload indexe scheme
* introduce consensus-level operation for creating payload index
* make operation_id optional in the UpdateResult
* set payload index in newly created shards
* upd api definitions
* include payload index schema into collection consensus state
* include payload index schema into shard snapshot
* review fixes
* create and connect discovery http and grpc interfaces
* add openapi tests
* fix bad rebase
* Add better descriptions
* remove numpy from openapi tests
* fix rebase artifact
* remove already addressed TODO
* add more tests
* 🤡🔫 (cfg batch handler)
* add timeout query param for discover requests
* More gRPC validation
* make fields pydantic_openapi_generator_v3 friendly
* `context_pairs` -> `context` with struct for pairs
* discovery api is only discovery or context,
move struct description to fields
---------
Co-authored-by: timvisee <tim@visee.me>
* add timeout query param for search requests
* enable timeout for recommend requests
* Add query timeout for group by requests
* update openapi models
* Don't decrease timeout after recommend preprocessing
* Add openapi test
* code review
* add timeout to individual group by requests, non-decreasing
* handle timeout for discover
* Update timeout field tag in SearchBatchPoints
message
* Clone inside blocks
* Add shard transfer method to distinguish between batching and snapshots
* Add stub method to drive snapshot transfer
* Store remote shard in forward proxy, merge unproxy methods
* On snapshot shard transfer, create a shard snapshot
* Unify logic for unproxifying forward and queue proxy
* Error snapshot transfer if shard is not a queue proxy
* Add remote shard function to request remote HTTP port
* Handle all specific shard types when proxifying
* Allow queue proxy for some shard holder snapshot methods
* Bring local and remote shard snapshot transfer URLs into transfer logic
* Expose optional shard transfer method parameter in REST and gRPC API
* Expose shard transfer method in list of active transfers
* Fix off-by-one error in queue proxy shard batch transfer logic
* Do not set max ack version for WAL twice, already set when finalizing
* Merge comment for two similar calls
* Use reqwest client to transfer and recover shard snapshot on remote
Using the reqwest client should be temporary. We better switch to a gRPC
call here eventually to use our existing channels. That way we don't
require an extra HTTP client (and dependency) just for this.
* Send queue proxy updates to remote when shard is transferred
* On shard queue transfer, set max WAL ack to last transferred
* Add safe queue proxy destructor, skip destructing in error
This adds a finalize method to safely destruct a queue proxy shard. It
ensures that all remaining updates are transferred to the remote, and
that the max acknowledged version for our WAL is released. Only then is
the queue proxy shard destructed unwrapping the inner local shard.
Our unproxify logic now ensures that the queue proxy shard remains if
transferring the updates fails.
* Clean up method driving shard snapshot transfer a bit
* Change default shard transfer method to stream records
This changes the default transfer method to stream records rather than
using a snaphsot transfer. We can switch this once snapshot transfer is
fully integrated.
* Improve error handling, don't panic but return proper error
* Do not unwrap in type conversions
* Update OpenAPI and gRPC specification
* Resolve and remove some TODOs
* During shard snapshot transfer, use REST port from config
* Always release max acknowledged WAL version on queue proxy finalize
* Rework queue unproxying, transform into forward proxy to handle errors
When a queue or forward proxy shard needs to be unproxified into a local
shard again we typically don't have room to handle errors. A queue proxy
shard may error if it fails to send updates to the remote shard, while a
forward proxy does not fail at all when transforming.
We now transfer queued updates before a shard is unproxified. This
allows for proper error handling. After everything is transferred the
shard is transformed into a forward proxy which can eventually be safely
unproxified later.
* Add trace logging for transferring queue proxy updates in batch
* Simplify snapshot method conversion from gRPC
* Remove remote shard parameter
* Add safe guard to queue proxy handler, panic in debug if not finalized
* Improve safety and architecture of queue proxy shard
Switch from an explicit finalized flag to an outer-inner architecture.
This improves the interface and robustness of the type.
* Do not panic on drop if already unwinding
* Make REST port interface in channel service for local node explicitly
* Recover shard on remote over gRPC, remove reqwest client
* Use shard transfer priority for shard snapshot recovery
* Remove obsolete comment
* Simplify qualified path with use
* Don't construct URLs ourselves as a string, use `parse` and `set_port`
* Use `set_path` when building shard download URL
* Fix error handling in queue to forward proxy transformation
Before, we didn't handle finalization errors properly. If this failed,
tie shard would be lost. With this change the queue proxy shard is put
back.
* Set default shard transfer method to stream records, eliminate panics
* Fix shard snapshot transfer not correctly aborting due to queue proxy
When a shard transfer fails (for any reason), the transfer is aborted.
If we still have a queue proxy shard it should also be reverted, and
collected updates should be forgotten. Before this change it would try
to send all collected updates to the remote, even if the transfer
failed.
* Review fixes
Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>
* Review fixes
Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>
* Initiate forward and queue proxy shard in specialized transfer methods
Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>
* Add consensus interface to shard transfer, repurpose dispatcher (#2873)
* Add shard transfer consensus interface
* Integrate shard transfer consensus interface into toc and transfer logic
* Repurpose dispatcher for getting consensus into shard transfer
* Derive clone
* Mark consensus as unused for now
* Use custom dispatcher with weak ref to prevent Arc cycle for ToC
* Add comment on why a weak reference is used
* Do exhaustive match in shard unproxy logic
* Restructure match statement, use match if
* When queue proxifying shard, allow forward proxy state if same remote
* Before retrying a shard transfer after error, destruct queue proxy
* Synchronize consensus across all nodes for shard snapshot transfer (#2874)
* Move await consensus commit functions into channel service
* Add shard consensus method to synchronize consensus across all nodes
* Move transfer config, channels and local address into snapshot transfer
* Await other nodes to reach consensus before finalizing shard transfer
* Do not fail right away awaiting consensus if still on older term
Instead, give the node time to reach the same term.
* Fix `await_commit_on_all_peers` not catching peer errors properly
* Change return type of `wait_for_consensus_commit` to `Result`
This is of course more conventional, and automatically sets `must_use`.
* Explicitly note number of peers when awaiting consensus
* Before consensus sync, wait for local shard to reach partial state
* Fix timeout error handling when waiting for replica set state
* Wait for replica set to have remote in partial state instead
* Set `(Partial)Snapshot` states for shard snapshot transfer through consensus (#2881)
* When doing a shard snapshot transfer, set shard to `PartialSnapshot`
* Add shard transfer method to set shard state to partial
It currently uses a naive implementation. Using a custom consensus
operation to also confirm a transfer is still active will be implemented
later.
* Add consensus snapshot transfer operation to change shard to partial
The operation `ShardTransferOperations::SnapshotRecovered` is called
after the shard snapshot is recovered on the remote and it progresses
the transfer further.
The operation sets the shard state from `PartialSnapshot` to `Partial`
and ensures the transfer is still active.
* Confirm consensus put shard into partial state, retry 3 times
* Get replica set once
* Add extensive shard snapshot transfer process docs, clean up function
* Fix typo
* Review suggestion
Co-authored-by: Luis Cossío <luis.cossio@qdrant.com>
* Add delay between consensus confirmation retries
* Rename retry timeout to retry delay
---------
Co-authored-by: Luis Cossío <luis.cossio@qdrant.com>
* On replicate shard, remember specified method
---------
Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>
Co-authored-by: Luis Cossío <luis.cossio@qdrant.com>
* Fix paste-bugs in `snapshot_service.proto`
* Add shard snapshot gRCP API definition
* Add validation to shard snapshot gRPC API definition
* Implement conversions between gRPC and `collection` types
* Extract shard snapshot API implementation into common sub-module
* Implement shard snapshot gRPC API
* Generate gRPC docs
* Refactor `ShardSnapshots` gRPC service to be internal API only
* fixup! Refactor `ShardSnapshots` gRPC service to be internal API only
Move `ShardSnapshotRecoverResponse` to `shard_snapshots_service.proto`
* fixup! fixup! Refactor `ShardSnapshots` gRPC service to be internal API only
Update `api/src/grpc/qdrant.rs`
* fixup! fixup! Refactor `ShardSnapshots` gRPC service to be internal API only
Update gRPC docs
* Switch `ShardSnapshots` gRPC service to use `validate_and_log` instead of `validate`