* Add initial gRPC call for requesting WAL recovery point for remote shard
* Add remote shard method to request WAL recovery point
* Add recovery point type in gRPC, use it in recovery point functions
* Add function to extend recovery point with missing clocks from clock map
* Add new gRPC type for recovery point clocks
* Remove atomic loading, because we use regular integers now
* Change order of actix web REST services to fix pattern matching
* Binary search over list of allowed read only patterns
* Add test to ensure the list of patterns is sorted
* first PR implementation (#2865)
- fetch offset id
- restructure tests
- only let order_by with numeric
- introduce order_by interface
cargo fmt
update openapi
calculate range to fetch using offset + limit, do some cleanup
enable index validation, fix test
Fix pagination
add e2e tests
make test a little more strict
select numeric index on read_ordered_filtered
add filtering test 🫨
fix filtering on order-by
fix pip requirements
add grpc interface, make read_ordered_filtered fallible
fmt
small optimization of `with_payload` and `with_vector`
refactor common logic of point_ops and local_shard_operations
Make filtering test harder and fix limit for worst case
update openapi
small clarity refactor
avoid extra allocation when sorting with offset
stream from numeric index btree instead of calculating range
use payload to store order-by value, instead of modifying Record interface
various fixes:
- fix ordering at collection level, when merging shard results
- fix offset at segment level, to take into account also value offset
- make rust tests pass
remove unused histogram changes
fix error messages and make has_range_index exhaustive
remove unused From impl
Move OrderBy and Direction to segment::data_types::order_by
Refactor normal scroll_by in local_shard_operations.rs
More cleanup + rename OrderableRead to StreamWithValue
empty commit
optimization for merging results from shards and segments
fix case of multi-valued fields
fix IntegerIndexParams name after rebase
precompute offset key
use extracted `read_by_id_stream`
Expose value_offset to user
- rename offset -> value_offset
- extract offset value fetching logic
* remove offset functionality when using order_by
* include order_by in ForwardProxyShard
* extra nits
* remove histogram changes
* more nits
* self review
* resolve conflicts after rebase, not enable order-by with datetime index schema
* make grpc start_from value extendable
* gen grpc docs
---------
Co-authored-by: kwkr <kawka.maciej.93@gmail.com>
Co-authored-by: generall <andrey@vasnetsov.com>
* Move CPU count function to common, fix wrong CPU count in visited list
* Change default number of rayon threads to 8
* Use CPU budget and CPU permits for optimizer tasks to limit utilization
* Respect configured thread limits, use new sane defaults in config
* Fix spelling issues
* Fix test compilation error
* Improve breaking if there is no CPU budget
* Block optimizations until CPU budget, fix potentially getting stuck
Our optimization worker now blocks until CPU budget is available to
perform the task.
Fix potential issue where optimization worker could get stuck. This
would happen if no optimization task is started because there's no
available CPU budget. This ensures the worker is woken up again to
retry.
* Utilize n-1 CPUs with optimization tasks
* Better handle situations where CPU budget is drained
* Dynamically scale rayon CPU count based on CPU size
* Fix incorrect default for max_indexing_threads conversion
* Respect max_indexing_threads for collection
* Make max_indexing_threads optional, use none to set no limit
* Update property documentation and comments
* Property max_optimization_threads is per shard, not per collection
* If we reached shard optimization limit, skip further checks
* Add remaining TODOs
* Fix spelling mistake
* Align gRPC comment blocks
* Fix compilation errors since last rebase
* Make tests aware of CPU budget
* Use new CPU budget calculation function everywhere
* Make CPU budget configurable in settings, move static budget to common
* Do not use static CPU budget, instance it and pass it through
* Update CPU budget description
* Move heuristic into defaults
* Fix spelling issues
* Move cpu_budget property to a better place
* Move some things around
* Minor review improvements
* Use range match statement for CPU count heuristics
* Systems with 1 or 2 CPUs do not keep cores unallocated by default
* Fix compilation errors since last rebase
* Update lib/segment/src/types.rs
Co-authored-by: Luis Cossío <luis.cossio@qdrant.com>
* Update lib/storage/src/content_manager/toc/transfer.rs
Co-authored-by: Luis Cossío <luis.cossio@qdrant.com>
* Rename cpu_budget to optimizer_cpu_budget
* Update OpenAPI specification
* Require at least half of the desired CPUs for optimizers
This prevents running optimizations with just one CPU, which could be
very slow.
* Don't use wildcard in CPU heuristic match statements
* Rename cpu_budget setting to optimizer_cpu_budget
* Update CPU budget comments
* Spell acquire correctly
* Change if-else into match
Co-authored-by: Luis Cossío <luis.cossio@qdrant.com>
* Rename max_rayon_threads to num_rayon_threads, add explanation
* Explain limit in update handler
* Remove numbers for automatic selection of indexing threads
* Inline max_workers variable
* Remove CPU budget from ShardTransferConsensus trait, it is in collection
* small allow(dead_code) => cfg(test)
* Remove now obsolete lazy_static
* Fix incorrect CPU calculation in CPU saturation test
* Make waiting for CPU budget async, don't block current thread
* Prevent deadlock on optimizer signal channel
Do not block the optimization worker task anymore to wait for CPU budget
to be available. That prevents our optimizer signal channel from being
drained, blocking incoming updates because the cannot send another
optimizer signal. Now, prevent blocking this task all together and
retrigger the optimizers separately when CPU budget is available again.
* Fix incorrect CPU calculation in optimization cancel test
* Rename CPU budget wait function to notify
* Detach API changes from CPU saturation internals
This allows us to merge into a patch version of Qdrant. We can
reintroduce the API changes in the upcoming minor release to make all of
it fully functional.
---------
Co-authored-by: Luis Cossío <luis.cossio@qdrant.com>
Co-authored-by: Luis Cossío <luis.cossio@outlook.com>
* WIP: `tokio::spawn` update API request handlers [skip ci]
* WIP: `cancel::future::spawn_cancel_on_drop` update API request handlers [skip ci]
* WIP: Make update API cancel-safe [skip ci]
TODO:
- Fix tests
- Evaluate and resolve TODOs
* Fix tests
* Fix benches
* WIP: Simplify cancel safety implementation
* Document and annotate cancel safety guarantees of update API
- Also fix tests after simplifying update API cancel safety impl
- And add a few `cancel::future::cancel_on_token` calls here and there
* Further simplify cancel safety implementation
No more cancellation tokens! 🎉
* Resolve cancel safety TODO
---------
Co-authored-by: timvisee <tim@visee.me>
* This changes the returning of StorageError::BadInput to StorageError::AlreadyExists.
* This adds AlreadyExists error type to StorageError enum
* This implements already_exists function to handle AlreadyExists StorageError.
* This adds error to status code for StorageError::AlreadyExists using tonic::Code::AlreadyExists
* This adds Error type for actix to handle AlreadyExists Storage error using Error::CONFLICT
* This adds using HttpResponse::Conflict for building HttpResp in case of StorageError
* This adds StatusCode and description for HttpError caused by StorageError
* fix integration test
* rename is_collection_exists -> collection_exists
---------
Co-authored-by: Luis Cossío <luis.cossio@outlook.com>
* Merge serde attributes
* Remove obsolete conversion
* Add integer type with parameters
* Make integer lookup and range parameters non-optional
* Add parameterized integer index types test
Co-authored-by: Di Zhao <diz@twitter.com>
* Cleanup
---------
Co-authored-by: Di Zhao <diz@twitter.com>
* add checksum to SnapshotDescription
* implement storing snapshot checksums in a file
* Don't serialize checksum if it's None for backwards compatibility
* Remove hex dependency, use Rust std formatter for this
* Do not error if we cannot remove checksum file for snapshot
Some snapshots may not have a corresponding checksum file. Maybe it was
created in an older Qdrant version that didn't have support for this, or
a user hasn't provided any.
* Add debug message when hashing snapshot, can be expensive on large files
* Inline debug messages
* Add checksum to shard snapshots
* If creating snapshot fails, delete snapshot target and checksum file
* Use Rust idiomatic ok() and improve debug messages
* Use correct snapshot checksum paths, clean up after shard snapshot
* Use better path type in get_checksum_path
---------
Co-authored-by: timvisee <tim@visee.me>
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
* feat: Expose git commit id in the health check endpoint
* fix: CI errors
* test: Add test for health check api
* feat: Add / endpoint to openapi schema
* Make git commit hash optional
* ci: Enable debugging setup-protoc action
* Install later protobuf compiler through GitHub Action
* Disable debug mode for setup-protoc job
* refactor: Use commit instead of commit_id
* fix: Use commit instead of commit_id gRPC docs
* test: Update ping API test
* refactor: Rename ping api to root api
---------
Co-authored-by: timvisee <tim@visee.me>
* Set high/low priority for consensus and HNSW threads
* Make setting thread priority Linux specific
* Fix compilation on non-Linux platforms
* Remove unused dependencies from Cargo.toml
* Rename function
* Add warning that setting lower nice is likely to fail
* Synchronize nodes after creating collection in distributed mode
* Give slow CI machines more time for collection churning
* Improve comment
* Do not double-wait synchronizing nodes when creating shard key
* Only sync nodes when creating collection, shardkey or changing aliases
* Update comments
* Read-only API keys
Co-authored-by: Luis Cossío <luis.cossio@outlook.com>
Correct placement of OpenAPI security
Place regex dep with actix/tonic
* Read-only API keys
* Replace with pytests
* API Key tests run on the same job
* Drop allow dead-code
* Rename setting key
* Containerized tests
* No special config files
* DRY
* refactor: re-use can_write method
* refactor: replace static by constants
* refactor: get PID from `$!`
* refactor: use explicit brackets on boolean condition
* style: fix identation
* small fixes + account for new APIs
* specify security in openapi
* small fix + chmod for .sh testfile
* add best-efford check for api consistency
---------
Co-authored-by: Amr Hassan <amr.hassan@gmail.com>
Co-authored-by: generall <andrey@vasnetsov.com>
* Extract shard transfer implementations into modules
* Extract shard transfer helpers into module
* Extract functions driving shard transfer into module
* Move implementation below struct definition
* remove duplicated search methods, introduced for compatibility in last version
* Use `with_capacity` rather than a manual reserve
* explicit Arc clones
* get rid of batching by runs of same strategy
* avoid refactor in group by
* more explicit arc clones, remove one .expect()
* little extra refactor on recommendations.rs
* refactor grouping_test.rs too
---------
Co-authored-by: timvisee <tim@visee.me>
Co-authored-by: Luis Cossío <luis.cossio@outlook.com>
* WIP: collection-level storage for payload indexe scheme
* introduce consensus-level operation for creating payload index
* make operation_id optional in the UpdateResult
* set payload index in newly created shards
* upd api definitions
* include payload index schema into collection consensus state
* include payload index schema into shard snapshot
* review fixes
* Ensure shard snapshot methods and API are cancel safe
* Remove resolved TODOs
* Extend comment on why we don't unproxify queue in a normal way
* fixup! Remove resolved TODOs
* Restructure `spawn_transfer_task`
* Update lib/collection/src/shards/transfer/shard_transfer.rs
Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>
* fixup! Restructure `spawn_transfer_task`
Make `spawn_transfer_task` more readable
---------
Co-authored-by: timvisee <tim@visee.me>
Co-authored-by: Tim Visée <tim+github@visee.me>
* WIP: Add proper HTTPS support for remote snapshot downloads to snapshot recover APIs
* Implement HTTPS client configuration
* fixup! Implement HTTPS client configuration
Move HTTPS client configuration from `actix/certificate_helpers.rs` to
`common/http_client.rs`
* Initialize and propagate HTTPS client to shard snapshot API
* Use reqwest client reference where possible
* Add integration test
* Simplify HTTP(S) client initialization
* fixup! Simplify HTTP(S) client initialization
* Fix lifetime conflicts after rebase on dev
* Add comment to elaborate on PEM concatenation
* fixup! Add integration test
Add `test_tls_snapshot_shard_transfer.sh` run to the existing TLS test job
instead of creating a new one
---------
Co-authored-by: timvisee <tim@visee.me>
* create and connect discovery http and grpc interfaces
* add openapi tests
* fix bad rebase
* Add better descriptions
* remove numpy from openapi tests
* fix rebase artifact
* remove already addressed TODO
* add more tests
* 🤡🔫 (cfg batch handler)
* add timeout query param for discover requests
* More gRPC validation
* make fields pydantic_openapi_generator_v3 friendly
* `context_pairs` -> `context` with struct for pairs
* discovery api is only discovery or context,
move struct description to fields
---------
Co-authored-by: timvisee <tim@visee.me>