* WIP: collection-level storage for payload indexe scheme
* introduce consensus-level operation for creating payload index
* make operation_id optional in the UpdateResult
* set payload index in newly created shards
* upd api definitions
* include payload index schema into collection consensus state
* include payload index schema into shard snapshot
* review fixes
* create and connect discovery http and grpc interfaces
* add openapi tests
* fix bad rebase
* Add better descriptions
* remove numpy from openapi tests
* fix rebase artifact
* remove already addressed TODO
* add more tests
* 🤡🔫 (cfg batch handler)
* add timeout query param for discover requests
* More gRPC validation
* make fields pydantic_openapi_generator_v3 friendly
* `context_pairs` -> `context` with struct for pairs
* discovery api is only discovery or context,
move struct description to fields
---------
Co-authored-by: timvisee <tim@visee.me>
* SparseVector implements VectorIndex
* dedicated telemetry
* conflict
* easy code reviews
* simplify tracking indexed points count
* move telemetry conversion to sparse file
* move max_result_count out of inverted index trait
* unify sparse vector fixtures
* simpler conversion
* add todo regarding OOM potential
* reuse check deleted from raw scorer with TODO
* change new to open to handle mmap index
* add timeout query param for search requests
* enable timeout for recommend requests
* Add query timeout for group by requests
* update openapi models
* Don't decrease timeout after recommend preprocessing
* Add openapi test
* code review
* add timeout to individual group by requests, non-decreasing
* handle timeout for discover
* Update timeout field tag in SearchBatchPoints
message
* wip: implement consensus operations to create and delete shard keys
* Fix typo in variable name
* wip: methods to add shard keys
* fmt
* api for submitting creating and deleting shard keys operations into consensus
* fmt
* apply shard mapping in with the collection state transfer
* handle single-to-cluster transition
* fix api structure and bugs
* fmt
* rename param
* review fixes
---------
Co-authored-by: timvisee <tim@visee.me>
* Clone inside blocks
* Add shard transfer method to distinguish between batching and snapshots
* Add stub method to drive snapshot transfer
* Store remote shard in forward proxy, merge unproxy methods
* On snapshot shard transfer, create a shard snapshot
* Unify logic for unproxifying forward and queue proxy
* Error snapshot transfer if shard is not a queue proxy
* Add remote shard function to request remote HTTP port
* Handle all specific shard types when proxifying
* Allow queue proxy for some shard holder snapshot methods
* Bring local and remote shard snapshot transfer URLs into transfer logic
* Expose optional shard transfer method parameter in REST and gRPC API
* Expose shard transfer method in list of active transfers
* Fix off-by-one error in queue proxy shard batch transfer logic
* Do not set max ack version for WAL twice, already set when finalizing
* Merge comment for two similar calls
* Use reqwest client to transfer and recover shard snapshot on remote
Using the reqwest client should be temporary. We better switch to a gRPC
call here eventually to use our existing channels. That way we don't
require an extra HTTP client (and dependency) just for this.
* Send queue proxy updates to remote when shard is transferred
* On shard queue transfer, set max WAL ack to last transferred
* Add safe queue proxy destructor, skip destructing in error
This adds a finalize method to safely destruct a queue proxy shard. It
ensures that all remaining updates are transferred to the remote, and
that the max acknowledged version for our WAL is released. Only then is
the queue proxy shard destructed unwrapping the inner local shard.
Our unproxify logic now ensures that the queue proxy shard remains if
transferring the updates fails.
* Clean up method driving shard snapshot transfer a bit
* Change default shard transfer method to stream records
This changes the default transfer method to stream records rather than
using a snaphsot transfer. We can switch this once snapshot transfer is
fully integrated.
* Improve error handling, don't panic but return proper error
* Do not unwrap in type conversions
* Update OpenAPI and gRPC specification
* Resolve and remove some TODOs
* During shard snapshot transfer, use REST port from config
* Always release max acknowledged WAL version on queue proxy finalize
* Rework queue unproxying, transform into forward proxy to handle errors
When a queue or forward proxy shard needs to be unproxified into a local
shard again we typically don't have room to handle errors. A queue proxy
shard may error if it fails to send updates to the remote shard, while a
forward proxy does not fail at all when transforming.
We now transfer queued updates before a shard is unproxified. This
allows for proper error handling. After everything is transferred the
shard is transformed into a forward proxy which can eventually be safely
unproxified later.
* Add trace logging for transferring queue proxy updates in batch
* Simplify snapshot method conversion from gRPC
* Remove remote shard parameter
* Add safe guard to queue proxy handler, panic in debug if not finalized
* Improve safety and architecture of queue proxy shard
Switch from an explicit finalized flag to an outer-inner architecture.
This improves the interface and robustness of the type.
* Do not panic on drop if already unwinding
* Make REST port interface in channel service for local node explicitly
* Recover shard on remote over gRPC, remove reqwest client
* Use shard transfer priority for shard snapshot recovery
* Remove obsolete comment
* Simplify qualified path with use
* Don't construct URLs ourselves as a string, use `parse` and `set_port`
* Use `set_path` when building shard download URL
* Fix error handling in queue to forward proxy transformation
Before, we didn't handle finalization errors properly. If this failed,
tie shard would be lost. With this change the queue proxy shard is put
back.
* Set default shard transfer method to stream records, eliminate panics
* Fix shard snapshot transfer not correctly aborting due to queue proxy
When a shard transfer fails (for any reason), the transfer is aborted.
If we still have a queue proxy shard it should also be reverted, and
collected updates should be forgotten. Before this change it would try
to send all collected updates to the remote, even if the transfer
failed.
* Review fixes
Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>
* Review fixes
Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>
* Initiate forward and queue proxy shard in specialized transfer methods
Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>
* Add consensus interface to shard transfer, repurpose dispatcher (#2873)
* Add shard transfer consensus interface
* Integrate shard transfer consensus interface into toc and transfer logic
* Repurpose dispatcher for getting consensus into shard transfer
* Derive clone
* Mark consensus as unused for now
* Use custom dispatcher with weak ref to prevent Arc cycle for ToC
* Add comment on why a weak reference is used
* Do exhaustive match in shard unproxy logic
* Restructure match statement, use match if
* When queue proxifying shard, allow forward proxy state if same remote
* Before retrying a shard transfer after error, destruct queue proxy
* Synchronize consensus across all nodes for shard snapshot transfer (#2874)
* Move await consensus commit functions into channel service
* Add shard consensus method to synchronize consensus across all nodes
* Move transfer config, channels and local address into snapshot transfer
* Await other nodes to reach consensus before finalizing shard transfer
* Do not fail right away awaiting consensus if still on older term
Instead, give the node time to reach the same term.
* Fix `await_commit_on_all_peers` not catching peer errors properly
* Change return type of `wait_for_consensus_commit` to `Result`
This is of course more conventional, and automatically sets `must_use`.
* Explicitly note number of peers when awaiting consensus
* Before consensus sync, wait for local shard to reach partial state
* Fix timeout error handling when waiting for replica set state
* Wait for replica set to have remote in partial state instead
* Set `(Partial)Snapshot` states for shard snapshot transfer through consensus (#2881)
* When doing a shard snapshot transfer, set shard to `PartialSnapshot`
* Add shard transfer method to set shard state to partial
It currently uses a naive implementation. Using a custom consensus
operation to also confirm a transfer is still active will be implemented
later.
* Add consensus snapshot transfer operation to change shard to partial
The operation `ShardTransferOperations::SnapshotRecovered` is called
after the shard snapshot is recovered on the remote and it progresses
the transfer further.
The operation sets the shard state from `PartialSnapshot` to `Partial`
and ensures the transfer is still active.
* Confirm consensus put shard into partial state, retry 3 times
* Get replica set once
* Add extensive shard snapshot transfer process docs, clean up function
* Fix typo
* Review suggestion
Co-authored-by: Luis Cossío <luis.cossio@qdrant.com>
* Add delay between consensus confirmation retries
* Rename retry timeout to retry delay
---------
Co-authored-by: Luis Cossío <luis.cossio@qdrant.com>
* On replicate shard, remember specified method
---------
Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>
Co-authored-by: Luis Cossío <luis.cossio@qdrant.com>
* Fix paste-bugs in `snapshot_service.proto`
* Add shard snapshot gRCP API definition
* Add validation to shard snapshot gRPC API definition
* Implement conversions between gRPC and `collection` types
* Extract shard snapshot API implementation into common sub-module
* Implement shard snapshot gRPC API
* Generate gRPC docs
* Refactor `ShardSnapshots` gRPC service to be internal API only
* fixup! Refactor `ShardSnapshots` gRPC service to be internal API only
Move `ShardSnapshotRecoverResponse` to `shard_snapshots_service.proto`
* fixup! fixup! Refactor `ShardSnapshots` gRPC service to be internal API only
Update `api/src/grpc/qdrant.rs`
* fixup! fixup! Refactor `ShardSnapshots` gRPC service to be internal API only
Update gRPC docs
* Switch `ShardSnapshots` gRPC service to use `validate_and_log` instead of `validate`
* Extend GeoPolygon to support interiors (#2315)
Per GeoJson, we should support polygon with exterior and interiors (holes on the surface) in Geo Filter by Polygon(#795). This commit extend current GeoPolygon filter to accept interiors. It includes:
1. changes to proto and internal GeoPolygon struct, and validation fn
2. add and refactor some tests
3. add integration test
* add gRPC geo_polygon validation
---------
Co-authored-by: Arnaud Gourlay <arnaud.gourlay@gmail.com>
* allow raw vectors in recommend
* add openapi test
* keep point id input separately forever in grpc
* minor fix
---------
Co-authored-by: generall <andrey@vasnetsov.com>
* Add naive wait-on-consensus-commit gRPC call with spinlock
* Add method for waiting on all other nodes to reach a consensus commit
* Fix codespel issue
* Consensus commit spin lock should use consensus tick interval
* Move consensus commit spin lock function into consensus manager module
* Demote function to wait on single peer
* Add explicit timeout to all consensus await commit methods
* Update gRPC docs
* Peer must be on the same term when waiting on a commit
* Update lib/storage/src/content_manager/toc.rs
* Update generated gRPC code
* add enum for vector query on segment search
* rename newly introduced types
* fix: handle QueryVector on async scorer
* handle QueryVector in QuantizedVectors impl
* fix async scorer test after refactor
* rebase + refactor on queue_proxy_shard.rs
* constrain refactor propagation to segment_searcher
* fmt
* fix after rebase
* Add healthz, livez and readyz endpoints
* Describe healthz, livez and readyz endpoints in OpenAPI definition
* Add integration test for healthz, livez and readyz endpoints
* Fix typo
* Add name to optimizers
* Track optimizer status in update handler
* Remove unused optimizer telemetry implementation
* Report tracked optimizer status in local shard telemetry
* Keep just the last 16 optimizer trackers and non successful ones
* Also eventually truncate cancelled optimizer statuses
* Fix codespell
* Assert basic optimizer log state in unit test
* Remove repetitive suffix from optimizer names
* Loosen requirements for optimizer status test to prevent flakyness
* Add gRPC endpoint to request remote HTTP port
* Add basic gRPC test for fetching remote HTTP port
* Extract gRPC endpoint for fetching HTTP port into internal service
* Add basic internal gRPC test
* Other minor improvements
* add ignore_plain_index to speed up search
* remove unnecessary & for vectors_batch
* format
* add special handle for proxy segments where the wrapped segment is plain
indexed
* review refactoring
* rollback changes in google.protobuf.rs
---------
Co-authored-by: Di Zhao <diz@twitter.com>
Co-authored-by: generall <andrey@vasnetsov.com>
* Add optional `tracing` crate dependency to all `lib/*` sub-crates
* Add optional alternative "tracing-logger"...
...that works a bit better with logs produced by the `tracing` crate
* Make `tracing-logger` reuse `log_level` config parameter (and `QDRANT__LOG_LEVEL` env-var)...
...instead of default `RUST_LOG` env-var only
* Replace `env_logger` with `tracing` (#2381)
* Replace `env_logger` with `tracing`/`tracing-subscriber`/`tracing-log`
TODO:
- Implement slog drain that would *properly* convert slog records into tracing events
(`slog-stdlog` is fine for now, though)
* fixup! Replace `env_logger` with `tracing`/`tracing-subscriber`/`tracing-log`
Fix `consensus` test
* fixup! fixup! Replace `env_logger` with `tracing`/`tracing-subscriber`/`tracing-log`
Forget to `unwrap` the result 🙈
* fixup! fixup! Replace `env_logger` with `tracing`/`tracing-subscriber`/`tracing-log`
`tracing_subscriber::fmt::init` initializes `tracing_log::LogTracer`
automatically with default `tracing-subscriber` features enabled 💁♀️
* Update `DEVELOPMENT.md` documentation regarding `tracing`
* Update default log level for development to make it less noisy
---------
Co-authored-by: timvisee <tim@visee.me>
* Implement "intelligent" default log-level filtering
---------
Co-authored-by: timvisee <tim@visee.me>
* Reoptimize segments on quantization mismatch
* Add collection quantization params to collection update REST endpoint
* Do not require rebuild of quantization for on_disk change
* Add collection quantization params to collection update gRPC endpoint
* Add option to update vector specific quantization config to REST API
* Add option to update vector specific quantization config to gRPC API
* Update OpenAPI and gRPC specification
* Fix quantization mismatch detecting not working as expected
* Fix config mismatch optimizer not using vector specific quantization
* Add unit test for quantization in config mismatch optimizer
* Add some quantization params to collection update integration test
* Reformat
* Apply suggestions from code review
Co-authored-by: Arnaud Gourlay <arnaud.gourlay@gmail.com>
* Remove if-statement, simply return
* Move quantization params in named vector in integration test
* Test updating quantization params in multivec integration test
* Fix gRPC docs
* Fix quantization rebuild on changing enabled state on indexed segment
* allow disabling quantization config
* fmt
* clippy
* compare all params of the quantization for rebuild
* explicitly use params on update
* test for disabling quantization
---------
Co-authored-by: Arnaud Gourlay <arnaud.gourlay@gmail.com>
Co-authored-by: generall <andrey@vasnetsov.com>
* unit-test for updating multivec config
* Remove single vector support from vector params update in REST API
With this change, a user is always required to specify vector names when
updating their parameters. Updating vector parameters in a collection
with a single vector is still possible by providing an empty name. This
is very reasonable as updating vector parameters on a single vector
collection is a bad pattern, since parameters can be set on the
collection itself.
* Update integration tests to reflect REST API change
* Update OpenAPI specification
* Remove obsolete VectorsConfigDiff functions, update docs
* Update OpenAPI specification
* Add name validation to update collection endpoint for vector params
---------
Co-authored-by: timvisee <tim@visee.me>
* Add update vector config type
* Add mutable getter to update vector config
* Add update vector config to update collection REST operation, reorder
* Add logic to update vector HNSW in collection
* Update OpenAPI specification
* If HNSW diff is empty, consider it as unset
* Update descriptions to clarify an empty HNSW diff unsets
* Make new vector HNSW diff update existing diff
This means that the existing vector HNSW diff is kept intact. Specified
fields are updated in the existing diff. An empty HNSW diff object may
be provided to unset the diff.
* Add update vector config to update collection gRPC operation
* Update gRPC docs
* Use vector specific HNSW config in integration test
* Extract & improve gRPC type conversions, fix collection update with None
* Validate vector specific HNSW config in collection update gRPC endpoint
* Explicitly test we do not rebuild on on_disk change
* Use new method to recreate optimizers
* Don't test on_disk change anymore
* Reuse collection write lock to save, do not relock
* Fix invalid update collection request body in OpenAPI test
* Simplify update collection test
* Add update collection test for HNSW parameters
* Extract vector param update logic into local functions
* Reformat
* Transform UpdateVectorParams into VectorParamsDiff
* Transform UpdateVectorsConfig into VectorsConfigDiff
* Combine serde attributes
* Make HNSW diff empty check generic over DiffConfig types
* Move DiffConfig implementation
* Fix typo
* charabia tokenizer for CJK, init
* add charabia to .proto
* generate grpc docs
* rename charabia to multilingual
* disable charabia default features
* ignore stopwords and separators in tokenizer
* fix rebase conflict
* fix test
* fix test
* Bump `charabia` crate version...
...and enable `charabia` features that does not require any additional dependencies by default
* Expose optional `charabia` features in the `segment` crate
* Expose optional `charabia` features in the `qdrant` crate
* fix codespell
* more regression tests
* more regression tests
* make language-specific tests feature-flagged
---------
Co-authored-by: Anatolii Smolianinov <zarkonesmall@gmail.com>
Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>
* Add github action to codespell master on push and PRs
* Add rudimentary codespell config
* some skips
* fix some ambigous typos
* [DATALAD RUNCMD] run codespell throughout
=== Do not change lines below ===
{
"chain": [],
"cmd": "codespell -w",
"exit": 0,
"extra_inputs": [],
"inputs": [],
"outputs": [],
"pwd": "."
}
^^^ Do not change lines above ^^^
* Add dev branch as target for the workflow