* update storage compat test to include quantized collections
* remove unused collection snapshot
* increase dims and range of vectors
* capture pid right after starting qdrant
* Remove extra space
* storage-compatibility.sh: capture PID right after instantiation
---------
Co-authored-by: timvisee <tim@visee.me>
* Refactor Dockerfile
- fix cross-compilation
- improve caching
TODO:
- check if `lld` is used/works for linkage (and enable if not used, or remove if doesn't work)
* Remove `aarch64` linker config from `.cargo/config.toml` (seems to be unnecessary)
* Expose `LINKER` argument and enable `lld` linker
* Add `mold` linker support
* Document Dockerfile
* Add closing ` in the Dockerfile comments
Co-authored-by: Tim Visée <tim+github@visee.me>
---------
Co-authored-by: Tim Visée <tim+github@visee.me>
* Add caching of docker layers in CI
Build required docker images for CI in a workflow step using buildkit's
gha cache type. This will populate the local layer cache from github
actions' cache. Builds in subsequent CI steps will be nearly instant,
because all layers can be reused.
* add minor change to see if build time is any faster
---------
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
* WIP: Start working on out-of-RAM errors handling [skip ci]
* Implement basic handling of out-of-RAM errors during Qdrant startup
* Try to fix CI fail by allowing both V1 and V2 cgroups
* Try to fix CI fail by improving cgroups handling
* Fix cgroups path detection/handling (+ some minor stylistic changes)
* fixup! Fix cgroups path detection/handling (+ some minor stylistic changes)
* Add test
* Enable low RAM test
* fixup! Add test
* free memory checks
* rm unused function
* Oom fallback script (#1809)
* add recover mode in qdrant + script for handelling OOM
* fix clippy
* reformat entrypoint.sh
* fix test
* add logging to test
* fix test
* fix test
---------
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
* Exclude deleted vectors from HNSW graph building stage
* When estimating query cardinality, use available points as baseline
We should not use the total number of points in a segment, because a
portion of it may be soft deleted. Instead, we use the available
(non-deleted) points as baseline.
* Add plain search check to unfiltered HNSW search due to deleted points
* Cardinality sampling on available points, ignore deleted named vectors
* Estimate available vectors in query planner, now consider deleted points
In the query planner, we want to know the number of available points as
accurately as possible. This isn't possible because we only know the
number of deletions and vectors can be deleted in two places: as point
or as vector. These deletions may overlap. This now estimates the number
of deleted vectors based on the segment state. It assumes that point and
vector deletions have an overlap of 20%. This is an arbitrary
percentage, but reflects an almost-worst scenario.
This improves because the number of deleted points wasn't considered at
all before.
* Remove unused function from trait
* Fix bench compilation error
* Fix typo in docs
* Base whether to do plain search in HNSW upon full scan threshold
* Remove index threshold from HNSW config, only use full scan threshold
* Simplify timer aggregator assignment in HNSW search
* Remove vector storage type from cardinality function parameters
* Propagate point deletes to all its vectors
* Check for deleted vectors first, this makes early return possible
Since point deletes are now propagated to vectors, deleted points are
included in vector deletions. Because of that we can check if the vector
is deleted first so we can return early and skip the point deletion
check.
For integrity we also check if the point is deleted, if the vector was
not. That is because it may happen that point deletions are not properly
propagated to vectors.
* Don't use arbitrary vector count estimation, use vector count directly
Before we had to estimate the number of vectors (for a named vector)
because vectors could be deleted as point or vector. Point deletes are
now propagated to vector deletes, that means we can simply use the
deleted vector count which is now much more accurate.
* When sampling IDs, check deleted vecs before deleted points
* On segment consistency check, delete vectors for deleted points
* Fix vector delete state not being kept when updating storage from other
* Fix segment builder skipping deleted vectors breaking offsets
* update segment to handle optional vectors + add test (#1781)
* update segment to handle optional vectors + add test
* Only update stored record when deleting if it wasn't deleted already
* Reformat comment
---------
Co-authored-by: timvisee <tim@visee.me>
* Fix missed vector name test, these are now marked as deleted
* upd test
* upd test
* Update consensus test
---------
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
* Change file mode of some .sh files to make them executable
* Remove whitespace on empty lines from YAML configurations
* Fix unused import warning on Windows
* Add user client certificate validation capabiltiy
* Add integration test for client side certificate validation
* Add server side ca certificate to verify
* Use store builder for client side certificate
* Use a trust store
* Fix config for TLS test
* Fix test, remove mTLS for external grpc endpoint
* Fix config comment
* Remove useless commit
* Fix style
* Fix internodal TLS test
* Simplify setting SSL verify mode
---------
Co-authored-by: timvisee <tim@visee.me>
* Add test script to ensure OpenAPI files are consistent with sources
* Add CI job to test OpenAPI file consistency
* Add CI task to test gRPC file consistency
* Tweak consistency scripts a bit, touch temp file to trigger gRPC rebuild
* Don't test .gitignored files
* Mention updating the OpenAPI specification is enforced by CI
* Update CI job configuration
* Also check consistency of gRPC docs
* Rename temporary files to have a .diff prefix
* Add docs to consistency checking scripts
* Fix incorrect metrics value for cluster commit
* Rewrite metrics logic, don't use registry, write values directly
* Only report REST timings for requests having HTTP 200 response
* Limit metrics reporting of endpoints to whitelist
The whitelist contains a selection of search, recommend and upsert endpoints.
* Add MetricsParam, remove detail level, keep anonymize
* Request metrics in basic API test
* Specify content type for metrics endpoint
* Add OpenAPI test for metrics endpoint, remove from basic API test
This test probes for some strings that must exist in the output
* Add note that metrics endpoint whitelist must be sorted
- Fix remote shards state recovery during Raft snapshot application
- Fix local shards data recovery during Raft snapshot application
- Refactor `test_collection_recovery` test
* Forward write request according to write consistency
* use gRPC internal update API by making shard_id optional
* improve and test
* mark leader peer as failed in case of service error during forward
* forward write ordering param
* Recover dead shards pro-actively
* reuse same ports as changing it requires a full consensus round
* refactoring
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
* WIP: Fix `Segment::take_snapshot`
TODO:
- This commit, probably, breaks snapshotting of segments with memmapped vector storage
- `ProxySegment::take_snapshot` seems to potentially similar bug
* WIP: Fix `Segment::take_snapshot`
- Fix snapshotting of `StructPayloadIndex`
- Fix snapshotting of segments with memmapped vector storage
- Temporarily break `ProxySegment::take_snapshot`
* Fix `ProxySegment::take_snapshot`
* Remove `copy_segment_directory` test
* nitpicking
* clippy fixes
* use OperationError::service_error
* Cleanup `TinyMap` trait bounds and derive `Debug`
* Fix `test_snapshot` test
- Derive `Debug` for `NamedVectors`
* Move utility functions from `segment.rs` to `utils` module
* Contextualize `segment::utils::fs::move_all` a bit more carefully
* Fix a typo
* add backward compatibility with old snapshot formats
* fmt
* add snapshot for compatibility test
* git lfs is a piece of shit
* Nitpicking
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
* WIP: introduce local state
* WIP: sync local state with consensus in idle
* fmt
* rm extensive debug
* update triple replication test
* update triple replication test
* rm debug logs
* fix check for established leader + only sync local if no proposals
* remove unused file
* test compatible with python 3.8
* test compatible with python 3.8
* longer wait for consensus
* longer wait for consensus
* extra sleep in test
* remove extra sleep
* Fixing missed leader inconsistency - transfers (#1298)
* explicit request timeout
* explicit request timeout
* explicit request timeout
* explicit request timeout
* explicit request timeout
* prevent double handelling of the transfer termination
* kill the process
* log on inconsistency
* log on inconsistency
* log on inconsistency
* debug
* revert debug in test
* forward updates to partial shards, abort transactions on dead node report
* fmt
* disable retry of transfer, if the transfer was cancelled + allow predictable ports in test
* fix import
* enable compression for gRPC
* add benchmarks
* fix grpc benchmark
* discardResponseBodies to reduce memory usage
* improve & split benchmarks
* do not use gzip for inter node communication as benchmarks are not conclusive