* nitpick: use CardinalityEstimation::exact() instead of explicit struct
* wip: implement binary index
* persist binary index
* add tests, refactor binary memory
* use bitflags for BinaryItem, use early return from @timvisee 's review
* fix iterator's end +1
* fix indexed_count, keep track of falses and trues counts
* Fix true/false counts if flag for point was set already
* Remove match in filter structure
* - rename `*_count()`
- fix set_or_insert to actually set
- add more tests
* review suggestions for binary index (#2186)
* small taste refactors
* move memory into its own module
---------
Co-authored-by: timvisee <tim@visee.me>
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
This is one of a series of commits for the new feature Geo Filter by Polygon(#795) that add function polygon_hashes that retrieves as-high-as-possible with maximum of max_regions number of geo-hash guaranteed to contain the whole polygon.
It also includes two helper functions: 1) check_polygon_intersection and 2) minimum_bounding_rectangle_for_polygon.
Test cases validate the functionality for different scenarios.
* Use unwrap_or_default rather than constructing it manually
* Validate oversampling parameter and surrounding structs
* Validate search limit to be 1 or higher
* Validate hnsw_ef parameter to be 4 or higher
* Also validate parameters in in gRPC
* Update OpenAPI specification
* Add unlikely to reach comment
* Remove hnsw_ef validation rules for now
* Update OpenAPI specification
This commit introduces the GeoPolygon struct, representing a polygon defined by a list of coordinates. It includes the check_point method to determine if a given point is inside the polygon. Test cases validate the functionality for different scenarios.
* Issue 1905: Configurable location for the tmp snapshot files
* Apply suggestions from code review
Co-authored-by: Tim Visée <tim+github@visee.me>
* fix code review suggestions
* clippy fix
* Propagate temp path, use configured dir for snapshot creation
* Use real temp dir in snapshot tests
* Mention default temporary snapshot file path in configuration
* Use temp everywhere rather than a mix of temp and tmp
* Use consistent naming for temporary snapshot directories
* Extract logic for temporary storage path into toc method
* Resolve clippy warnings
* Apply suggestions from code review
Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>
---------
Co-authored-by: Tim Visée <tim+github@visee.me>
Co-authored-by: timvisee <tim@visee.me>
Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>
* Before inserting vector update, check for correct dimensions
* Improve vector name and data checking on segment level
* Optimize vector compatibility checking
* Remove obsolete vector dimensionality check
* Add test to ensure all segment functions catch bad input
* Test vector checking functions directly as well
* wip: add rest param
* WIP: Refactor `search_with_graph` to use oversampling
* fixup! WIP: Refactor `search_with_graph` to use oversampling
Fix `search_with_graph` quantization config handling
* fixup! WIP: Refactor `search_with_graph` to use oversampling
Use `truncate` instead of `shrink_to` 🤦♀️
* schrink -> truncate
* protection against wrong oversampling value
* Revert second "fixup! WIP: Refactor `search_with_graph` to use oversampling"
Accidentally commited the wrong file 🤦♀️
* fix max method notation
* rollbacK: max notation
* Add `oversampling` field to `QuantizationSearchParams` gRPC type
* Simplify `oversampled_top` calculation
* `./tools/generate_grpc_docs.sh`
* `./tools/generate_openapi_models.sh`
---------
Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>
* Correctly count vectors in segment info for normal segment
* Correctly count vectors in segment info for proxy segment
* Simplify available point count method
* Minor improvements
* Add unit test for point and vector counts in segment
* Add unit test for point and vector counts in proxy segment
* Improve vector counting for proxy segment
* async raw scorer
* fmt
* Disable `async_raw_scorer` on non-Linux platforms
* Refactor `async_raw_scorer.rs`
* Conditionally enable `async_raw_scorer` in `segment` crate
* Add `async_scorer` config parameter to the config...
...and enable `async_raw_scorer`, if config parameter is set to `true`
* fixup! Add `async_scorer` config parameter to the config...
Fix tests
* Add basic `async_raw_scorer` test
* Extend `async_raw_scorer` tests to be more extensive
* Async uring vector storage io uring (#2041)
* replace tokio-uring with just low level io-uring
* fnt
* minor fixes
* add sync
* wip: try to use less submissions
* fmt
* fix uring size
* larger buffer
* check for overflow
* submit with re-try
* mmap owns uring context
* large disk parallelism
* rollbacK: large disk parallelism
* fix windows build
* explicitly panic on uring fail
* fix windows build again
* use async scorer in the quantization re-scoring
* refactor
* rename UringReader
* refactor buffers
* fix for windows
* error checking
* fix handing
---------
Co-authored-by: Roman Titov <ffuugoo@users.noreply.github.com>
* Skip applied WAL operations on shard load
* fixup! Skip applied WAL operations on shard load
- Skip WAL operation up to the *minimal* persisted version
- Add explanation comment
* fixup! Skip applied WAL operations on shard load
- "beautify" `skip_while` expression
- handle version `0` corner case (and add explanation)
* Update `Segment::handle_version` comment
* Simplify `Segment::handle_version` a bit
* Refactor `check_unprocessed_points`
* fixup! Refactor `check_unprocessed_points`
Optimized `check_unprocessed_points`
* Revert `LocalShard::load_from_wal` and `check_unprocessed_points` changes
* Downgrade failure in `check_unprocessed_points` to a warning
* Persist/track the first un-truncated index in `SerdeWal`...
...and use it as first index for WAL operations
* fixup! Downgrade failure in `check_unprocessed_points` to a warning
Downgrade failure in `check_unprocessed_points` to a warning *during `LocalShard::load_from_wal`* 🙈
* fixup! Persist/track the first un-truncated index in `SerdeWal`...
Words...
* Fix a bug in `SerdeWal::truncated_prefix_entries_num` and add explanation comment
* Fix a bug in `SerdeWal::ack`
* prevent unnesessary savings, use json, never decrease ack index
---------
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
* feat: add lookup_ids function
- rename GroupId -> PseudoId
- impl try_from PseudoId to PointIdType
- create lookup_ids function
- add tests
* fix: keep `GroupId` name in the openapi schema
* refactor: make new type for PseudoId
* refactor: make new `Lookup` output for lookup_ids()
* changes from review, thanks @agourlay