* fix(local): skip bool group-by keys to match server GroupId semantics
* fix: do not hard code types, move tests to test group search
---------
Co-authored-by: feiiiiii5 <feiiiiii5@users.noreply.github.com>
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* fix(local): honor nested json-path keys in delete_payload
In local mode delete_payload only removed top-level dict keys, so a key
given as a json path (`a.b`, `location[0].name`, `location[].name`) never
matched and the delete was a silent no-op. The server deletes nested keys
via dot notation and preserves the rest of the payload, so local mode
diverged from it.
set_payload and filters already resolve these paths through
parse_json_path; delete_payload was the one payload operation ignoring
them. Add a delete_value_by_key helper next to set_value_by_key that walks
the same JsonPathItem path and removes the leaf (a missing path is a
no-op, siblings are preserved), and use it from delete_payload.
* fix(local): match server semantics for indexed payload deletion
delete_value_by_key deleted terminal array elements by index and
honored Python-style negative indices, but the server does neither:
it treats a terminal array-index delete as a no-op (not idempotent)
and addresses elements with an unsigned index, so a negative index
cannot be represented. Both cases diverged from the server this path
exists to mirror.
Make a terminal array index a no-op and require a non-negative,
in-range index for nested traversal. Add local and congruence
coverage for terminal and negative indices.
* fix: update json path parser, do not apply partial updates in delete by key, add tests
* fix: remove new redundant top level directory
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* fix: slice index error message reports an off-by-one upper bound
The guard accepts `0 <= index < total`, so the largest valid index is
`total - 1`, but the message advertises the range as `0..{total}`. With
`total=4`, rejecting `index=4` reports "Slice index must be in 0..4, got 4",
naming the rejected value as though it were allowed.
Report `0..{total - 1}` instead.
* fix slice error message in local mode
---------
Co-authored-by: George <george.panchuk@qdrant.tech>
* fix(local): match core's MMR tie-breaking
Local mode ordered MMR results differently from core whenever two
candidates tied exactly, on relevance or on MMR score. The MMR score
itself was already correct; the selection rules around it were not:
* the first point was seeded from `candidate_ids[0]`, i.e. whatever
`search` happened to return first, instead of the most relevant
candidate;
* `np.argmax` picked the *first* maximum, while core's `max_by_key`
returns the *last* one on ties;
* pending candidates were kept in an order-preserving list, while core
holds them in an `IndexSet` and drops the selected one with
`swap_remove`, which moves the last candidate into the freed slot and
therefore changes the order candidates are visited in - and so which
one wins a tie.
Reproduce the three rules in `_mmr`. The divergence was spotted on a
MAX_SIM multivector field with DOT, but it is specific to neither:
plain dense vectors and EUCLID diverge the same way once an exact tie
is constructed.
The added congruence tests keep relevance scores distinct on purpose:
core orders equally relevant candidates by search order, which is not
stable, so only ties in the MMR score can be asserted on.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* refactor: move utils from collection, move tests
* tests: update comments
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix: fix missing multivector placeholder in local mode
* fix: do not use deleted vectors in recommend, etc
* fix: fix mypy complaints in point id vector resolution
* fix: regen async
Recommend, discovery, context and relevance-feedback queries score points
from the internal core distance, which is oriented so that a higher score
is always better regardless of the collection's distance metric. The raw
distance order therefore does not describe how their scores should be
sorted or thresholded.
Two places got this wrong on Euclid/Manhattan collections in local mode:
- Result ordering already special-cased the recommend/discovery/context
family, but omitted NaiveFeedbackQuery, so relevance-feedback queries
returned the farthest points first instead of the nearest.
- score_threshold filtering keyed off the raw distance order for every
query type, so for the whole higher-is-better family it compared in the
wrong direction and dropped all results (or applied no filtering) instead
of removing the low-scoring points.
Derive a single higher_score_is_better flag from the query type and use it
for both the ordering and the threshold, and add NaiveFeedbackQuery to the
family. Adds regression tests covering both the ordering and the threshold
on Euclid and Manhattan.
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* fix(local): apply score_threshold strictly to match server semantics
The Qdrant server keeps only points whose score is *better* than
score_threshold (strict inequality): a point whose score equals the
threshold is excluded. Local mode used non-strict comparisons, so such
boundary points were incorrectly kept for Cosine/Dot/Euclid/Manhattan.
Fusion and formula post-filters remain inclusive, matching observed
server behavior for those paths.
Includes parametrized regression tests covering all four distance metrics.
* tests: update tests
* tests: update tests to include other query points ways
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* Fix TypeError sorting heterogeneous facet values and mixed-type point ids in local mode
facet() broke a count tie with the raw facet value and _search_distance_matrix
sorted samples by raw point id; both raise TypeError when the values span
types (e.g. int vs str, or int vs UUID id). Route point-id sorting through the
existing _universal_id helper, and give facet values a dedicated type-safe key
that also keeps equal-count ties deterministic across colliding types (e.g. the
string "" and the int 0). Adds regression tests for both.
* fix: narrow down the fix to search matrix only
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* fix: local mode accepts min_should min_count values the server rejects
Local mode evaluates min_should as `matches >= min_count`. Any value at
or below zero is therefore trivially true for every point, so the filter
returns the entire collection instead of being refused.
The server refuses these outright: 422 Unprocessable Entity for 0, and
400 Bad Request for negatives. So a query that a developer tests against
local mode passes there and fails in production - and until it fails, it
silently returns everything, which for a filter is the worst direction
to be wrong in.
flt = Filter(min_should=MinShould(conditions=[...], min_count=0))
client.scroll("collection", scroll_filter=flt)
# local mode: every point in the collection
# server: 422 Unprocessable Entity
Validation runs once in calculate_payload_mask, before the scan, and
recurses into nested filters since a bad min_count inside a nested must
clause is just as invalid. Raises ValueError with the same shape as the
existing limit validation in qdrant_local.py.
Known limitation, called out rather than hidden: an empty collection
short-circuits in LocalCollection.scroll before any filter code runs, so
an invalid filter against an empty collection is still accepted. Fixing
that means validating in each entry point, which is where #1339 is
already working - happy to move it there instead if preferred.
Verified against Qdrant 1.19.0 in Docker. Full local suite: 87 passed.
* fix: validate filters reached through a NestedCondition
* fix: update filter validation, update tests
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* fix: apply root filters after local fusion
* fix: filter prefetch sources before fusion to match server scores
* fix: push the root filter into prefetch search
* fix: fix local mode filters in prefetch, update tests
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* fix: mirror server token-aware text/phrase matching on unindexed fields
qdrant/qdrant#10341 (dev) changed MatchText and MatchPhrase on fields
without a text index from a substring scan to token-aware matching via
the default word tokenizer: every query token must appear as a whole
document token (text, order-independent; consecutive for phrase), empty
queries match nothing. Local mode still substring-scanned, so congruence
tests randomly failed whenever the filter generator drew a MatchText
whose word is a substring of another fixture word ("fly" in "butterfly",
"ant" in "elephant"). MatchTextAny keeps substring semantics, matching
the server.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R25zh9xS78xMHgPcoFaUdw
* fix: match null elements inside arrays in local IsNull condition
qdrant/qdrant#10101 (dev) made the unindexed IsNull check inspect array
elements: a value like [null, 1] now satisfies IsNull (one level deep).
Local mode only matched values that were null themselves. This was the
second divergence behind the congruence CI failures, previously masked
by the MatchText one because pytest runs with -x.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R25zh9xS78xMHgPcoFaUdw
* fix: close local client before reopening storage in persistence tests
The persistence tests released the storage lock with `del local_client`,
relying on garbage collection timing; when the lock outlived the del,
reopening the same directory raised "Storage folder is already accessed
by another instance". test_query.py was already fixed to call close()
(90913f8); apply the same fix to the remaining five persistence tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R25zh9xS78xMHgPcoFaUdw
* fix: bound remote group hits by exact local hits instead of equality
Server-side grouping is best-effort within a request budget (qdrant
lib/shard/src/grouping/driver.rs): once the budget is spent, a group may
be filled with worse points than its true best, or stay below
group_size. Local mode groups exhaustively, so asserting exact per-rank
score equality of deep group hits randomly failed when the fill budget
missed a group member (test_query_group, local 0.6926 vs remote 0.6798
at rank 4). Compare one-sided instead: at any rank the remote hit may be
worse than the exact local one, never better; the top hit stays strict.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R25zh9xS78xMHgPcoFaUdw
* test: move local text-match and is-null tests to their conventional homes
The two new test files sat at the tests/ root. Local-mode behavior belongs in
qdrant_client/local/tests, and filter corner cases in
tests/congruence_tests/test_complex_filters.py.
- the check_match assertions mirroring the server's unindexed_text_match_test.rs
move into qdrant_client/local/tests/test_payload_filters.py, next to the other
filter unit tests
- the client-level cases become congruence tests in test_complex_filters.py, so
they compare local against a real server instead of asserting local behavior
alone: text/phrase/text-any matching on an unindexed field, and IsNull over
arrays holding a null
Both congruence tests fail against the pre-fix payload_filters and pass with it,
against qdrant 1.19.1-dev.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* tests: add non-consecutive case for match filter
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* Fix in-place mutation of inputs in cosine_similarity
cosine_similarity normalized its `query` and `vectors` arguments in place
via `/=`, which (1) mutated caller-owned arrays as a side effect and
(2) raised UFuncTypeError on integer-dtype inputs, unlike the dot,
euclidean, and manhattan distance functions. Switch to out-of-place
division so the inputs are left untouched. Computed results are unchanged.
Add regression tests asserting the query/vectors arguments are not mutated
(1D and 2D query paths) and that integer-dtype inputs are accepted.
* test: exercise 2D cosine query path with multiple rows
Use a two-row 2D query so the batched per-row normalization path is
verified, and assert the full distance matrix in addition to input
immutability.
* Keep vectors normalization in place per review, fix query only
@joein noted that vectors is always already normalized when reaching
cosine_similarity through the client API (cosine collections are
normalized on upsert), so copying it is unnecessary overhead on what
can be a large candidate set. Revert vectors to in-place normalization
and keep only the query-side fix, which he agreed is worth the
(minimal) copy cost since queries are fresh, user-supplied input each
call and aren't guaranteed to be pre-normalized.
Update tests to match: drop the vectors-not-mutated assertions and the
vectors integer-dtype case (vectors is always float32 in real usage),
keep the query-side mutation and integer-dtype coverage.
grpc.KeywordPrefixParams is an empty message: presence is the only
signal, so an explicit prefix=False cannot be represented in gRPC.
It is sent as absent (same server-side semantics, disabled) and is
recovered as None. Document this at both conversion sites and pin
the behavior with a reverse-direction (rest->grpc->rest) test.
`show_warning_once` defaults to `stacklevel=1`, which makes `warnings.warn`
attribute the warning to `qdrant_client/common/client_warnings.py:7` -- inside
the client -- instead of the caller's construction site.
Every other `show_warning_once` call in the package passes an explicit
stacklevel (4, 5, 6 or 10, depending on nesting); this is the only site that
omits it. Its two immediate siblings in the same `__init__` -- the
`api-key`-in-headers warning and the `grpc.primary_user_agent` warning -- both
pass `stacklevel=4`, so this brings it in line with them.
Before:
.../qdrant_client/common/client_warnings.py:7: UserWarning: `User-Agent` ...
After:
.../qdrant_client/qdrant_client.py:134: UserWarning: `User-Agent` ...
which is where the sibling warnings already point.
`async_qdrant_remote.py` is generated from the sync source; the hunk there
matches what `tools/generate_async_client.sh` emits (verified by running the
generator with ruff pinned to 0.4.3).
* fix: local mode matches wrong points on null and empty-should filters
* tests: move tests to complex filters
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* new: 1.19.0 updates
* fix: fix search params as a dict in local mode
* fix: update qdrant backward compatibility version
* fix: add version check to the test
* fix: add version check to the test
* Fix local mode filters cross-matching booleans and integers
Python treats bool as a subclass of int (True == 1, False == 0), but Qdrant
keeps booleans and integers as distinct payload value types. Local mode
compared them with a plain `==` / `in` / `isinstance(value, (int, float))`, so:
- MatchValue(value=1) matched a payload of True, and MatchValue(value=True)
matched a payload of 1 (same for 0 / False)
- MatchAny / MatchExcept cross-matched the same way
- Range matched booleans as if they were 0 / 1
The server never cross-matches these (its ValueVariants keeps Integer and Bool
distinct, and booleans are not numeric for range conditions). Add a type-aware
equality helper used by the value-match conditions, and exclude booleans from
range checks. Adds an in-memory regression test.
* Cover MatchExcept in the bool/int cross-match test
MatchExcept also routes through values_match, so assert that except=[1]
keeps the True payload (bool is not the integer 1).
* Add isolated MatchAny and range asserts to the bool/int cross-match test
Lock the single-value MatchAny path and the check_range bool guard against
regressions, in addition to the existing combined-condition coverage.
* fix: handle floats in cross-match local mode filters, add congruence tests
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* fix: spurious async client tests failures
* skip cluster-only test when server is standalone
* increase timeout for unit test performing multiple snapshot operations
* clean up stale snapshots left by previous runs
* fix: remove deleted methods, add/update cluster checks
* fix: remove unused import
* fix: remove redundant indent
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* fix: check_match() raises TypeError when MatchText applied to non-string field
* tests: move non-string match test to test_nested_filter, cover MatchText and MatchTextAny
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* fix: update poetry lock
* fix: add type annotations, update poetry.lock
* fix: fix local persistence tests
* fix: replace del client with client.close in local mode persistence tests