Files
qdrant/tests/openapi/test_sparse_idf_corpus.py
Andrey Vasnetsov 8a7325ad5d Remove deprecated search endpoints from OpenAPI, deprecate them in gRPC (#9982)
* Remove deprecated search/recommend/discover endpoints from OpenAPI

Remove deprecated REST API endpoint definitions from the OpenAPI
generator. These endpoints were deprecated in v1.13.3 (`f4ced2567`,
#5907, 2025-01-30) in favor of the universal `/points/query` endpoint:

- POST /points/search
- POST /points/search/batch
- POST /points/search/groups
- POST /points/recommend
- POST /points/recommend/batch
- POST /points/recommend/groups
- POST /points/discover
- POST /points/discover/batch

Also removes the corresponding request types from the schema generator
and updates the expected API count in the consistency check.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Migrate OpenAPI integration tests to /points/query

The deprecated /points/search, /points/recommend and /points/discover
endpoints (along with their /batch and /groups variants) were removed
from the OpenAPI spec, which caused validation failures in the Python
integration test harness.

This commit migrates the affected tests to the universal /points/query
endpoint:

- Delete tests dedicated to the deprecated endpoints:
  test_recommend.py, test_discover.py, test_multicollection_reco.py,
  test_recommendation_multivector.py
- Refactor remaining tests to call /points/query (and /query/batch,
  /query/groups), translating request bodies (vector -> query / using,
  positive/negative -> query.recommend, target/context -> query.discover)
  and unwrapping the new result.points response shape.
- Drop equivalence assertions against the now-removed legacy endpoints.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Relax non-empty assertions in migrated recommend/discover tests

The previous migration added `len(...) > 0` assertions to tests that
previously only checked equivalence between the deprecated and new
API. These assertions are too strict because the parametrized
`query_filter` cases legitimately produce empty result sets.

Drop the `> 0` assertion and rely on `request_with_validation` to
verify the response is well-formed and HTTP OK.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Migrate remaining OpenAPI tests off deprecated search endpoints

Tests added to dev after the original migration was written still call
/points/search and /points/recommend/groups through
`request_with_validation`, which resolves the endpoint against the
OpenAPI spec and therefore breaks once the endpoint is not in the spec:

- test_turbo4_storage.py, test_sparse_idf_corpus.py, test_validation.py:
  translate /points/search to /points/query (vector{name,vector} ->
  query + using, result -> result.points).
- test_group.py: drop the /points/recommend/groups half of the
  lookup_from validation test in favour of the query equivalent.

test_sparse_idf_corpus.py's test_query_api_supports_idf_corpus goes
away: with the helper on /points/query every test in the file now
exercises what it asserted.

Also record why test_recommend_group cannot assert on its groups: it
uses every point in the collection as a recommend example, so all of
them are excluded and the result is legitimately empty.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Regenerate openapi.json without the deprecated search endpoints

Drops the 8 deprecated paths and the request schemas that only they
referenced: Search/Recommend/Discover request (+Batch, +Groups) types
and their exclusive dependencies (NamedVector, NamedSparseVector,
NamedVectorStruct, UsingVector, RecommendExample, ContextExamplePair).

Regenerated output is a strict subset of the previous spec, and every
remaining $ref still resolves.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Deprecate the search/recommend/discover RPCs in gRPC

The REST counterparts have carried `deprecated: true` since v1.13.3 and
are now gone from the OpenAPI spec, while the gRPC RPCs never got any
deprecation annotation at all. Mark all 8 with `option deprecated = true`
so generated clients warn, and point each doc comment at its `Query`
replacement.

tonic puts `#[deprecated]` on the generated client methods only; the
server trait gets the doc comment alone, so our own `impl` is unaffected.
The RPCs keep serving traffic — this is annotation only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Restore the deleted recommend/discover suites on /points/query

The earlier migration deleted these four files outright, but the
query-side tests it left behind are all shallow smoke tests
(`len(result) > 0`, `"points" in result[0]`). The deleted ones carried
invariants with no query-API equivalent anywhere, so deleting them was a
real loss of coverage rather than de-duplication:

- test_recommend.py: default strategy equals average_vector; batch
  results identical to sequential singles across six request shapes;
  best_score with only negatives yields all-negative scores; best_score
  with a single positive orders identically to a nearest query; raw
  vectors as examples equal ids as examples.
- test_discover.py: context-only scores are all <= 0; target-only orders
  identically to a nearest query but scores differently; with a fixed
  context the integer part of the score is stable while the decimal part
  moves, and vice versa with a fixed target; batch equals singles;
  lookup_from by id equals by vector.
- test_multicollection_reco.py: cross-collection lookup_from, plus
  wrong-vector-size, unknown-collection and unknown-vector rejections.
- test_recommendation_multivector.py: the same recommend invariants over
  a max_sim multivector collection, which the query suite never covered.

Only test_recommend_missing_lookup_from_collection_with_raw_vector is
dropped as genuinely redundant — test_query.py's
test_query_missing_lookup_from_collection covers query, query/batch and
prefetch.

Two request-shape differences the translation had to absorb:

- Giving no examples at all is 422 (a RecommendInput validation rule),
  where the legacy API reported 400 from the query itself. A malformed
  example, such as an empty vector, is still 400.
- DiscoverInput requires the `context` key and accepts only an explicit
  null to mean "no context", so target-only discover must spell it out.
  The legacy API let it be omitted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 18:09:43 +02:00

240 lines
7.5 KiB
Python

import math
import pytest
from .helpers.collection_setup import drop_collection
from .helpers.helpers import request_with_validation
SPARSE_VECTOR_NAME = "text"
@pytest.fixture(autouse=True)
def setup(collection_name):
corpus_collection_setup(collection_name=collection_name)
yield
drop_collection(collection_name=collection_name)
# Deterministic layout, mirrored by the Rust integration tests:
#
# | point | tenant | sparse dims |
# |-------|--------|-------------|
# | 0 | a | 0, 1 |
# | 1 | a | 0 |
# | 2 | b | 0, 1, 2 |
# | 3 | b | 1 |
POINTS = [
("a", [0, 1]),
("a", [0]),
("b", [0, 1, 2]),
("b", [1]),
]
QUERY = {"indices": [0, 1, 2], "values": [1.0, 1.0, 1.0]}
def expected_idf(n, df):
"""Advanced IDF formula used by the engine for the `idf` modifier."""
return math.log((n - df + 0.5) / (df + 0.5) + 1.0)
def tenant_filter(value):
return {"must": [{"key": "tenant", "match": {"value": value}}]}
def corpus_collection_setup(collection_name):
response = request_with_validation(
api='/collections/{collection_name}',
method="DELETE",
path_params={'collection_name': collection_name},
)
assert response.ok
response = request_with_validation(
api='/collections/{collection_name}',
method="PUT",
path_params={'collection_name': collection_name},
body={
# Dense vector without IDF, to check `idf` param rejection
"vectors": {
"dense": {
"size": 2,
"distance": "Dot",
}
},
"sparse_vectors": {
SPARSE_VECTOR_NAME: {
"modifier": "idf",
}
},
},
)
assert response.ok
response = request_with_validation(
api='/collections/{collection_name}/points',
method="PUT",
path_params={'collection_name': collection_name},
query_params={'wait': 'true'},
body={
"points": [
{
"id": idx,
"vector": {
"dense": [1.0, 0.0],
SPARSE_VECTOR_NAME: {
"indices": dims,
"values": [1.0] * len(dims),
},
},
"payload": {"tenant": tenant, "idx": idx},
}
for idx, (tenant, dims) in enumerate(POINTS)
]
},
)
assert response.ok
def search(collection_name, filter=None, params=None, limit=10):
body = {
"query": QUERY,
"using": SPARSE_VECTOR_NAME,
"limit": limit,
}
if filter is not None:
body["filter"] = filter
if params is not None:
body["params"] = params
response = request_with_validation(
api='/collections/{collection_name}/points/query',
method="POST",
path_params={'collection_name': collection_name},
body=body,
)
assert response.ok, response.json()
return {point['id']: point['score'] for point in response.json()['result']['points']}
def test_explicit_global_matches_default(collection_name):
default_scores = search(collection_name)
explicit_global = search(collection_name, params={"idf": "global"})
assert default_scores == explicit_global
# Global statistics: N = 4, df = [3, 3, 1]
idf = [expected_idf(4, 3), expected_idf(4, 3), expected_idf(4, 1)]
assert default_scores[0] == pytest.approx(idf[0] + idf[1])
assert default_scores[1] == pytest.approx(idf[0])
assert default_scores[2] == pytest.approx(idf[0] + idf[1] + idf[2])
assert default_scores[3] == pytest.approx(idf[1])
def test_corpus_decoupled_from_retrieval_filter(collection_name):
# IDF over tenant a only (N = 2, df = [2, 1, 0]), retrieval unfiltered:
# every point is scored against tenant a's term statistics.
scores = search(
collection_name,
params={"idf": {"corpus": tenant_filter("a")}},
)
idf = [expected_idf(2, 2), expected_idf(2, 1), expected_idf(2, 0)]
assert scores[0] == pytest.approx(idf[0] + idf[1])
assert scores[1] == pytest.approx(idf[0])
assert scores[2] == pytest.approx(idf[0] + idf[1] + idf[2])
assert scores[3] == pytest.approx(idf[1])
def test_filter_tightening_does_not_move_scores(collection_name):
# With a fixed corpus, adding or tightening the retrieval filter narrows
# the result set but must not change any returned score.
params = {"idf": {"corpus": tenant_filter("b")}}
broad = search(collection_name, params=params)
narrow = search(collection_name, filter=tenant_filter("b"), params=params)
assert set(narrow) == {2, 3}
for point_id, score in narrow.items():
assert score == broad[point_id]
def test_empty_corpus_never_falls_back_to_global(collection_name):
# A corpus matching nothing yields degenerate but corpus-scoped scores:
# idf(0, 0) = ln(2) per term, never another population's statistics.
scores = search(
collection_name,
params={"idf": {"corpus": tenant_filter("missing")}},
)
ln2 = expected_idf(0, 0)
assert scores[0] == pytest.approx(2 * ln2)
assert scores[1] == pytest.approx(ln2)
assert scores[2] == pytest.approx(3 * ln2)
assert scores[3] == pytest.approx(ln2)
def search_expecting_error(collection_name, query, using, params, expected_status):
response = request_with_validation(
api='/collections/{collection_name}/points/query',
method="POST",
path_params={'collection_name': collection_name},
body={
"query": query,
"using": using,
"params": params,
"limit": 10,
},
)
assert not response.ok, response.json()
assert response.status_code == expected_status
return response.json()['status']['error']
def test_corpus_accepts_any_filter(collection_name):
# The corpus is a regular filter: any filter shape the search API accepts
# works, not just `match` conjunctions.
# `should` over both tenants matches every point — same statistics as
# global.
global_scores = search(collection_name)
should_scores = search(
collection_name,
params={
"idf": {
"corpus": {
"should": [
{"key": "tenant", "match": {"value": "a"}},
{"key": "tenant", "match": {"value": "b"}},
]
}
}
},
)
assert should_scores == global_scores
# A range corpus over `idx < 2` selects points {0, 1} — the same
# population as tenant a, so the same statistics: N = 2, df = [2, 1, 0].
range_scores = search(
collection_name,
params={"idf": {"corpus": {"must": [{"key": "idx", "range": {"lt": 2}}]}}},
)
tenant_a_scores = search(
collection_name,
params={"idf": {"corpus": tenant_filter("a")}},
)
assert range_scores == tenant_a_scores
idf = [expected_idf(2, 2), expected_idf(2, 1), expected_idf(2, 0)]
assert range_scores[2] == pytest.approx(idf[0] + idf[1] + idf[2])
def test_idf_params_require_idf_modifier(collection_name):
# On a dense vector the `idf` param cannot apply — it is rejected rather
# than silently ignored.
error = search_expecting_error(
collection_name,
[1.0, 0.0],
"dense",
{"idf": {"corpus": tenant_filter("a")}},
expected_status=400,
)
assert "idf" in error