Files
qdrant/tests/consensus_tests/test_write_ordering.py
qdrant-cloud-bot 242c7dfa0a ci: parallelize consensus tests with pytest-xdist (#8717)
* ci: parallelize consensus tests with pytest-xdist

Enable pytest-xdist for consensus_tests to run tests across multiple
workers in parallel, significantly reducing CI wall time (~20min → ~5-7min).

Changes:
- Add `-n auto --dist=loadfile` to the consensus test pytest invocation
- Remove hardcoded port_seed from tests that don't need fixed ports for
  restart/rejoin (test_order_by, test_consensus_compaction,
  test_named_vector_crud, test_listener_node)
- Give test_cluster_rejoin its own PORT_SEED=15000 to avoid port
  conflicts with auth tests (PORT_SEED=10000)
- Derive restart ports from killed PeerProcess objects instead of
  hardcoded arithmetic where possible
- Add xdist_group("auth") marker to auth test files to ensure they
  run on the same worker (they share PORT_SEED=10000)

Made-with: Cursor

* fix: remove remaining hardcoded port_seed=20000 causing parallel test conflicts

8 test files were using port_seed=20000 as a positional argument to
start_cluster(), which was missed in the initial change. When running
in parallel with pytest-xdist, multiple workers would try to bind to
the same port range (20000-20x02), causing port conflicts and cascading
test failures.

Also remove port_seed=23000 from test_snapshot_recovery_kill.py since
it doesn't need fixed ports for restart.

Made-with: Cursor

* fix: use saved port for restart in test_two_follower_nodes_down

The test was restarting killed peers on hardcoded ports (20200/20100)
that previously matched port_seed=20000. After switching to random
ports, the restart ports no longer match the original peer ports,
causing raft state URI mismatches and peer startup failures.

Save the p2p_port from the killed PeerProcess and reuse it for restart.

Made-with: Cursor

* Reuse p2p ports when restarting killed peers in consensus tests

When a peer is killed and restarted with random ports, it gets a new
consensus URI. The cluster needs a Raft operation to update this URI,
which under CPU contention from parallel test workers can exceed the
30-second timeout. Fix by capturing each peer's p2p_port before killing
and reusing it on restart, so the URI stays the same and no consensus
update is needed.

Made-with: Cursor

* A few improvements for parallel runs (#8731)

* fix: make auth tests' PORT_SEED per-worker to avoid port collisions
* ci: improve failure visibility for parallel consensus tests
* Three small changes to make hangs, interleaved output, and coverage runs behave predictably under pytest-xdist
* fix: two test bugs surfaced by parallel runs and revert drop PR_SET_PDEATHSIG helper
* fix: wait for count convergence in test_triple_replication
* fix: clean leaked peer processes at test start

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Cursor Agent <agent@cursor.com>
Co-authored-by: tellet-q <166374656+tellet-q@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-08 13:47:29 +02:00

140 lines
4.2 KiB
Python

import pathlib
from .fixtures import create_collection
from .utils import *
from .assertions import assert_http_ok
N_PEERS = 3
N_SHARDS = 1
N_REPLICA = 3
# Test update write ordering guarantees.
def test_write_ordering(tmp_path: pathlib.Path):
assert_project_root()
# seed port to reuse the same port for the restarted nodes
peer_api_uris, peer_dirs, bootstrap_uri = start_cluster(tmp_path, N_PEERS)
create_collection(peer_api_uris[0], shard_number=N_SHARDS, replication_factor=N_REPLICA)
wait_collection_exists_and_active_on_all_peers(
collection_name="test_collection",
peer_api_uris=peer_api_uris
)
max_peer_url = fetch_highest_peer_id(peer_api_uris)
url_index = peer_api_uris.index(max_peer_url)
# Kill update leader peer
print(f"Stopping update leader peer with url {max_peer_url} at index {url_index}")
p = processes.pop(url_index)
restart_port = p.p2p_port
p.kill()
peer_api_uris.pop(url_index)
# write ordering weak (will trigger the detection of dead replicas)
r = requests.put(
f"{peer_api_uris[0]}/collections/test_collection/points?wait=true&ordering=weak", json={
"points": [
{
"id": 1,
"vector": [0.05, 0.61, 0.76, 0.74],
"payload": {
"city": "Berlin",
"country": "Germany",
}
}
]
})
assert_http_ok(r)
# Assert that there are dead replicas
wait_for_some_replicas_not_active(peer_api_uris[0], "test_collection")
# write ordering medium
r = requests.put(
f"{peer_api_uris[0]}/collections/test_collection/points?wait=true&ordering=medium", json={
"points": [
{
"id": 1,
"vector": [0.05, 0.61, 0.76, 0.74],
"payload": {
"city": "Munich",
"country": "Germany",
}
}
]
})
# succeeds as Medium selects an Alive peer
assert_http_ok(r)
# Write ordering with filter update operation
r = requests.post(
f"{peer_api_uris[0]}/collections/test_collection/points/delete?wait=true&ordering=medium", json={
"filter": {
"must": [
{
"key": "a",
"match": {
"value": 100
}
}
]
}
}
)
# succeeds as Medium selects an Alive peer
assert_http_ok(r)
# write ordering strong
r = requests.put(
f"{peer_api_uris[0]}/collections/test_collection/points?wait=true&ordering=strong", json={
"points": [
{
"id": 1,
"vector": [0.05, 0.61, 0.76, 0.74],
"payload": {
"city": "Dresden",
"country": "Germany",
}
}
]
})
# fails as highest peer is dead
assert r.status_code == 500
# Restart peer
new_url = start_peer(peer_dirs[url_index], f"peer_{url_index}_restarted.log", bootstrap_uri, port=restart_port)
peer_api_uris.append(new_url)
# Wait for peers to be online
wait_for_peer_online(new_url)
time.sleep(1)
timeout = 10
while True:
if timeout < 0:
raise Exception("Timeout waiting for all replicas to be active")
all_active = True
points_counts = set()
for peer_api_uri in peer_api_uris:
res = check_collection_cluster(peer_api_uri, "test_collection")
points_counts.add(res['points_count'])
if res['state'] != 'Active':
all_active = False
if all_active:
if len(points_counts) != 1:
for peer_api_uri in peer_api_uris:
res = requests.post(f"{peer_api_uri}/collections/test_collection/points/count", json={"exact": True})
print(res.json())
assert False, f"Points count is not equal on all peers: {points_counts}"
break
time.sleep(1)
timeout -= 1