Files
qdrant/lib
7e7f19fc41 Keep consensus operation awaiters alive for concurrent waiters (#10338)
* Keep consensus operation awaiters alive for concurrent waiters

Callers proposing an identical consensus operation deduplicate onto one
broadcast channel, and the map holds its only sender. Removing the entry on
timeout therefore closed the channel for every other waiter, failing their
still in-flight operation with "Channel sender dropped".

Only remove the entry once no receiver is left, dropping our own receiver
first so the last caller out cleans up. Apply the same to
await_for_multiple_operations, which registered awaiters but never
deregistered them when it timed out.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Add debug assert to ensure we clean up consensus operation waiters

* Deregister consensus operation awaiters when the waiter is dropped

Dispatcher::submit_collection_meta_op registers the expected operations before
proposing, then drops that future unpolled whenever the proposal itself fails.
The awaiters stayed in the map with no receiver left, so the next identical
request deduplicated onto a dead entry and never heard back. This is what
tripped the new debug assert in CI: a rejected create-collection left a
SetShardReplicaState awaiter behind, and the next run of the same test hit it.

Move registration into an OperationAwaiters guard that deregisters on drop, so
timeout, drop-before-poll and request cancellation are all covered by one path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Close the race the awaiter debug assert trips on

The assert is sound only if no one can observe an entry whose receivers are all
gone. Both cleanup sites dropped their receiver before taking the map lock, so a
concurrent register could see exactly that and panic. Drop the receiver while
holding the lock instead, and take that lock once per batch rather than once per
operation: creating a collection registers an awaiter per replica, on the mutex
the consensus thread needs for every entry it applies.

Collect the awaiters into the guard as we go, so giving up part way still
deregisters the ones already registered, and only build the broadcast channel
when the operation is not already in-flight.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: timvisee <tim@visee.me>
2026-08-27 14:52:05 +02:00
..
2026-07-30 20:59:31 +00:00
2026-07-30 20:59:31 +00:00
2026-07-30 20:59:31 +00:00
2026-07-30 20:59:31 +00:00