mirror of
https://github.com/qdrant/qdrant.git
synced 2026-09-30 09:58:03 -05:00
* Keep consensus operation awaiters alive for concurrent waiters Callers proposing an identical consensus operation deduplicate onto one broadcast channel, and the map holds its only sender. Removing the entry on timeout therefore closed the channel for every other waiter, failing their still in-flight operation with "Channel sender dropped". Only remove the entry once no receiver is left, dropping our own receiver first so the last caller out cleans up. Apply the same to await_for_multiple_operations, which registered awaiters but never deregistered them when it timed out. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Add debug assert to ensure we clean up consensus operation waiters * Deregister consensus operation awaiters when the waiter is dropped Dispatcher::submit_collection_meta_op registers the expected operations before proposing, then drops that future unpolled whenever the proposal itself fails. The awaiters stayed in the map with no receiver left, so the next identical request deduplicated onto a dead entry and never heard back. This is what tripped the new debug assert in CI: a rejected create-collection left a SetShardReplicaState awaiter behind, and the next run of the same test hit it. Move registration into an OperationAwaiters guard that deregisters on drop, so timeout, drop-before-poll and request cancellation are all covered by one path. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Close the race the awaiter debug assert trips on The assert is sound only if no one can observe an entry whose receivers are all gone. Both cleanup sites dropped their receiver before taking the map lock, so a concurrent register could see exactly that and panic. Drop the receiver while holding the lock instead, and take that lock once per batch rather than once per operation: creating a collection registers an awaiter per replica, on the mutex the consensus thread needs for every entry it applies. Collect the awaiters into the guard as we go, so giving up part way still deregisters the ones already registered, and only build the broadcast channel when the operation is not already in-flight. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: timvisee <tim@visee.me>