mirror of
https://github.com/qdrant/qdrant.git
synced 2026-09-28 17:07:49 -05:00
* Check shard transfer requests against the registered transfer source A sender that restarts replays committed but unapplied consensus entries before it has joined consensus, so a `Start` for a transfer that has since been aborted still spawns its driver. The receiver accepted such a stale sender: `initiate_shard_transfer` only required *some* transfer into the shard, and the pre-download clear in `recover_shard_snapshot` only required the replica not to be a source of truth. A `Partial` replica being populated by another transfer passed both, and got wiped. Have the sender identify itself in the internal `initiate` and shard snapshot `recover` requests, and have the receiver refuse both unless a transfer from that peer into the shard is registered. The field is optional, so requests from older peers keep the previous behavior. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Hold off a shard transfer driver until a consensus leader is established On startup, committed but unapplied consensus entries are replayed before the peer joins consensus. A `Start` or `Restart` replayed this way spawns the transfer driver right away, on a view of consensus that predates the restart: the transfer may already be aborted, and the peer is about to abort every transfer it is part of anyway once a leader is established (`cancel_related_transfers`). That driver still contacted the receiver, which acted on the stale transfer. Make the driver task wait until this peer knows a consensus leader before it touches the remote, so the entries that come with the leader, such as an abort of this transfer, are applied first. Stopping the task cancels the wait like any other stage of the transfer. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Test that only the registered source may clear a shard for snapshot recovery Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Test a restarted sender replaying an aborted transfer start on a real cluster A staging-only delay in the sender's `Start` apply, before it registers anything, lets a test kill the sender while the entry is committed but not yet applied. That is the state the incident's sender crashed in: on restart, `Consensus::new` replays the entry before the peer rejoins consensus and spawns the driver of a transfer the cluster has aborted and replaced meanwhile. The test marks the sender `Dead` with an update while it is down, so the transfer aborts and the receiver is recovered from the other replica, then restarts the sender while the receiver sits in `Partial`. Without the fixes on this branch the replayed driver clears the receiver, its download is cut when the sender catches up, the dummy is marked `Active` by the legitimate `Finish`, and the receiver's consensus dies on the next transfer entry. With them the cluster converges. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Treat a dummy local shard as nothing to un-proxify Transfer restarts and aborts un-proxify the sender's local shard when applied. A dummy shard was reported as an unexpected type with a service error, and consensus apply treats a service error as fatal: consensus stops, and on the next start the replay of the same committed entry fails again before the peer opens raft networking, so it never comes back. That is how the receiver of a stale snapshot transfer ended up in a crash loop: its shard had been cleared under a transfer, the dummy was marked `Active` by the legitimate `Finish`, and the next transfer entry naming it as the sender was fatal. A dummy was never proxified, so there is nothing to revert. Return without error for it, and leave every other case as it was. A peer already stuck this way then starts, applies the abort of that transfer or receives a consensus snapshot, and its `Dead` replica is recovered through the normal path, which recreates the shard behind the dummy. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Nuance in comment * Require a registered transfer to clear a shard for snapshot recovery The check only ran when the sender identified itself. A sender running an older version doesn't, and neither does the REST endpoint, so both cleared the shard unchecked. Hold them to *some* transfer into this shard being registered, which is what `initiate_shard_transfer` required of them before. The REST endpoint cannot select the shard transfer priority that the clear hangs off, so nothing user-facing changes. * Test that clearing for snapshot recovery needs a registered transfer * On snapshot recovery, source target shard ID from remote shard * Reject immediately if another transfer is ongoing * Add test --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>