Files
Tim ViséeandClaude Fable 5.1 81b9a84bd9 Fix stale snapshot transfer breaking cluster (#10643)
* Check shard transfer requests against the registered transfer source

A sender that restarts replays committed but unapplied consensus entries
before it has joined consensus, so a `Start` for a transfer that has
since been aborted still spawns its driver. The receiver accepted such a
stale sender: `initiate_shard_transfer` only required *some* transfer
into the shard, and the pre-download clear in `recover_shard_snapshot`
only required the replica not to be a source of truth. A `Partial`
replica being populated by another transfer passed both, and got wiped.

Have the sender identify itself in the internal `initiate` and shard
snapshot `recover` requests, and have the receiver refuse both unless a
transfer from that peer into the shard is registered. The field is
optional, so requests from older peers keep the previous behavior.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Hold off a shard transfer driver until a consensus leader is established

On startup, committed but unapplied consensus entries are replayed
before the peer joins consensus. A `Start` or `Restart` replayed this
way spawns the transfer driver right away, on a view of consensus that
predates the restart: the transfer may already be aborted, and the peer
is about to abort every transfer it is part of anyway once a leader is
established (`cancel_related_transfers`). That driver still contacted
the receiver, which acted on the stale transfer.

Make the driver task wait until this peer knows a consensus leader
before it touches the remote, so the entries that come with the leader,
such as an abort of this transfer, are applied first. Stopping the task
cancels the wait like any other stage of the transfer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Test that only the registered source may clear a shard for snapshot recovery

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Test a restarted sender replaying an aborted transfer start on a real cluster

A staging-only delay in the sender's `Start` apply, before it registers
anything, lets a test kill the sender while the entry is committed but
not yet applied. That is the state the incident's sender crashed in: on
restart, `Consensus::new` replays the entry before the peer rejoins
consensus and spawns the driver of a transfer the cluster has aborted
and replaced meanwhile.

The test marks the sender `Dead` with an update while it is down, so
the transfer aborts and the receiver is recovered from the other
replica, then restarts the sender while the receiver sits in `Partial`.
Without the fixes on this branch the replayed driver clears the
receiver, its download is cut when the sender catches up, the dummy is
marked `Active` by the legitimate `Finish`, and the receiver's consensus
dies on the next transfer entry. With them the cluster converges.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Treat a dummy local shard as nothing to un-proxify

Transfer restarts and aborts un-proxify the sender's local shard when
applied. A dummy shard was reported as an unexpected type with a service
error, and consensus apply treats a service error as fatal: consensus
stops, and on the next start the replay of the same committed entry
fails again before the peer opens raft networking, so it never comes
back. That is how the receiver of a stale snapshot transfer ended up in
a crash loop: its shard had been cleared under a transfer, the dummy
was marked `Active` by the legitimate `Finish`, and the next transfer
entry naming it as the sender was fatal.

A dummy was never proxified, so there is nothing to revert. Return
without error for it, and leave every other case as it was. A peer
already stuck this way then starts, applies the abort of that transfer
or receives a consensus snapshot, and its `Dead` replica is recovered
through the normal path, which recreates the shard behind the dummy.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Nuance in comment

* Require a registered transfer to clear a shard for snapshot recovery

The check only ran when the sender identified itself. A sender running an
older version doesn't, and neither does the REST endpoint, so both cleared
the shard unchecked.

Hold them to *some* transfer into this shard being registered, which is
what `initiate_shard_transfer` required of them before. The REST endpoint
cannot select the shard transfer priority that the clear hangs off, so
nothing user-facing changes.

* Test that clearing for snapshot recovery needs a registered transfer

* On snapshot recovery, source target shard ID from remote shard

* Reject immediately if another transfer is ongoing

* Add test

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-28 17:56:49 +02:00
..