feat: replication: reset target backoff on transfer - #2145
Merged
Merged
Conversation
# Summary Applications can end the active AppendEntries backoff for selected replication targets, or every target when no node ids are given. # Details A follower may be reachable again while its replication stream still waits after an RPC failure. Calling `reset_backoff()` before `Trigger::transfer_leader()` lets replication resume without that delay. A reset ends the current or next wait of the active backoff and drops its `Backoff` iterator. `BackoffState::rank` remains until an RPC succeeds, so another failure can start backoff again. A reset sent before backoff starts does not affect the later backoff. Snapshot transfer has a separate retry backoff. The reset receiver goes to `ReplicationCore::spawn()`. It stays out of `ReplicationContext`, which is also used by `SnapshotTransmitter`. This avoids giving snapshot tasks a receiver whose sender is dropped. `ExternalCommandName` gains `ResetBackoff` for metrics. - Fix: databendlabs#2143
# Summary Leadership transfer now ends the target's active AppendEntries backoff so a restarted node can catch up before the transfer request times out. # Details A reachable transfer target may still be waiting through a retry delay after an earlier RPC failure. `RaftCore` signals that target's reset channel before it starts the leadership transfer. `Config::reset_backoff_on_transfer_leader` is a new `Option<bool>`. `None` and `Some(true)` enable the reset; `Some(false)` keeps the existing backoff delay. Configs written before this field was added also enable the reset.
xp-trumpet
approved these changes
Sep 29, 2026
xp-trumpet
left a comment
Collaborator
There was a problem hiding this comment.
@xp-trumpet reviewed 15 files and all commit messages.
Reviewable status:complete! all files reviewed, all discussions resolved (waiting on drmingdrmer).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Changelog
feat: replication: reset target backoff on transfer
Summary
Leadership transfer now ends the target's active AppendEntries backoff
so a restarted node can catch up before the transfer request times out.
Details
A reachable transfer target may still be waiting through a retry delay
after an earlier RPC failure.
RaftCoresignals that target's resetchannel before it starts the leadership transfer.
Config::reset_backoff_on_transfer_leaderis a newOption<bool>.NoneandSome(true)enable the reset;Some(false)keeps theexisting backoff delay. Configs written before this field was added
also enable the reset.
feat: replication: add Trigger::reset_backoff()
Summary
Applications can end the active AppendEntries backoff for selected
replication targets, or every target when no node ids are given.
Details
A follower may be reachable again while its replication stream still
waits after an RPC failure. Calling
reset_backoff()beforeTrigger::transfer_leader()lets replication resume without that delay.A reset ends the current or next wait of the active backoff and drops
its
Backoffiterator.BackoffState::rankremains until an RPCsucceeds, so another failure can start backoff again. A reset sent
before backoff starts does not affect the later backoff. Snapshot
transfer has a separate retry backoff.
The reset receiver goes to
ReplicationCore::spawn(). It stays out ofReplicationContext, which is also used bySnapshotTransmitter.This avoids giving snapshot tasks a receiver whose sender is dropped.
ExternalCommandNamegainsResetBackofffor metrics.This change is