Skip to content

feat: replication: reset target backoff on transfer - #2145

Merged
drmingdrmer merged 2 commits into
databendlabs:mainfrom
drmingdrmer:reset-backoff
Sep 29, 2026
Merged

drmingdrmer merged 2 commits into
databendlabs:mainfrom
drmingdrmer:reset-backoff

Conversation

@drmingdrmer

@drmingdrmer drmingdrmer commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Changelog

feat: replication: reset target backoff on transfer

Summary

Leadership transfer now ends the target's active AppendEntries backoff
so a restarted node can catch up before the transfer request times out.

Details

A reachable transfer target may still be waiting through a retry delay
after an earlier RPC failure. RaftCore signals that target's reset
channel before it starts the leadership transfer.

Config::reset_backoff_on_transfer_leader is a new Option<bool>.
None and Some(true) enable the reset; Some(false) keeps the
existing backoff delay. Configs written before this field was added
also enable the reset.

feat: replication: add Trigger::reset_backoff()

Summary

Applications can end the active AppendEntries backoff for selected
replication targets, or every target when no node ids are given.

Details

A follower may be reachable again while its replication stream still
waits after an RPC failure. Calling reset_backoff() before
Trigger::transfer_leader() lets replication resume without that delay.

A reset ends the current or next wait of the active backoff and drops
its Backoff iterator. BackoffState::rank remains until an RPC
succeeds, so another failure can start backoff again. A reset sent
before backoff starts does not affect the later backoff. Snapshot
transfer has a separate retry backoff.

The reset receiver goes to ReplicationCore::spawn(). It stays out of
ReplicationContext, which is also used by SnapshotTransmitter.
This avoids giving snapshot tasks a receiver whose sender is dropped.

ExternalCommandName gains ResetBackoff for metrics.



This change is Reviewable

# Summary

Applications can end the active AppendEntries backoff for selected
replication targets, or every target when no node ids are given.

# Details

A follower may be reachable again while its replication stream still
waits after an RPC failure. Calling `reset_backoff()` before
`Trigger::transfer_leader()` lets replication resume without that delay.

A reset ends the current or next wait of the active backoff and drops
its `Backoff` iterator. `BackoffState::rank` remains until an RPC
succeeds, so another failure can start backoff again. A reset sent
before backoff starts does not affect the later backoff. Snapshot
transfer has a separate retry backoff.

The reset receiver goes to `ReplicationCore::spawn()`. It stays out of
`ReplicationContext`, which is also used by `SnapshotTransmitter`.
This avoids giving snapshot tasks a receiver whose sender is dropped.

`ExternalCommandName` gains `ResetBackoff` for metrics.

- Fix: databendlabs#2143
# Summary

Leadership transfer now ends the target's active AppendEntries backoff
so a restarted node can catch up before the transfer request times out.

# Details

A reachable transfer target may still be waiting through a retry delay
after an earlier RPC failure. `RaftCore` signals that target's reset
channel before it starts the leadership transfer.

`Config::reset_backoff_on_transfer_leader` is a new `Option<bool>`.
`None` and `Some(true)` enable the reset; `Some(false)` keeps the
existing backoff delay. Configs written before this field was added
also enable the reset.

@xp-trumpet xp-trumpet left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@xp-trumpet reviewed 15 files and all commit messages.
Reviewable status: :shipit: complete! all files reviewed, all discussions resolved (waiting on drmingdrmer).

@drmingdrmer
drmingdrmer merged commit 4f2d135 into databendlabs:main Sep 29, 2026
52 of 53 checks passed
@drmingdrmer
drmingdrmer deleted the reset-backoff branch September 29, 2026 14:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Availability: TransferLeader timeout due to Replication Backoff

2 participants