Interactive Raft — 5-server cluster

Click a server to take it offline (click again to restart it). Shift-click a server to force its election timeout. Watch elections, heartbeats and log replication keep the replicated state machines consistent.

follower candidate leader offline RequestVote AppendEntries reply (hollow) client cmd
The ring around a follower/candidate is its randomized election timer. Small dots under a candidate are votes received. Numbers on AppendEntries messages are the count of log entries carried (none = heartbeat).
AppendEntries and catching up a server

The number inside a blue AppendEntries dot is how many log entries that message carries, at most 5 per message. A small dot with no number carries no entries: it is a heartbeat, which tells followers the leader is alive and passes on the leader's commit index. Every AppendEntries also carries prevIndex/prevTerm, the index and term of the entry just before the new ones. The follower accepts the message only if its own log has an entry there with that same term.

How a server is brought up to speed after being offline:

  1. On restart it keeps its persisted log, term and vote, but it has forgotten its commit index and state machine. It starts as a follower and waits.
  2. The leader keeps a nextIndex for each follower: its guess for the next entry to send (the leader row shows next/match). It sends AppendEntries starting from there.
  3. If the follower's log is missing the prevIndex entry, or holds it with a different term, it rejects the message (red hollow reply). The leader lowers nextIndex by one and retries right away. This repeats until the leader reaches the last entry both logs agree on.
  4. From that point the follower accepts. Any conflicting entries after it are truncated, and the leader's entries are appended. The leader sends batches of up to 5 entries back to back until the follower has every entry.
  5. Each reply updates matchIndex. The leaderCommit field lets the follower advance its own commit index and replay the committed entries into its state machine, rebuilding the key/value store.

Try it: take a server offline, send a burst of client requests, then bring it back. You will see a few rejections, then numbered dots streaming the missing entries.

Replicated logs

Each cell is a log entry, colored and labelled by the term it was created in (hover for the command). Solid = committed on that server, faded/dashed = not yet known to be committed. The leader row shows next/match indices for each follower.

State machines (key → value)

Events

Things to try
  • Take the leader offline: followers' timers run out, one becomes candidate, wins a majority and becomes the new leader in a higher term.
  • Take two servers offline: the remaining 3 are still a majority, so requests still commit.
  • Take three offline: the leader can append entries but can never commit them (no majority). Bring one back and they commit.
  • Send requests, kill the leader before entries replicate, then bring it back: its uncommitted entries from the old term get overwritten by the new leader.
  • Restart a server: its state machine is rebuilt from the persisted log once it learns the commit index from the leader.
  • Note a new leader does not commit entries from earlier terms by counting replicas (§5.4.2); they commit once an entry from its own term commits.