All resources
Daily Learning · Software Architecture: The Hard Parts

Split Brain: Why "Fast Failover" Can Break Your Database in Two

Friday, 3 July 2026
Split Brain: Why "Fast Failover" Can Break Your Database in Two

🎯 You'll understand how automatic database failover works, why trusting a single failure signal can create two "primary" servers at once, and why waiting a few extra seconds for agreement is worth it.

Let's learn one of the most humbling lessons in running big systems: sometimes the tool you built to save you is the very thing that hurts you. We'll build this up slowly, one piece at a time.

By the end, you'll understand a real failure that hit GitHub on October 21, 2021, and the elegant idea that fixed it.


Step 1: The boring, wonderful setup

Let's start with how many big websites store their data.

        Writes
          │
          ▼
   ┌─────────────┐
   │   PRIMARY   │  ◄── the one server that accepts changes
   └─────────────┘
          │  copies its changes
          ▼
   ┌─────────────┐   ┌─────────────┐
   │  REPLICA 1  │   │  REPLICA 2  │  ◄── read-only copies
   └─────────────┘   └─────────────┘

GitHub ran on MySQL for years like this.

What is MySQL?

MySQL is a database — a program that stores your data in an organized way (users, orders, comments) and lets you read and write it. Think of it as a giant, very reliable notebook that many programs share.

What is a primary and a replica?

  • The primary is the one server allowed to accept changes (writes). Example: "user signed up" gets written here.
  • A replica is a live copy of the primary. It only serves reads (like "show me this profile").

Why this setup? It's battle-tested and boring in the best way. Boring is a compliment here — it means predictable. It spreads out the reading load, and half the team already knows how to debug it at 3am.

The golden rule to remember: only one server accepts writes at a time. Everything else copies from it. Hold onto that rule — the whole story hinges on it.


Step 2: What happens when the primary dies?

Servers crash. Hardware fails. When the primary dies, nobody can write anymore — signups fail, orders stop. You need a new primary, fast.

The obvious move: promote a replica to become the new primary.

What is failover?

Failover = automatically switching from a broken component to a healthy one, so the system keeps working. Like a co-pilot taking the controls the moment the pilot passes out.

You could page a sleepy engineer to promote a replica by hand. But that's slow and error-prone. So GitHub used a tool to do it automatically: Orchestrator.

What is Orchestrator?

Orchestrator is a topology manager. A topology is just the map of which server is primary and which are replicas. Orchestrator watches that map, and when the primary dies, it promotes a healthy replica to take over.

   Orchestrator watching...
          │  "Is the primary alive?"
          ▼
   ┌─────────────┐
   │   PRIMARY   │  ✗ no response!
   └─────────────┘
          │
          ▼
   Orchestrator: "Primary is dead. Promote Replica 1."

So far, this is a good design. Automated recovery beats a human paging themselves at 3am. What could go wrong?


Step 3: The trap — a network that lies

Here's where reality gets messy. Networks don't always fail cleanly. Sometimes a server is isolated — perfectly alive, but unreachable for a moment.

What is a network partition?

A network partition is when the network splits so some machines can't talk to others — even though every machine is still running fine. Imagine two rooms in a building where the phone line between them goes dead. Both rooms are full of people, working normally. They just can't call each other.

A partial partition is even sneakier: only some connections drop, not all.

Now watch what Orchestrator sees during a partial partition:

   Orchestrator's view:
          │  "Primary, are you there?"
          ▼
   ┌─────────────┐
   │   PRIMARY   │   (silence — the line is cut,
   └─────────────┘    but the server is ALIVE)

Orchestrator can't reach the primary. From its one viewpoint, the primary looks dead. So it does exactly what it was told: it promotes a replica.

But the "dead" primary wasn't dead. It was isolated. It's still happily accepting writes from any app that can still reach it.


Step 4: Split brain — two servers both think they're the boss

Now we've broken the golden rule from Step 1. Two servers believe they are the writable primary at the same time.

   App group A ──writes──►  ┌─────────────┐
                            │ OLD PRIMARY │  "I'm the primary!"
                            │  (isolated) │
                            └─────────────┘

   App group B ──writes──►  ┌─────────────┐
                            │ NEW PRIMARY │  "No, I'M the primary!"
                            │  (promoted) │
                            └─────────────┘

What is split brain?

Split brain is when a system that should have one leader ends up with two, each acting independently. Like a company where two people both think they're CEO, both signing contracts — you end up with contradictory decisions and a legal mess.

Now the damage begins. Writes land on both servers. When the network heals and the two try to reconcile, they can't agree on what actually happened.

What is a replication position?

Replicas keep track of how far along they are in copying changes — a bookmark called the replication position. When two servers each have their own conflicting history of writes, their bookmarks don't line up.

The replication threads choked on conflicting positions. Reconciling them took hours of manual work — not seconds.

Read that again. A tool built to reduce downtime extended it. Instead of a quick recovery, engineers had to untangle two diverging versions of the truth by hand.

Why? Because Orchestrator trusted a single failure signal in a situation with multiple failure modes. It asked one question — "can I reach the primary?" — and treated one silent answer as proof of death.

A failover system that trusts a single signal isn't resilient. It's just a faster way to break both halves.

Step 5: The fix — stop trusting one pair of eyes

The root problem: decisions were made from one vantage point. One observer, one health check. If that one view was fooled, the whole system was fooled.

The fix is beautifully simple in spirit: don't act until several independent observers agree.

What is consensus-based failure detection?

Consensus means agreement among multiple parties. Instead of one probe deciding "the primary is dead," you ask several independent probes, placed in different parts of the network. Only if enough of them agree does the promotion happen.

Think of a jury. You don't convict on one person's opinion — you require a group to agree first.

   Probe A: "Can't reach primary" ✗
   Probe B: "Can't reach primary" ✗
   Probe C: "Can't reach primary" ✗
        │
        ▼
   Enough agree → SAFE to promote

Versus the old, dangerous way:

   Probe A: "Can't reach primary" ✗
        │
        ▼
   One vote → promote anyway → SPLIT BRAIN risk

In config terms, GitHub set something like a minimum number of probes that must agree:

min_detection_consensus = <require several independent probes to agree>

If only one probe can't reach the primary but the others can, that's a clue the problem is the network to that one probe — not a dead primary. So you don't promote. Crisis avoided.


Step 6: Fencing — make sure the old boss truly steps down

Consensus stops you from promoting wrongly. But you also need to guarantee the old primary can't keep writing after you promote a new one. That's fencing.

What is fencing?

Fencing means forcibly cutting off the old primary so it cannot accept writes anymore — demoting it "hard." It's like changing the locks and revoking the old CEO's badge the instant the new one is named. Even if the old server wakes back up, it can't touch the data.

   Promote NEW primary
        │
        ▼
   FENCE the old one ──► "You are read-only now. No more writes."

Together, consensus (agree before acting) plus fencing (make the old boss powerless) closes the door on split brain.


Step 7: The honest trade-off nobody likes to mention

Here's the part most explanations skip — and it's the most important lesson of all.

Consensus made failover slower.

You now have to wait for multiple nodes to agree before promoting. That adds seconds to recovery in the normal case. When the primary really is dead, you're sitting in read-only mode a little longer while the observers confirm.

So you're trading:

   OLD:  Fast promotion  →  but risks HOURS of tangled writes (rare, catastrophic)
   NEW:  A few seconds slower  →  but split brain is prevented (common, safe)

GitHub took that trade with their eyes open. A few extra seconds of read-only time beats hours of corrupted, divergent data.

Slower and correct wins over fast and wrong.

This is the whole game in resilient systems. Resilience isn't about never failing — it's about failing safely. A system that recovers fast but sometimes recovers wrong isn't resilient at all.


Step 8: The big picture in one line

Here's the chain of ideas to carry with you:

   ONE SIGNAL  =  SPLIT BRAIN  →  MANY AGREE  →  SAFE FAILOVER
  • One signal → trusting a single health check.
  • Split brain → that single signal gets fooled, two primaries appear.
  • Many agree → require consensus from independent observers.
  • Safe failover → recovery that's a touch slower, but always correct.

The core lessons

  1. Only one primary should ever accept writes. The entire nightmare comes from breaking that one rule.
  2. A silent server isn't always a dead server. Networks partition — a machine can be alive but isolated. Don't confuse "unreachable" with "dead."
  3. Never make a big, destructive decision from a single viewpoint. One probe can be fooled. Require consensus — agreement from multiple independent observers — before promoting.
  4. Fence the old primary hard. Demoting the old boss so it cannot write is what actually closes the door on split brain.
  5. Speed can be a trap. Automation that trusts one signal isn't resilient; it's just a faster way to break both halves.
  6. Slower and correct beats fast and wrong. A few extra seconds of read-only time is a cheap price to avoid hours of tangled, contradictory data.

Resilience is not the absence of failure. It's the discipline to fail carefully.

Want this kind of thinking applied to your business?

Book a free 30-minute discovery call — we’ll show you your highest-value first automation, no jargon, no obligation.

Book a discovery call