Step 1: The Surprising Villain — Slowness, Not Death
Here's a truth that sounds backwards at first:
A slow dependency is more dangerous than a dead one.
Why? Because dead things fail fast. If a service is completely down, your call to it fails in milliseconds. You get an error, you handle it, you move on.
But a slow service? It doesn't say no. It says "hold on... hold on... hold on..." — and while you hold on, you're stuck. Your thread is stuck. And threads, as we're about to see, are a resource you can run out of.
What is a dependency?
A dependency is any other system your code needs to call to do its job — a database, a payment service, a recommendations service. If your service calls it, you depend on it.
What is a thread?
A thread is a worker inside your program. A typical web server has a limited pool of them — say, 200 threads. Each incoming request is handled by one thread. If a thread is waiting for a slow dependency to answer, it can't do anything else. It's like a waiter standing frozen at the kitchen window, waiting for one dish, unable to serve any other table.
Step 2: How One Slow Service Kills a Healthy One
Let's watch the failure happen in slow motion. Imagine a simple chain:
User ──▶ Service A ──▶ Service B ──▶ Service C (getting slow)
Service C starts taking 30 seconds to respond instead of 30 milliseconds. Nothing has crashed. Here's what happens next:
- Service B's threads start waiting on C. Each request to B grabs a thread, calls C, and waits.
- Threads pile up. New requests keep arriving at B, each one grabbing a fresh thread... which also gets stuck waiting on C.
- B's thread pool drains. All 200 of B's threads are now frozen, waiting on C. B has no free workers left.
- Now B looks slow too. Service A calls B, and B can't respond — no free threads. So A's threads start waiting on B.
- The failure climbs. A's pool drains. Then A's callers drain. One sick service at the bottom, and the whole call graph goes down.
The Netflix engineers who lived through this described it perfectly: the failure climbs the call graph like a fire in a stairwell.
What is a thread pool?
A thread pool is the fixed set of worker threads a service keeps ready. Fixed is the key word. When they're all busy, new work waits in line — or gets dropped. Example: a restaurant with exactly 10 waiters. When all 10 are frozen at the kitchen window, nobody greets new customers.
What is a call graph?
The call graph is the map of who-calls-whom. A → B → C is a tiny one. At Netflix, a single user request fans out across dozens of microservices, out of hundreds total. That means one slow service can poison dozens of paths at once.
And here's the cruel part: the architecture diagram might say "high availability" all over it. Redundant servers, multiple zones, the works. None of that helps, because nothing crashed. The system died from waiting.
Step 3: The Fix Has a Name — the Circuit Breaker
Netflix's answer was a library called Hystrix: it wraps every remote call in a circuit breaker (with timeouts and bulkheads alongside — more on those in a moment). They open-sourced it, and the pattern went mainstream.
What is a circuit breaker?
A circuit breaker in software is borrowed from the one in your house's electrical panel. When your home wiring detects a dangerous overload, the breaker trips — it cuts the circuit instantly so the house doesn't burn down. The software version does the same thing: when calls to a dependency keep failing, the breaker trips and stops even trying, so the failure can't spread.
The idea in one line:
If a dependency is sick, stop calling it. Fail instantly instead of waiting.
That sounds harsh — refusing to even try? But remember Step 2. The waiting is what kills you. An instant "no" frees your thread immediately. Your service stays alive, keeps serving everything else, and can show the user a fallback ("Recommendations unavailable right now") instead of hanging forever.
What is a bulkhead?
A bulkhead is a companion pattern: give each dependency its own small, separate pool of threads, like watertight compartments in a ship. If the "Service C" compartment floods (all its threads stuck), the other compartments stay dry. Hystrix pairs bulkheads with breakers so one bad dependency can't drain the whole pool while the breaker is still deciding whether to trip.
Step 4: The Three States (With Hystrix's Real Numbers)
A circuit breaker is a tiny state machine with exactly three states. Here they are, using Hystrix's actual defaults:
error rate > ~50%
in 10-sec window
┌────────┐ ─────────────────────▶ ┌────────┐
│ CLOSED │ │ OPEN │
│(normal)│ ◀──────┐ │(reject │
└────────┘ │ │ all) │
▲ │ └───┬────┘
│ probe succeeds │ after ~5 sec
│ │ │ sleep window
│ ┌────┴──────┐ │
└───────│ HALF-OPEN │◀───────────┘
probe fails │ (1 probe) │
→ back OPEN └───────────┘
State 1: Closed — everything is fine
Closed means the circuit is connected — traffic flows normally (yes, it's the electrical naming; closed = current flows). While closed, the breaker quietly does bookkeeping: it counts successes and failures over a rolling window of roughly 10 seconds.
What is a rolling window?
A rolling window means "the last N seconds, always." The breaker doesn't remember all of history — just the most recent ~10 seconds of results. Old results continuously fall out the back as new ones arrive. Like a security camera that only keeps the last 10 seconds of footage.
State 2: Open — the breaker has tripped
If the error rate crosses about 50% inside that 10-second window, the breaker trips to Open.
Now every single call to that dependency fails instantly. The breaker doesn't even attempt the network call. In pseudocode:
def call_dependency(request):
if breaker.state == OPEN:
raise FailFast() # returns in microseconds
# → no thread waits
# → no queue builds
# → caller gets fallback immediately
result = actually_call_service_c(request) # only when closed/probing
breaker.record(result)
return result
This is the whole magic: no thread waits, no queue builds. Fail fast, on purpose. Your service stays healthy even while its dependency is on fire.
State 3: Half-Open — testing the waters
The breaker can't stay open forever — Service C will eventually recover, and you want traffic back. But you can't just fling the door wide open either (imagine 10,000 queued-up requests slamming into a service that's barely back on its feet).
So after a sleep window of about 5 seconds, the breaker moves to Half-Open and lets exactly one probe request through:
- Probe succeeds → the dependency seems healthy → circuit closes, normal traffic resumes.
- Probe fails → still sick → circuit snaps back open, and the 5-second sleep starts again.
Fail fast, recover deliberately. That's the whole design.
One request as a scout, not an army. That's the "deliberately" part.
Step 5: The Part People Skip — the Real Cost
Here's the honest bit that rarely makes it into design reviews.
An open breaker rejects requests you might have been able to serve. When the error rate is 50%, the other 50% of calls were succeeding! The moment the breaker trips, those would-have-succeeded calls get rejected too. You are deliberately throwing away good traffic to protect the rest of the system.
That's a real cost. Real users get real errors that didn't strictly have to happen. The right move is to say this out loud in the design review, not to pretend the breaker is free:
The trade you're making is total collapse for explicit, bounded rejection. Take that trade every time. Just know you're making it.
Total collapse: everyone waits, everything drains, the whole platform goes down. Bounded rejection: some users get a fast, clean failure on one feature, and everything else keeps working.
The second one is clearly better. But it's a trade, not a free lunch.
Step 6: Tuning — Where Good Intentions Go Wrong
The breaker has knobs: the error threshold (~50%), the window (~10s), the sleep window (~5s). Tuning them is genuinely hard, and there are two classic failure modes.
Trap 1: Threshold too tight
Set the trip threshold too sensitive (say, trip at 10% errors over 2 seconds), and a brief blip — one garbage-collection pause, one dropped packet burst — trips the breaker for no reason. Now you're rejecting good traffic because of noise.
Trap 2: Sleep window too short — the breaker "flaps"
This one is nastier. Set the sleep window too short and you get this loop:
OPEN ──▶ probe ──▶ CLOSED ──▶ full traffic slams
▲ succeeds the recovering service
│ │
└────────── TRIPS AGAIN ◀───────────┘
open → probe → close → hammer → trip → open → ...
The dependency was just starting to recover. Your probe squeaks through, the breaker closes, and full traffic instantly buries the fragile service again. It falls over, the breaker trips, sleeps briefly, probes, closes, hammers...
You can DDoS your own struggling service with recovery attempts.
What is DDoS?
DDoS = Distributed Denial of Service — overwhelming a service with more traffic than it can handle. Usually it's an attack from outside. Here, hilariously and tragically, you're doing it to yourself with your own retry-and-recover machinery. Like a crowd rushing back into a shop the second the "reopening soon" sign appears, trampling the staff before they've restocked a single shelf.
The fix is patience in the config: a sleep window long enough for genuine recovery, and (in more advanced setups) ramping traffic back gradually instead of all at once.
Step 7: What a Breaker Actually Buys You
It's worth being precise about what this pattern does and doesn't do:
- It does not fix the slow dependency. Service C is still sick.
- It does not prevent failure. Users of the C-powered feature still see errors.
- It does draw a firebreak. The failure stops at the breaker instead of climbing the stairwell.
A circuit breaker doesn't prevent failure. It decides where failure stops.
That reframe is the deepest lesson here. In a large system, failure is not optional — some dependency, somewhere, is always having a bad day. The design question is never "how do we prevent all failure?" It's "when failure happens, where does it stop, and who decides?" A circuit breaker is you deciding, in advance, on purpose.
The core lessons
- Slow is worse than dead. Dead dependencies fail fast; slow ones hold your threads hostage until your whole thread pool drains and the failure climbs the call graph.
- A circuit breaker turns slow failures into fast ones. Closed = normal traffic with bookkeeping; Open = instant rejection (trips at ~50% errors over a ~10-second rolling window); Half-Open = one probe after a ~5-second sleep to test recovery.
- Fail fast, recover deliberately. Instant rejection protects your threads; a single probe protects the recovering dependency.
- An open breaker throws away good traffic. Some rejected calls would have succeeded. That's the price — name it out loud in your design review.
- Tuning is hard in both directions. Too tight and blips trip it for nothing; too short a sleep window and it flaps — open, probe, close, hammer, trip — until you've DDoS'd your own recovering service.
- You're trading total collapse for explicit, bounded rejection. Take that trade every time. Just know you're making it.
- A circuit breaker doesn't prevent failure — it decides where failure stops. In systems that survive production, someone chose that spot on purpose.
