All resources
Daily Learning · Release It!: Design and Deploy Production-Ready Software

Canary releases: how to ship software without breaking the whole world

Thursday, 9 July 2026
Canary releases: how to ship software without breaking the whole world

🎯 You'll understand what a canary release is, why smart teams roll changes out slowly through stages, and how one skipped step caused a global internet outage — teaching us why a canary you're allowed to skip isn't really a canary at all.

Imagine you cook one new dish and serve it to the entire city at dinner. If the recipe is wrong, everyone gets sick at once. A wiser cook tastes it first, then serves one table, then a few tables, then the whole restaurant. That slow, careful path is the heart of today's lesson.

Let's build the idea up piece by piece.


Step 1: What "deploying" even means

When engineers deploy, they push a new version of their software so real users start using it.

What is a deploy?

A deploy = the act of putting new code (or new settings) into the live system that real customers touch.

The scary part: the moment you deploy, real people feel the result. If it's broken, they feel that too.

So the whole question becomes: how do we deploy without hurting everyone at once if we're wrong?


Step 2: The safe answer — roll it out slowly

Instead of flipping the change on for 100% of users instantly, you turn it on for a tiny slice first. You watch. If nothing breaks, you widen it. Step by step.

This is called a staged rollout.

What is a staged rollout?

A staged rollout = releasing a change to bigger and bigger groups over time, checking health at each stage before moving on.

Here's the exact path a well-run company (Cloudflare) used:

DOG  →  PIG  →  CANARY  →  GLOBAL

Let's define each stage, because the names are cute but the meaning matters.

  • DOG — deploy to your own traffic first. Like a chef eating their own dish. If it poisons anyone, it poisons you.
  • PIG — deploy to a small slice of real customers. A first small bite of the outside world.
  • CANARY — deploy to a somewhat larger sample and watch closely for problems.
  • GLOBAL — everyone on the planet gets it.

What is a canary?

A canary = a small early group that gets the change first, so if something's wrong, they show the symptoms before the whole population does.

The name comes from old coal mines. Miners carried a canary bird underground. If poisonous gas built up, the bird got sick first — warning the miners to get out. The bird was an early alarm.

A canary release is an early-warning bird for your software.

Step 3: Why you wait between stages

Between each stage, you let the change "bake."

What is a bake / bake time?

Bake time = a waiting period where the new version runs, while you watch dashboards and error rates to see if anything is going wrong.

Deploy to CANARY
   │
   ▼
Watch for 30 min... errors normal? CPU normal? latency normal?
   │
   ▼ (only if healthy)
Promote to GLOBAL

If the canary group shows trouble, you stop. You never reach GLOBAL. The damage is contained to a tiny slice.

This is the whole safety net. Now here's where it gets interesting — because sometimes people skip the net on purpose.


Step 4: The tempting shortcut — "this one's too urgent to wait"

Cloudflare's main job is protecting websites. One tool they use is a WAF.

What is a WAF?

A WAF = Web Application Firewall. It inspects incoming web requests and blocks malicious ones before they reach a website.

A WAF works using rules — patterns that describe "this request looks like an attack, block it."

Now, some of those rules exist to stop zero-day attacks.

What is a zero-day attack?

A zero-day = a brand-new attack that's being used against people right now, before there's a fix. "Zero days" of warning.

Here's the dilemma. A staged rollout takes hours. But a zero-day is hurting customers this second. Waiting hours to slowly bake a protection rule means customers get hacked while you wait.

So Cloudflare made a deliberate decision:

WAF rules skip the canary path and ship to the entire planet in seconds.

They did this using a fast distribution system called Quicksilver.

What is Quicksilver / a KV layer?

Quicksilver was Cloudflare's KV distribution layer. KV = Key-Value — a simple store that maps a name (key) to a value, like a giant dictionary. Its job here: push a new rule to every datacenter worldwide almost instantly.

The reasoning was honest and reasonable. Security rules are supposed to protect people now. But remember the safety net they gave up: no bake, no canary, straight to global.


Step 5: The day it all broke

July 2, 2019, 13:42 UTC. Someone shipped one WAF rule. Inside it was a pattern that looked roughly like this:

.*(?:.*=.*)

That's a regex.

What is a regex?

A regex (regular expression) = a mini-language for describing text patterns. For example, the pattern \d{3} means "three digits in a row." WAFs use regex to describe what an attack looks like.

To use a regex, the computer runs a regex engine that checks each request against the pattern.

What is a backtracking regex engine?

A backtracking engine tries one way to match the pattern; if it fails, it goes back and tries another way, and another, and another. For most patterns this is fine. But some patterns cause it to try an explosion of combinations.

That's exactly what .(?:.=.*) did. It caused catastrophic backtracking.

What is catastrophic backtracking?

Catastrophic backtracking = when a regex pattern makes the engine try so many combinations that the work explodes. Here the cost grew quadratically — meaning if the input doubles, the work roughly quadruples. On certain inputs, matching a single request could burn enormous CPU.

What is CPU?

The CPU = the "brain" of a computer that does the actual work. When it's "pinned at 100%," it's completely maxed out — like a chef so overwhelmed by one impossible order that no other orders get cooked.

Here's what happened, step by step:

New rule → shipped GLOBALLY in seconds (no canary)
   │
   ▼
Every incoming request runs the bad regex
   │
   ▼
Regex backtracks catastrophically → CPU hits 100%
   │
   ▼
Every CPU core, in every datacenter, worldwide → pinned
   │
   ▼
No CPU left to serve real traffic → 502 errors everywhere

What is a 502?

A 502 = "Bad Gateway." It's an HTTP error meaning the server that was supposed to answer couldn't. To users, the internet just... broke.

Global traffic dropped about 82% within minutes. A company whose entire product is keeping the internet up had taken a big chunk of the internet down. Painful irony.


Step 6: Stopping the bleeding

The panic fix came at 14:02 UTC: a global WAF kill switch — one big button to turn the whole WAF off everywhere.

What is a kill switch?

A kill switch = an emergency control that instantly shuts something off. Not a fix — just a way to stop the bleeding fast.

Total damage: 27 minutes of worldwide outage. All from skipping the canary for one rule.

13:42  Bad rule deployed globally
   │   ...CPUs pinned, 502s everywhere, ~82% traffic gone
14:02  Kill switch flipped — WAF disabled
   │   Traffic recovers
Total: ~27 minutes

Step 7: The real fix (and its hidden price)

The kill switch was a bandage. The real fixes were structural.

Fix 1 — Every WAF rule now walks the full path.

BEFORE:  WAF rule ────────────────────► GLOBAL   (skipped everything)

AFTER:   WAF rule → DOG → PIG → CANARY → GLOBAL   (no more skipping)

Now a bad rule would burn CPU in a tiny group first, get caught, and never reach the world.

Fix 2 — A safer regex engine. They moved toward a non-backtracking engine.

What is a non-backtracking engine?

A non-backtracking regex engine guarantees that match time is bounded by the length of the input, not by how nasty the pattern is. In plain words: no single request can blow up the CPU no matter how the pattern is written. The pathology becomes impossible.


Now the part most people skip: the fix wasn't free.

By forcing security rules through the canary path, an active exploit now has to wait behind the slow rollout. That's a real cost — attackers get a little more time.

So Cloudflare kept an emergency bypass — but wrapped it in guardrails:

  • Mandatory second review — another human must approve.
  • Simulated deploys — test the rule against real traffic without actually enabling it, to catch the CPU explosion before it happens.

What is a simulated deploy?

A simulated deploy = running the change in a "practice mode" that mimics production and measures its effect, but doesn't actually affect real users. Like a fire drill instead of a real fire.

They didn't get rid of the trade-off. They moved it — from "risk of taking the whole world down" to "a little slower to respond to attacks, plus some process friction." That's a conscious, eyes-open choice. Not a magic win.

Every safety decision is a trade. The goal isn't to remove all risk — it's to move risk to where it hurts least, on purpose.

Step 8: The one line to tattoo on your brain

A canary you're allowed to skip isn't a canary. It's a suggestion.

The whole point of a canary is that it's unskippable. The moment there's a special lane that bypasses it "just this once for urgent stuff," that lane becomes the exact path your worst outage travels.

So the useful question to ask about your own systems:

What's the deploy path that still skips the canary — and what's the justification? If the justification is "it's urgent," remember: urgent changes are the most dangerous ones to ship untested, not the least.


The core lessons

  • Never flip a change on for everyone at once. Roll it out in stages — a small group first — so mistakes hurt a few, not all.
  • The canary is an early-warning bird. Give the change to a tiny slice, watch it "bake," and only promote it if it stays healthy.
  • A skippable canary isn't a canary. The bypass lane you build for "urgent" changes is exactly where your biggest outage will travel.
  • Settings and rules are code too. Cloudflare's outage came from a config rule, not app code. If it can reach production, it should be staged like production.
  • Beware catastrophic backtracking. A single bad regex can pin every CPU. Prefer engines whose match time is bounded by input length, not pattern cleverness.
  • A kill switch stops the bleeding; it doesn't heal. Have one ready — but the real fix is structural.
  • Every safety choice is a trade-off. You don't erase risk; you move it somewhere less catastrophic — and you do it with your eyes open.

Want this kind of thinking applied to your business?

Book a free 30-minute discovery call — we’ll show you your highest-value first automation, no jargon, no obligation.

Book a discovery call