All resources
Daily Learning · Release It!: Design and Deploy Production-Ready Software

How to Stop Your Servers from Attacking Themselves: The Magic of Deadline Budgets

Monday, 6 July 2026
How to Stop Your Servers from Attacking Themselves: The Magic of Deadline Budgets

🎯 Discover how Twitter cured its notorious "Fail Whale" outages by replacing rigid, static timeouts with dynamic deadline budgets that protect databases from being crushed by their own retry attempts.

Imagine you are running a highly popular restaurant.

A customer walks in and orders a gourmet burger. Your server takes the order to the kitchen. Normally, the kitchen takes 10 minutes to make a burger. But tonight, the kitchen is super busy, and the wait stretches to 11 minutes.

At exactly the 10-minute mark, the server gets impatient. Instead of waiting another minute, they yell, "Forget that first burger, it's taking too long!" and immediately submit a second order for the exact same burger to the kitchen.

Now, the kitchen has to keep cooking the first burger (because they already started) and start cooking a second, identical burger.

As you can guess, this makes the kitchen even slower. More servers get impatient, more duplicate orders get fired, and soon the kitchen is completely buried under orders for food that will be thrown in the trash.

This is exactly what happens to websites when they use bad timeout settings. It is how Twitter, in its early days, constantly crashed and showed users the infamous "Fail Whale" error screen.

Let's look step-by-step at how this happens and how engineers fixed it using a brilliant concept called Deadline Budgets.


Step 1: The Chain of Commands

In a large system like Twitter, one single click on your screen triggers a chain reaction of computers talking to each other.

[ User Feed ] ──(Request)──> [ Frontend Service ] ──(RPC)──> [ Database ]

To make this happen smoothly, Twitter migrated its system to run on the JVM and built a special communication tool called Finagle.

What is JVM?

JVM stands for Java Virtual Machine. It is a engine that runs software programs. It allows systems to handle thousands of tasks at the exact same time.

What is Finagle?

Finagle is an open-source library built by Twitter. It acts like a highly organized postal service, managing how different servers talk, share data, and handle errors with one another.

What is RPC?

RPC stands for Remote Procedure Call. It is when one server tells another server on the network to run a specific task. Example: Your Frontend server sends an RPC to your Database server saying, "Give me the latest tweets for Sarah."

To keep the website feeling snappy, engineers set a rule on the Frontend server: If a database request takes longer than 1000 milliseconds (1 second), give up.

This is called a static timeout.


Step 2: The Trap of Static Timeouts

A static timeout seems like a great idea. If a database query is stuck, you don't want your user staring at a loading spinner forever.

What is a Timeout?

A timeout is a safety timer. If a server does not get a response within a set limit, it stops waiting and assumes something went wrong. Example: "If the database doesn't answer in 1000ms, abort the connection."

But look what happens when the database gets slightly busy. Let's say a query that usually takes 500ms slows down to 1050ms under heavy holiday traffic.

Here is the play-by-play of the disaster:

Time:   0ms                    1000ms        1050ms
Req 1:  [=======================(Timeout)====(Finished, but abandoned!)]
Req 2:                         [========================================...]
  1. At 0ms: The Frontend asks the Database for data (Request 1).
  2. At 1000ms: The Frontend's timer hits its 1000ms limit. The Frontend says, "Too slow! I'm canceling this request!"
  3. At 1001ms: The Frontend immediately fires a Retry (Request 2) to try again.
  4. At 1050ms: The Database finally finishes processing Request 1. But the Frontend has already abandoned it! The Database just wasted precious CPU cycles doing useless work.
  5. Meanwhile, the Database is now starting to work on Request 2.

What is a Retry?

A retry is an automatic action where a system tries a failed task again. Example: If a web page fails to load, your browser automatically tries to load it again a split second later.


Step 3: The Thundering Herd

When thousands of users are on the app, this creates a catastrophic loop.

Because the first requests are slightly slow, the frontend cancels them and fires thousands of brand-new retries. The database is now working on the old, abandoned requests and the new, duplicate retry requests at the same time.

This triggers a Thundering Herd.

Database CPU Load:
正常 (Normal):    [████░░░░░░] 40%
重试风暴 (Storm):  [██████████] 100% (Saturated!)

What is a Thundering Herd (or Retry Storm)?

A thundering herd occurs when many client systems all request resources at the exact same time, overwhelming the server and causing it to crash.

What is Thread Saturation?

Thread saturation is when all of a computer's active workers (threads) are completely occupied with tasks, leaving no room to process any new incoming work.

The database spends all its power running queries that have already been abandoned by the frontend. The system has self-sabotaged.


Step 4: The Solution — Deadline Budgets

To save their platform, Twitter's engineers realized they needed a way to share a single "clock" across the entire chain of servers. They implemented Context-Propagated Deadline Budgets.

Instead of every service having its own independent 1000ms timer, they gave the entire request a single overall budget of time.

What is Context Propagation?

Context propagation is the act of passing metadata (extra information) along with a request as it travels from one microservice to another. Example: Passing a "Tracking ID" or a "User ID" through three different servers so they all know who triggered the action.

What is a Deadline Budget?

A deadline budget is a dynamic countdown timer passed from service to service. It tells downstream systems exactly how much time is left before the overall user request will time out. Example: context.deadline = current_time + remaining_budget

Let's look at how this works in practice.

Imagine a user request starts with a total budget of 1000ms.

  1. The request spends 995ms waiting in line at the Frontend.
  2. The Frontend passes the request to the Database via an RPC. Because of context propagation, it attaches a header saying: "Hey Database, you only have 5ms left in our budget to finish this."
  3. The Database looks at this budget:
// Check the remaining time in the deadline budget
remainingTime := ctx.RemainingBudget()

if remainingTime < 5 * time.Millisecond {
    // Fail Fast! Do not start the heavy work.
    return Error("Out of time! Rejecting request.")
}

The database immediately sees that 5ms is not enough time to run a query.

Instead of starting a heavy query that will just get abandoned anyway, it fails fast and rejects the request instantly.

What is Failing Fast?

Failing fast is an engineering practice where a system stops an operation immediately when it knows success is impossible, rather than wasting resources trying to finish it.

By failing fast, the database saves its CPU and memory to work on requests that actually have a chance of finishing on time!


The Trade-off: Graceful Degradation

Implementing deadline budgets forced Twitter to make a hard decision.

By rejecting requests that were out of budget, some users would occasionally see an error message (like a small "Could not load tweets" box) instead of waiting.

Engineers sometimes resist showing errors because they want a perfect user experience. But shielding your users with infinite retries is a illusion. It guarantees that a small, temporary slowdown will snowball into a total, site-wide crash.

By choosing to let a few requests fail gracefully, Twitter kept the overall platform fast, healthy, and alive.


## The core lessons

  • Timeouts must be shared: Hardcoded, static timeouts on individual services are a trap. They cause services to work on requests that have already been abandoned.
  • Pass the budget: Use context propagation to send a dynamic deadline budget (context.deadline) down your entire chain of servers.
  • Fail fast: If a downstream service receives a request with a remaining budget that is too small to succeed, it should reject it immediately.
  • Protect the database: A database under heavy load should never waste CPU cycles processing dead, timed-out requests.
  • Partial failure is better than total collapse: It is always better to show a fast, graceful error to 5% of your users than to let a retry storm crash your entire system for 100% of your users.

Want this kind of thinking applied to your business?

Book a free 30-minute discovery call — we’ll show you your highest-value first automation, no jargon, no obligation.

Book a discovery call