Step 1: The Simple Starting Point
Imagine every message as a direct phone call. A client sends a request to a central broker, the broker delivers the message, and the client replies with “got it.” The broker waits for that reply before it can handle the next thing.
This is called synchronous communication.
What is synchronous communication?
It means the sender stops and waits for an answer before doing anything else. Example: calling a friend and staying on the line until they say hello.
In code it looks like this:
client → HTTP POST /send → broker
broker delivers message
broker ← 200 OK ← client
Everything happens in one straight line. When daily active users were in the low millions, this worked fine.
Step 2: The Hidden Cost of Waiting
Each online user keeps a persistent connection open. Every delivery attempt blocks a worker thread until the client acknowledges.
What is a thread?
A thread is a small slice of a CPU’s attention. Think of it as one waiter in a restaurant who can only serve one table at a time.
Under load the number of blocked threads grows. The server measured the exact limit: roughly 200k concurrent connections before context switches and garbage collection pressure collapsed throughput. New connections start queuing and p99 latency jumps from hundreds of milliseconds into multiple seconds.
Worker threads
[ busy ] [ busy ] [ busy ] ... [ 200k limit reached ]
new connections wait here →
Step 3: Switching to an Asynchronous Model
WhatsApp replaced the direct call with a message queue. Clients now receive messages through a non-blocking event loop. The broker simply stores the message and moves on.
What is asynchronous communication?
The sender drops the message into a queue and immediately continues working. It does not wait for the receiver to finish.
The new flow looks like this:
client → enqueue(topic: (user_id, device_id)) → queue
broker later delivers
They built this on Erlang’s lightweight processes. These are tiny, cheap units of work that can number in the millions on a single machine.
Before (sync) After (async)
200k connections ~2 million connections
one thread per call event loop + lightweight processes
The same hardware now handled ten times the connections.
Step 4: What You Give Up
Immediate feedback disappears. Delivery now needs a separate acknowledgment path. Partial outages become harder to notice, so extra monitoring is required. The team accepted this cost because raw connection capacity was the constraint that would otherwise have capped the business.
Synchronous calls feel safe until the number of waiting threads exceeds your cores.
The core lessons
- Synchronous communication blocks a thread for every in-flight request.
- Thread count is limited by CPU cores and context-switch overhead; WhatsApp hit ~200k connections.
- An asynchronous queue lets the broker enqueue and forget, freeing the thread immediately.
- Lightweight processes (as in Erlang) make millions of concurrent “waiters” affordable.
- You trade instant visibility for dramatically higher capacity and must add monitoring to compensate.
