Step 1: Why global billing needs stronger guarantees
Imagine two datacenters, one in Europe and one in Oregon. A user clicks an ad. The system must debit an account exactly once, even if one site loses power. If the same click is counted twice, money is lost or overcharged.
Early systems used Bigtable with eventual consistency.
What is eventual consistency?
Eventual consistency means replicas may temporarily disagree, but they will match after enough time passes with no new writes. Example: Two copies of a counter may read 5 and 7 for a few seconds, then both settle on 12.
This worked for most analytics because a few seconds of lag was acceptable and replication stayed simple.
Step 2: The hidden cost of loose ordering
When the same approach was tried for financial debits, two problems appeared.
First, two-phase commit (2PC) across continents hit 150 ms round-trip time. Coordinators could not safely decide commit order without risking a double spend.
Second, machine clocks drifted by several milliseconds. A transaction could appear committed in one datacenter while still in-flight in another.
Adding synchronous Paxos groups across regions pushed tail latency past 400 ms. The system became too slow for production.
Step 3: Spanner's new building blocks
Spanner solves the ordering problem with two ideas that work together: the TrueTime API and commit-wait.
What is TrueTime?
TrueTime is an API that returns a time interval [earliest, latest] instead of a single clock reading. The interval captures the maximum possible error from atomic clocks and GPS receivers.
What is commit-wait?
Commit-wait means the coordinator deliberately pauses until the uncertainty window has passed before declaring the transaction committed.
Step 4: How a write actually flows
Here is the sequence for one transaction:
Client
│
â–¼ 1. Ask TrueTime for [earliest, latest]
Coordinator
│
â–¼ 2. Run Paxos to agree on data
│
â–¼ 3. Wait until wall time > latest
│ (the commit-wait step)
│
â–¼ 4. Write chosen timestamp + data
Replicas
The chosen timestamp is stored in a new column on every tablet. All future reads and writes can now compare these timestamps directly.
Because every replica sees the same globally ordered timestamps, the system can guarantee that no two transactions will ever be reordered across datacenters.
Step 5: The price you pay on purpose
In production the uncertainty window was roughly 7 ms. Every write therefore waits those 7 ms plus the Paxos round trips. Reads stay fast because they can still hit the nearest replica.
The engineers accepted the extra wait rather than risk double-spends. As the source material states:
Strong consistency at global scale is mostly just waiting out your uncertainty bound until the clocks say it is safe.
The core lessons
- Eventual consistency is simple but allows temporary disagreements that break financial correctness.
- Wall-clock time cannot be trusted across datacenters because of clock skew measured in milliseconds.
- TrueTime returns an interval that bounds uncertainty instead of a single point in time.
- Commit-wait forces the coordinator to pause until the entire uncertainty window has passed, creating a safe global order.
- The cost is deliberate latency (about 7 ms per write in early Spanner) traded for correctness that two-phase commit alone could not deliver at global scale.
