The Art of Balancing Reliability and Consistency in Distributed Systems
When Amazon's retail platform began, it was designed as a monolith.
What is a Monolith?
A monolith is a large single application where all the components are interconnected and interdependent. Imagine it like a single, huge machine where all parts work closely together.
Initially, having everything in one place worked great for Amazon. Quick and easy changes were possible because everything was connected. But as they grew, the system had to support many independent services like catalog, checkout, recommendations, and the shopping cart.
What is a Distributed System?
A distributed system is one where components are located on different networked computers but work together as if they were a single system.
Let's picture a library spread across a city. Each branch has its own section of books, but they all function as one giant library to the visitors.
Moving Towards Flexibility
As Amazon's business expanded, they broke the monolithic system into service-owned capabilities.
What are Service-Owned Capabilities?
These are independent components managed by different teams. Each service can evolve separately without affecting others.
This is like each library branch managing its own collection and staff but still sharing books.
The shift to distributed systems allowed individual teams more freedom. However, it also brought new challenges, especially for Amazon's shopping cart.
The Shopping Cart Problem
At the heart of any online shopping experience is the "Add to Cart" button. It’s crucial in the buying process — a delay or failure here means unhappy customers. Amazon found that managing the cart was not a simple database issue anymore.
The Dynamo Approach
Amazon created a system called Dynamo to tackle these issues. Dynamo aimed to provide high availability, ensuring the system remains operational, even in the face of failures.
What is High Availability?
High availability refers to a system's ability to operate continuously without fail for a designated period of time. It is crucial for systems that need to be reliable and responsive continuously.
Imagine a bustling restaurant that never closes and always has enough staff to serve everyone quickly, even during rush hours.
To achieve this, Dynamo used several techniques:
- Consistent Hashing: A way to distribute and store data across servers evenly.
- Replication: Keeping multiple copies of data on different servers to ensure availability.
- Sloppy Quorum: A flexible way of managing how many data copies must agree for a transaction to be considered successful.
- Hinted Handoff: Temporarily storing data on alternative nodes if the primary ones aren’t available.
- Merkle-Tree Anti-Entropy: A method of efficiently checking for differences in data between servers.
- Vector Clocks: Tracking changes over time to help manage data consistency.
What is Consistent Hashing?
In simple terms, consistent hashing is a method to evenly distribute data across many servers. Think of it like spreading books evenly across a series of library shelves.
Each of these methods helped Dynamo keep the cart system quick and reliable despite potential failures.
Dealing with Conflict
A challenge with distributed systems is conflict resolution. Two updates might happen simultaneously, leading to divergent versions of data.
What is Conflict Resolution?
Conflict resolution involves determining how to resolve conflicts between different data changes to build a consistent view.
Amazon chose to resolve conflicts within the application, letting the system decide how to merge changes automatically without blocking user actions, like adding socks to a cart.
The Results and Realizations
By making these architectural choices, Amazon ensured that temporary system issues didn’t prevent customers from shopping. But this meant that sometimes the complexity moved from the database to the application.
If something like a lost cart item reappeared, it was all part of Amazon choosing reliability and uptime over perfection.
The Core Lessons
- Distributed systems allow services to be independent and flexible but come with new challenges related to consistency and reliability.
- High availability techniques ensure systems continue working correctly under stress, but they can introduce complexity elsewhere.
- Configurations like Dynamo’s ensure systems are resilient by using spreading data, replication, and smart conflict handling strategies.
- Choosing where to manage complexity — from the database to the application layer — is crucial in distributed system design.
Ultimately, distributed systems require careful decisions on where to place the pain points and trade-offs between reliability and consistency.
