Let's start with a puzzle that sounds backwards.
Spotify built one of the most admired software systems in the tech world. And yet, at their peak, their own engineers couldn't find things inside it. They'd spend hours asking each other, "Wait — who even owns this piece?"
How does a great design lead to that? To understand it, we need to build the idea up slowly. Let's go one step at a time.
Step 1: The starting point — one big program
Most software begins life as a monolith.
What is a monolith?
A monolith is one single program that does everything — login, payments, search, recommendations — all bundled together and deployed as one unit.
Think of it like a single giant restaurant kitchen where every cook shares the same counter, the same fridge, and the same stove.
┌─────────────────────────┐
│ THE MONOLITH │
│ login | payments | │
│ search | recommendations│
└─────────────────────────┘
deployed as ONE thing
This works well at first. Everything is in one place. But as the company grows, a problem appears.
Step 2: The monolith's real problem — coupling
When everyone shares one kitchen, they get in each other's way.
What is coupling?
Coupling means different parts are tangled together, so changing one part can break another.
Imagine the payments cook needs to swap the stove. But the search cook is using that same stove right now. Nobody can change anything without checking with everybody.
In software terms:
- To ship a tiny fix, you must redeploy the whole thing.
- One team's bug can crash a totally unrelated feature.
- Teams have to wait in line to release — a release train.
What is a release train?
A release train is when all teams' changes go out together on a fixed schedule, whether they're ready or not — like a train that leaves at 5 PM no matter who's on it.
The monolith's problem is coupling: everything is stuck to everything else.
Step 3: The fix — break it into microservices
So Spotify did what many fast-growing companies do. They broke the giant kitchen into many small, independent food trucks.
What is a microservice?
A microservice is a small program that does one job and runs on its own. Instead of one big kitchen, you have a payments truck, a search truck, a login truck — each with its own stove, its own cook, its own schedule.
┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐
│ login │ │payments│ │ search │ │ recs │
│ service│ │ service│ │ service│ │ service│
└────────┘ └────────┘ └────────┘ └────────┘
each deploys, scales, and fails ON ITS OWN
Now each team can:
- Deploy independently — ship a fix without touching anyone else.
- Scale independently — give the search truck more power on a busy day without touching payments.
- Own their piece end to end — no waiting for the release train.
At Spotify this was organized into a squad model.
What is the squad model?
A squad is a small, self-sufficient team that fully owns one or more services — building them, running them, and fixing them. Autonomous means "makes its own decisions."
For a company hiring engineers very fast, this was genuinely the right call. It removed the coupling problem entirely.
Step 4: But autonomy compounds
Here's the subtle part. When every team is free to build its own thing, the number of things grows fast. Nobody planned it; it just piled up.
By 2020, when Spotify released a tool called Backstage, their system had grown to:
2,000+ backend services
300+ websites
4,000+ data pipelines
Nobody sat down and decided to build 2,000 services. Each squad made a sensible local choice. Added together, over years, you get an ocean of little programs.
And now a brand-new problem appears — one nobody saw coming.
Step 5: The new problem — you can't find anything
This failure mode is sneaky. It's not a crash. Nothing pages you at 3 AM. It's slower and quieter:
- Engineers burn hours in Slack asking "who owns this thing?"
- Teams build a duplicate service because nobody knew one already existed.
- A new hire is dropped into a system no single human can hold in their head.
Picture 2,000 food trucks scattered across a city with no map, no phone book, no signs. You need a "refund" truck. Is there one? Who runs it? Where is it parked? You have no idea — so you just build another refund truck. Now there are two.
The monolith's problem was coupling. The microservices problem is finding anything at all.
This is the big idea:
Microservices don't delete complexity. They trade deployment complexity for discovery complexity.
What is deployment complexity vs. discovery complexity?
- Deployment complexity = the pain of changing and shipping software (the monolith's problem).
- Discovery complexity = the pain of finding out what exists and who owns it (the microservices problem).
And here's why discovery complexity is so dangerous: it never pages you. A crash screams for attention. But discovery cost is silent — it just quietly taxes every engineer, every day, in lost hours.
Step 6: The fix — a software catalog
Spotify's answer started as an internal tool called System Z and grew into Backstage.
What is a software catalog?
A software catalog is a single searchable directory of every service in the company — like a phone book or a library card catalog, but for software. You type "refunds" and it tells you: yes it exists, here's who owns it, and here's whether it's safe to use.
The clever trick is how the catalog gets filled in. Every service must ship a small file describing itself:
# catalog-info.yaml
apiVersion: backstage.io/v1alpha1
kind: Component
metadata:
name: refund-service
spec:
owner: squad-payments
lifecycle: production
Let's read that file line by line, because this is the whole heart of the solution.
What is catalog-info.yaml?
It's a plain text file that lives inside each service's own code. YAML is just a simple, human-readable format for writing facts as key: value pairs.
kind: Component— says what this is (here, a Component means a piece of software).spec.owner: squad-payments— says which team owns it. This is the magic line.spec.lifecycle: production— says what stage it's in (production = live and real, versus experimental or deprecated).
Now the answer to "who owns this thing?" isn't stuck in one veteran engineer's memory. It's written down, right next to the code, in a form a computer can read.
The key shift: ownership becomes machine-readable metadata, not tribal memory.
What is tribal memory?
Tribal memory is knowledge that only lives in people's heads — "ask Priya, she knows." It vanishes when Priya goes on vacation or quits. Metadata (data about the data) written in a file never forgets and never leaves.
Here's how it fits together:
Each service: Backstage reads them all:
┌──────────────────┐
│ refund-service │──┐
│ catalog-info.yaml│ │ ┌────────────────────────┐
└──────────────────┘ │ │ SOFTWARE CATALOG │
┌──────────────────┐ ├──────▶│ search: "refund" │
│ search-service │ │ │ → refund-service │
│ catalog-info.yaml│ │ │ owner: squad-payments│
└──────────────────┘ │ │ lifecycle: production│
┌──────────────────┐ │ └────────────────────────┘
│ login-service │──┘
│ catalog-info.yaml│
└──────────────────┘
The payoff was real and measurable. Spotify measured new-hire onboarding by time to tenth pull request (how long until a new engineer has shipped 10 chunks of code — a sign they're truly productive). With the catalog, that time dropped by more than half.
Step 7: The honest catch — the fix has a cost
Here's the part most people leave out. The catalog wasn't free, and it wasn't purely technical.
To make it work, Spotify had to add:
- Mandatory metadata — every service must ship that
catalog-info.yaml. No exceptions. - Golden paths — recommended, standardized ways to build things.
- A standing platform org — a permanent team whose whole job is maintaining this shared system.
What is a golden path?
A golden path is a pre-approved, well-supported way to do a common task — "here's the blessed recipe for building a new service." It reduces chaos, but it also reduces freedom.
Do you see the tension? All of that is standardization — rules everyone must follow. But the original reason for microservices was autonomy — every squad doing its own thing. The fix pushes directly against the very freedom that justified the design.
They accepted it with eyes open: a fixed platform tax beats every engineer paying a discovery tax every single day.
What is a platform tax vs. a discovery tax?
- A platform tax is a known, fixed cost — the effort of maintaining standards and the catalog. You pay it once, up front, in a predictable way.
- A discovery tax is a hidden, endless cost — hours quietly lost by everyone, forever, hunting for things.
A fixed tax you can see and plan for beats an invisible tax that bleeds you slowly.
The core lessons
- Monoliths suffer from coupling — everything is tangled, so changing anything is hard.
- Microservices fix coupling by splitting software into small, independent pieces each owned by one team.
- But complexity is never deleted — only moved. Microservices trade deployment pain for discovery pain.
- Discovery pain is dangerous because it's silent. It never crashes; it just taxes everyone's time, every day.
- A software catalog (like Backstage) solves discovery by making every service declare itself in a machine-readable file —
owner,lifecycle, and more — so ownership stops being tribal memory. - The cure requires standardization, which pushes back against the autonomy that microservices promised. That's a real trade-off, not a free win.
- Choose your tax on purpose. A predictable, one-time platform tax beats an invisible, forever discovery tax.
