All resources
Daily Learning · Release It!: Design and Deploy Production-Ready Software

Why a Green Dashboard Can Lie: From Averages to Observability

Friday, 10 July 2026
Why a Green Dashboard Can Lie: From Averages to Observability

🎯 You'll understand how classic metrics like averages and p90s are built, why they can show "healthy" during a real outage, and how per-request, high-cardinality observability fixes that — at a cost you accept on purpose.

Step 1: The world where this problem was born

In 2013, engineers at Etsy (the online marketplace) were shipping code to production up to 50 times a day.

That sounds terrifying. Fifty chances a day to break the site. But it wasn't terrifying — and the reason was deliberately boring:

They measured everything. Their philosophy: if it moves, graph it. If it matters, alert on it.

Every deploy was followed by a glance at a graph. If a line jumped, you rolled back. Measurement was the safety net that made speed possible.

What is "production"?

Production is the live version of your software — the one real customers use. A bug on your laptop is annoying. A bug in production costs money.

To make measurement effortless, Etsy built and open-sourced a small tool called StatsD. It let any engineer record a measurement with one line of code:

statsd.increment('checkout.success')

That's it. A developer adds that line, deploys, and a graph appears in the dashboard within minutes. No ticket. No asking the ops team for permission. No meeting.

What is StatsD?

StatsD is a tiny background program (a daemon — a program that runs continuously, like a doorman who never goes home). Your application throws quick little messages at it — "a checkout succeeded", "this request took 240 milliseconds" — and StatsD collects them and forwards summaries to a graphing system called Graphite.

The genius was making instrumentation so cheap that nobody thought twice about adding it. And for a while, this was the gold standard: thousands of dashboards, deploy, watch the graph, react.

Then the graphs got too good. To see why, we need to look at what StatsD actually does with those measurements.


Step 2: What aggregation really means

Here's the key mechanical detail. StatsD does not keep every individual measurement. By default, it collects measurements for 10 seconds, squashes them into a summary, ships the summary, and throws the raw data away.

Your app (10 seconds of traffic):
  request 1: 40ms
  request 2: 55ms
  request 3: 38ms
  ...
  request 9,999: 51ms
  request 10,000: 12,000ms   ← one seller's request, painfully slow

          │
          ▼  StatsD "flush" every 10s
          
Shipped to Graphite (the summary):
  count: 10,000
  mean:  ~52ms
  p90:   ~60ms
          
          ▼
  Raw data: DELETED

What is aggregation?

Aggregation means replacing many individual numbers with a few summary numbers — like a teacher reporting "the class average was 78%" instead of listing all 30 students' scores. It's compact and cheap. It's also how information disappears.

What is a percentile (like p90)?

Line up all 10,000 request times from fastest to slowest. The p90 (90th percentile) is the time at the 90% mark — 90% of requests were faster than it. Percentiles were the industry's upgrade over plain averages, because averages get dragged around by outliers. p90 or p99 tells you about the typical worst experience.

This aggregation is exactly what made StatsD so cheap. Ten thousand requests become a handful of numbers. Storage costs almost nothing. Graphs draw instantly.

And that same aggregation is exactly what hides the pathological request.


Step 3: How a healthy dashboard lies during an incident

Here's the failure mode most explanations skip over. Walk through it slowly.

Imagine your marketplace serves 10,000 requests every 10 seconds. Now one specific seller — say, a big shop in one specific region — starts hitting a slow database query path. Their pages take 12 seconds instead of 50 milliseconds. For them, the site is down.

What does the dashboard show?

p90 Latency                              ● Healthy
────────────────────────────────────────────────
 ~~~~~~~~~~~~~~~~~ flat green line ~~~~~~~~~~~~~
────────────────────────────────────────────────
 00:00     06:00     12:00     18:00     24:00

Flat. Green. Healthy. Why?

Do the math. That seller's traffic is maybe 10 requests out of 10,000 — 0.1% of all requests:

  • The mean absorbs it. 9,990 fast requests plus 10 slow ones barely moves the average. (9,990 × 50ms + 10 × 12,000ms) ÷ 10,000 ≈ 62ms. Looks fine.
  • The p90 hides it. The p90 only looks at the 90% mark in the sorted line. The disaster lives in the last 0.1% — way past the p90's line of sight. The p90 doesn't budge.
The mean absorbs it. The tail hides it. You're staring at a healthy dashboard during an active incident.

What is "the tail"?

The tail is the slowest sliver of your requests — the far right end of the sorted lineup: the slowest 1%, 0.1%, 0.01%. Real outages often live only in the tail, because failures usually hit one customer, one region, one code path — not everyone at once.

A restaurant analogy. Your restaurant serves 1,000 meals a night. Average wait time on the report: 12 minutes. Wonderful. Meanwhile, every party seated in the back corner waits two hours, because that one section's kitchen printer is jammed. The average never shows it. Those customers are furious, leaving one-star reviews — and your nightly report says everything is fine.

And here's the truly cruel part: even if some alarm eventually fires, your data can't answer the next question. You'd want to ask: "Which requests were slow? What did they have in common?" But StatsD threw the raw requests away 10 seconds after they happened. All you have is p90 = 60ms. The evidence was destroyed at write time.


Step 4: The fix — stop aggregating at write time

The industry's answer was not fancier averages or more percentiles. It was a change in when you summarize.

Old way (metrics): summarize at write time. Store only the summary. Cheap, but the detail is gone forever.

New way (observability): store the raw event for every request, with lots of detail attached. Summarize later, at query time, only when you ask a question.

Facebook built an internal system called Scuba to do exactly this: query raw, "wide" events at interactive speed. Later, Charity Majors and the team at Honeycomb turned this from an internal trick into a public discipline with a name: observability.

What is a wide event?

A wide event is one record per request, carrying many descriptive fields:

{
  "timestamp":   "2024-03-01T14:22:07Z",
  "endpoint":    "/checkout",
  "seller_id":   "seller_84291",
  "region":      "eu-west-1",
  "db_query":    "get_listing_variants",
  "duration_ms": 12043,
  "status":      200
}

One of these per request. Nothing thrown away.

What is high cardinality?

Cardinality = how many distinct values a field can have. region might have 20 values — low cardinality. seller_id might have 5 million values — high cardinality. Old metrics systems choke on high-cardinality fields (a separate counter per seller would mean millions of graphs). Event stores like Scuba and Honeycomb are built to handle them.

Now compare what each approach lets you ask:

METRICS (pre-aggregated)          EVENTS (raw, wide)
─────────────────────────         ─────────────────────────────────
"What's the average               "Show me the slowest 0.1% of
 checkout latency?"                requests, and what they have
                                   in common."
Answer: 62ms. Healthy!            Answer: all 10 slow requests are
                                  seller_84291, eu-west-1, and all
(dead end — no detail left)       hit get_listing_variants.
                                  Root cause in one query.

That second question — "show me the slowest 0.1% and why" — is the entire point. You slice the raw events by any dimension: seller, region, endpoint, query. The needle stops being invisible, because you never melted it into the haystack.

Library analogy. Metrics are like a library that, every night, burns the day's returned books and keeps only a note: "412 books returned, average length 280 pages." Observability keeps every book on the shelf, so tomorrow you can ask, "Which books over 900 pages were returned late, and by whom?" You can't ask that of a burn pile.


Step 5: The bill — and why you pay it anyway

Here's the twist that rarely makes it onto the conference slide: high-cardinality observability is expensive. Deliberately so.

Compare the storage math:

Aggregated counter:
  10,000 requests / 10s  →  ~5 summary numbers
  
Wide events:
  10,000 requests / 10s  →  10,000 full records,
  each with seller_id, region, endpoint, duration...

That's roughly a thousand-fold explosion in what you write and store, and queries over raw events cost far more compute than reading a pre-rolled line off a graph. There's no clever trick that makes this free. It's a real bill.

So why pay it?

Because a cheap metric that lies is worse than an expensive one that tells the truth.

Think about what the cheap dashboard actually cost during the incident: hours of engineers staring at green graphs while a real customer was down, unable even to ask the question that mattered. The expensive event store answers "who is slow and why" in one query. You accept the storage bill with your eyes open, as the price of being able to see.

In practice, teams do both: cheap aggregated metrics for broad health and alerting, raw events for the investigation — knowing which tool can answer which question.

And it's worth sitting with the one-line summary of the whole trap:

An aggregate is an average of the things you needed to see individually.

The core lessons

  • Measurement enables speed. Etsy deployed 50 times a day in 2013 because instrumentation was one line of code (statsd.increment(...)) and a graph appeared minutes later.
  • Aggregation is a trade. StatsD flushes every 10 seconds and keeps only summaries (means, p90s). That's what makes it cheap — and what destroys the evidence.
  • Averages and percentiles hide the tail. One seller's 12-second requests are 0.1% of traffic; the mean absorbs them, the p90 never sees them. A dashboard can be green during an active incident.
  • Observability = don't aggregate at write time. Store one wide event per request with high-cardinality fields (seller_id, region, endpoint), then summarize at query time. Ask "show me the slowest 0.1% and why," not "what's the average."
  • It costs real money, on purpose. Per-request events explode storage and query cost versus one rolled-up counter. You pay it because a cheap metric that lies is worse than an expensive one that tells the truth.
  • Remember the trap in one line: an aggregate is an average of the things you needed to see individually.

Want this kind of thinking applied to your business?

Book a free 30-minute discovery call — we’ll show you your highest-value first automation, no jargon, no obligation.

Book a discovery call