All resources
Daily Learning · Reflection Series

How surveillance systems get built by accident — and how to design so they can't be

Sunday, 26 July 2026
How surveillance systems get built by accident — and how to design so they can't be

🎯 You'll understand how ordinary engineering habits (logging everything, optimizing engagement, giving AI agents broad access) quietly assemble a surveillance system, and the four concrete design changes that prevent it.

Step 1: Nobody builds a surveillance system on purpose

Here's the uncomfortable truth at the heart of this lesson:

The most dangerous thing about AI surveillance is that nobody decided to build it. It emerged, one convenient logging pipeline at a time.

No villain wrote a design doc titled "Track Everyone." Instead, a thousand small, reasonable-sounding decisions stacked up:

  • "Let's log every event — storage is cheap."
  • "Let's keep every user prompt — we might need it for training later."
  • "Let's store every embedding — deleting is more work than keeping."

Each decision is defensible on its own. Together, they produce a system that can read, infer, and profile everyone.

What is an embedding?

An embedding is a list of numbers that represents the meaning of something (a sentence, an image, a user's behavior). Example: the sentences "I'm expecting a baby" and "I'm pregnant" get very similar number lists, so a computer can tell they mean the same thing — even if the user never used the word "pregnant."

That last point matters. Embeddings mean your stored data can reveal things the user never explicitly said. Keep that in mind — we'll come back to it.


Step 2: A fourteen-century-old design principle

This isn't a new problem, and one of the sharpest answers to it is very old.

Classical Muslim jurists prohibited something called tajassus — probing into people's private matters. It was part of a broader duty called hifz al-'ird: protecting human dignity.

What is tajassus?

Tajassus = spying or prying into someone's private life without justification. Analogy: your neighbor keeping their curtains closed is not an invitation to bring a ladder.

Here's the striking part — how far the jurists took it:

Evidence obtained by unwarranted intrusion into private life was refused even when it uncovered real wrongdoing.

Read that again. Even if the spying worked — even if it caught a genuine crime — the evidence was thrown out. Why would anyone design a rule that lets wrongdoers off the hook?

The reasoning was structural, not sentimental: normalising surveillance corrodes the whole society more than any single catch is worth. If everyone knows they might be watched, everyone changes how they live. That cost is permanent and society-wide. Catching one wrongdoer is a one-time gain. The trade is bad.

Roman law and English common law circled the same problem from different angles. This is not a niche religious concern — it's a pattern that multiple legal traditions independently discovered.

Why this matters for engineers

Because it gives us a rule with actual teeth:

Don't collect what you have no right to see, even when collecting is trivial and the findings are useful. "We could" has never implied "we may."

Fourteen centuries ago, that was jurisprudence. Today, it should be a design review checklist item. Let's turn it into one — piece by piece.


Step 3: Failure — The schema that collects too much

Every application has a data model: the set of fields you record about each event or user.

What is a schema?

A schema is the blueprint for what data you store — the columns of your database tables. Example: a login_event table with columns user_id, timestamp, ip_address.

Now here's the trap. Compare these two event schemas for a "user clicked play on a video" event:

-- Minimal schema: only what the feature needs
CREATE TABLE play_events (
    user_id     UUID,
    video_id    UUID,
    played_at   TIMESTAMP
);

-- "Just in case" schema: what teams actually ship
CREATE TABLE play_events (
    user_id       UUID,
    video_id      UUID,
    played_at     TIMESTAMP,
    ip_address    TEXT,        -- "for fraud, maybe"
    device_model  TEXT,        -- "might be useful"
    geo_lat       FLOAT,       -- "was in the SDK anyway"
    geo_lon       FLOAT,
    battery_level INT,         -- "came free with the payload"
    wifi_ssid     TEXT         -- "why not"
);

The second table can reconstruct where someone lives, works, sleeps, and who they're near — from a video play button. And here's the key insight:

If your event model captures fields no feature requires, you have built a surveillance system that merely hasn't been queried yet.

The surveillance isn't in the query. It's in the collection. The data sits there like a loaded question, waiting for someone — an employee, an acquirer, an attacker, a subpoena — to ask it.

The fix: data minimisation at the schema level. Every field must justify its existence by pointing at a real feature that needs it. No feature, no field.

Design review question for every column:
┌─────────────────────────────────────────────┐
│  "Which feature breaks if we delete this?"  │
│                                             │
│  Answer names a feature  → keep it          │
│  Answer is "might need it later" → delete   │
└─────────────────────────────────────────────┘

Restaurant analogy: a waiter needs your order and your table number. A waiter who also writes down your car's license plate, who you arrived with, and what you whispered — "in case it's useful later" — isn't a waiter anymore. He's an informant. The notebook makes him one, even if nobody ever reads it.


Step 4: Failure — Objective functions that reward inference

This one is subtler, because no human collects anything. The model does the prying.

What is an objective function?

An objective function is the score a machine learning model is trained to maximize. Example: a recommender system might be trained to maximize "probability the user clicks." The model will learn anything that helps it raise that score.

Here's the failure mode:

Objective: maximize clicks
                │
                ▼
Model discovers: pregnant users click
   baby content at high rates
                │
                ▼
Model learns internal signal:
   "this user is probably pregnant"
                │
                ▼
User never said it. Nobody asked for it.
   The model infers it anyway —
   because inference was PROFITABLE
   and never made COSTLY.

A recommender optimized purely on engagement will learn to infer pregnancy, illness, grief — whatever predicts clicks. This is one of the oldest documented failure patterns in applied ML. The model profiles people because profiling was never made costly.

Notice what's happening: nobody wrote code that says "detect pregnancy." The objective function rewarded the inference, so gradient descent found it. The surveillance is an emergent property of the incentive.

What is an emergent property?

An emergent property is behavior that arises from a system without anyone designing it directly. Example: no single driver decides to create a traffic jam — the jam emerges from many drivers each making small decisions.

The fix: make sensitive inference costly in the objective, or structurally impossible in the features. If clicks are the only thing you reward, don't be surprised when the model becomes an expert at knowing things people never told it.


Step 5: Failure — Agents with broad read scopes

Now the newest version of the problem: LLM agents.

What is an LLM agent?

An LLM agent is an AI language model that's been given tools and permissions to act — read documents, query databases, send messages — not just chat. Example: an assistant that can search your company's files to answer questions.

The common shortcut looks like this:

# Convenient (and dangerous)
agent: helpdesk-assistant
permissions:
  documents: read: org-wide    # "for context"
  email:     read: all-users
  hr-system: read: all-records

Why do teams do this? Because scoping is work, and broad access makes the agent seem smarter. But look at what you've built:

An LLM agent given org-wide document access "for context" is tajassus as a service.

Every question anyone asks the agent becomes a potential probe into everyone else's private material. The agent doesn't need bad intent. The scope is the intrusion.

The fix: scope to the task, not the tenant.

# Scoped (the discipline that protects everyone)
agent: helpdesk-assistant
permissions:
  documents:
    read:
      - folder: /helpdesk/kb/          # only the knowledge base
      - requester_owned: true          # only the asking user's files
  email:     none
  hr-system: none

What does "task, not tenant" mean?

A tenant is a whole organization's space in a system (all users, all data). Scoping "to the tenant" means the agent can see everything in the org. Scoping "to the task" means it can see only what this specific request requires. Analogy: a locksmith you hire to open your front door should get your front-door key — not the master key to the whole apartment building.

Tenant scope:                Task scope:
┌───────────────┐            ┌───────────────┐
│ ███████████████│            │ ░░░░░░░░░░░░░ │
│ ███ AGENT  ████│  vs.       │ ░░┌──────┐░░░ │
│ ███ SEES   ████│            │ ░░│AGENT │░░░ │
│ ███ ALL    ████│            │ ░░│SEES  │░░░ │
│ ███████████████│            │ ░░└──────┘░░░ │
└───────────────┘            └───────────────┘

Step 6: Failure — Evals that never test for it

Last piece. AI teams run evals before shipping.

What is an eval?

An eval (evaluation) is an automated test suite for AI behavior. Example: feed the model 1,000 prompts trying to make it say something toxic, and measure how often it fails.

Today, teams routinely eval for:

  • ✅ Toxicity (does it say harmful things?)
  • ✅ Jailbreaks (can users trick it past its rules?)

But almost nobody evals for this:

  • ❌ Does the system infer attributes users never disclosed?

That's a huge blind spot, because — as we saw in Step 4 — undisclosed inference is one of the oldest failure modes in ML. What could such an eval look like?

# Sketch: a privacy-inference eval
test_user = build_profile(
    disclosed = ["likes cooking videos", "based in London"],
    undisclosed = ["pregnant", "recently bereaved"],
)

response = system.recommend(test_user)

# The check: does the system's behavior reveal
# it has inferred what was never disclosed?
assert not shows_targeting(response, attribute="pregnancy")
assert not shows_targeting(response, attribute="grief")

The principle: if you don't test for it, you've silently decided it's acceptable. Your eval suite is a statement of what you care about. Toxicity and jailbreaks made the list. Uninvited inference should too.


Step 7: Putting it together — the design review checklist

Let's assemble the four failures and four fixes into one picture:

WHERE SURVEILLANCE SNEAKS IN         THE FIX
─────────────────────────────       ─────────────────────────────
1. Schema collects "just in    →    Data minimisation: every
   case" fields                     field maps to a real feature

2. Objective rewards clicks,   →    Make sensitive inference
   model learns to profile          costly, not free

3. Agent has org-wide read    →    Scope to the task,
   access "for context"             not the tenant

4. Evals test toxicity but    →    Eval for inference of
   never inference                  undisclosed attributes

And over all four, the old rule from Step 2:

"We could" has never implied "we may." Cheap collection is not consent. It never was.

Notice how the classical logic and the engineering logic are the same logic. The jurists refused ill-gotten evidence even when it was useful, because normalized surveillance costs more than any single find is worth. The engineer refuses the extra column even when storage is cheap, because the collected-but-unqueried data is a surveillance system waiting for its first query.

A practical exercise to end on — one worth actually doing this week:

Open your current data model and find one field that exists only because it was easy to capture. Then delete it, or write down the feature that justifies it. If you can't do either, you've found your surveillance system.

The core lessons

  • Surveillance emerges; nobody has to decide to build it. It accumulates one convenient logging pipeline, one "just in case" field, one broad permission at a time.
  • Collection is the harm, not just the query. Data your features don't need is a surveillance system that merely hasn't been queried yet. Minimise at the schema level.
  • Models infer whatever the objective rewards. An engagement-optimized recommender will learn pregnancy, illness, and grief unless profiling is made costly. This is one of the oldest failure patterns in applied ML.
  • Scope agents to the task, not the tenant. Org-wide read access "for context" is prying built into the architecture — tajassus as a service.
  • Eval for undisclosed inference, not just toxicity and jailbreaks. What you don't test, you've silently approved.
  • "We could" has never implied "we may." Fourteen centuries of legal thinking agree: cheap collection is not consent, and privacy is an architecture choice — not an afterthought.

Want this kind of thinking applied to your business?

Book a free 30-minute discovery call — we’ll show you your highest-value first automation, no jargon, no obligation.

Book a discovery call