Step 1: A model that works perfectly — and is still doing something wrong
Let's start with a story that has played out in real companies for decades.
A bank builds a credit model. Its job: predict whether someone will repay a loan. The team feeds it historical data — income, repayment history, employment, address — and trains it.
The model performs beautifully. Accuracy is high. Profit goes up. The metrics dashboard is all green.
Then someone looks inside and finds this: one of the strongest signals the model learned is postcode. Where you live.
INPUTS MODEL OUTPUT
───────── ┌─────────┐
Income ────────▶ │ │
Repayment history ────────▶ │ Credit │ ────────▶ "Approve"
Employment ────────▶ │ Model │ or
Postcode ────────▶ │ │ "Deny"
└─────────┘
▲
└── quietly one of the strongest signals
Here's the uncomfortable truth at the heart of this lesson:
A model can be accurate, profitable, and still be doing something indefensible.
Why? Because accuracy only measures whether predictions match past outcomes. It says nothing about whether the model's reasons are legitimate. And postcode, in many countries, is a near-perfect stand-in for race, ethnicity, and class — because neighborhoods are historically segregated.
This pattern has a name. It's called redlining — and the model just reinvented it.
What is redlining?
Redlining is the old practice of banks drawing red lines around certain neighborhoods on a map and refusing loans to anyone inside them — regardless of that person's actual finances. It's discrimination by address. When a model learns postcode as a strong credit signal, it is redlining, just dressed up as "pattern recognition."
Step 2: The trap — the model didn't do anything "against the rules"
Here's what makes this hard. Nobody told the model to discriminate. Nobody put "race" or "ethnicity" in the training data. The team may have deliberately removed those columns.
It doesn't matter. The model found postcode, and postcode carries the same information.
This is called a proxy.
What is a proxy?
A proxy is a feature that stands in for another attribute, because the two are correlated. Example: you delete "age" from your data, but keep "year of graduation." The model can still figure out roughly how old everyone is. Graduation year is a proxy for age.
Think of it like a restaurant with a "no children" policy that instead bans "anyone under 140cm tall." They never said "children." But that's exactly who they're excluding. The height rule is a proxy.
So here's the problem no fairness dashboard actually solves:
**For any given decision, which differences between people is a model allowed to use?**
Let's line some up for a loan decision:
Feature Allowed to use it?
────────────────── ──────────────────
Income Probably yes
Repayment history Yes — clearly relevant
Postcode ...?
Name ...??
The shape of a surname ...???
Income and repayment history feel legitimate — they're directly about the person's ability and track record of paying money back. But somewhere down that list, a "legitimate signal" turns into a proxy smuggling in something you would never write into your objective function on purpose.
What is an objective function?
The objective function is the goal you give the model during training — the thing it tries to optimize. For a credit model: "predict loan repayment as accurately as possible." Nobody writes "and penalize people from immigrant neighborhoods" into that goal. But if postcode helps accuracy, the model will use it — and achieve that penalty anyway, as a side effect.
Step 3: An ancient legal test for exactly this problem
Here's the surprising part: this exact question — which attribute is really doing the work? — was studied intensely, for centuries, long before computers.
Classical Muslim jurists faced a version of it constantly. Their legal tradition allowed extending an existing ruling to a new case by analogy. But an analogy is only valid if you extend it on the right attribute. So they developed a discipline for finding what they called the ʿilla.
What is the ʿilla?
The ʿilla (pronounced roughly "ILL-ah") is the operative cause of a ruling — the specific attribute that actually drives the ruling, as opposed to attributes that are merely present and correlated. Classic example: wine is prohibited. Why? Not because it's a liquid, not because it's made from grapes — because it intoxicates. Intoxication is the ʿilla. So the ruling extends to other intoxicants (even non-grape ones), and does not extend to grape juice.
The jurists' rule was strict:
Extend a ruling on an incidental attribute — one that's merely correlated, along for the ride — and the whole inference is invalid. Full stop.
Roman jurists chased the very same distinction under the name ratio legis — "the reason of the law." Two independent legal traditions, same insight: before you generalize from a case, you must isolate why the case has its outcome.
Now watch what happens when we map this onto machine learning:
JURIST'S QUESTION ML ENGINEER'S QUESTION
───────────────────────────── ─────────────────────────────
Ruling on a known case ↔ Outcome in the training data
Extending it to a new case ↔ Making a prediction
The ʿilla (operative cause) ↔ A legitimately causal feature
An incidental attribute ↔ A proxy feature
Invalid analogy ↔ Invalid inference
A model making a prediction is an analogy machine. It says: "this new applicant resembles past applicants who defaulted, therefore deny." The question is: resembles them in what? In repayment behavior — or in address?
Step 4: "Postcode does not repay a loan. People do."
That sentence is the whole test compressed into eight words. Let's unpack it slowly.
Ask, for each feature: is this the operative cause of the outcome, or a stand-in for something morally irrelevant?
- Repayment history → Does a person's track record of repaying debts cause (in a real, mechanical sense) their likelihood of repaying the next one? Yes. Habits, financial discipline, existing obligations — there's a genuine causal story. This is the ʿilla. Operative.
- Postcode → Does living at a certain address cause loan repayment? No. A postcode has never repaid a loan. People repay loans. Postcode is correlated with repayment only because it's correlated with other things — wealth, historical discrimination, local job markets. Incidental.
Here's the picture:
Repayment history ●━━━━━━━━━━━━━━━━━━━━▶ ◎ OUTCOME
(operative cause) solid causal loan repaid
connection
▲
╱
┈┈┈⊗┈┈┈ ◀── blocked: this link
╱ should never be drawn
Postcode ●┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈
(incidental — a stand-in)
And the crucial punchline:
When the model leans on postcode, you have extended a ruling on the wrong attribute — and no amount of AUC makes that valid.
What is AUC?
AUC (Area Under the Curve) is a standard score for how well a model separates good outcomes from bad ones, from 0.5 (coin flip) to 1.0 (perfect). It's a favorite metric in ML. The point here: AUC measures predictive power, not legitimacy of reasoning. A model can have a gorgeous AUC precisely because it exploits a proxy. The metric will happily reward the redlining.
This reframes something we usually treat as a purely technical chore:
Feature selection is a moral discipline. Choosing which columns your model sees is choosing which differences between people you consider fair grounds for treating them differently.
Step 5: Who has to prove what? The second principle
The jurists had a second commitment that gives the first one teeth. They called it ʿiṣma.
What is ʿiṣma?
ʿiṣma (roughly "ISS-mah") means the equal inviolability of persons. Every person starts with the same protected standing — the same default right not to be harmed or treated as lesser. It's a starting point, not something you earn.
Why does this matter for models? Because it settles the question of burden of proof.
Without ʿiṣma, the default is what most companies do today:
Model draws a distinction ──▶ ships to production
│
harmed person notices (maybe, someday)
│
harmed person must PROVE they were wronged
With ʿiṣma, the arrow flips:
Model wants to draw a distinction
│
▼
YOU must justify that distinction ──▶ only then does it ship
The burden falls on the builder to justify every distinction the model draws — not on the affected person to prove they were wronged after the fact.
Everyday analogy: a doctor who wants to give you a drug must justify it before prescribing. We don't say "take everything, and if you're harmed, sue us and prove it." Consequential automated decisions deserve the same direction of burden.
Step 6: What this changes in practice — audit the cause, not just the output
Most fairness work today checks outcome disparity: after the fact, do approval rates differ across groups? That's worth doing, but it's checking the output. This discipline checks the reasoning.
Here's the concrete process. For every high-weight feature, force an answer to one question — operative, or proxy? — and log the justification.
What is a high-weight feature?
A high-weight feature is one the model relies on heavily. You can measure this with feature-importance tools (e.g., SHAP values, permutation importance). If removing the feature barely changes predictions, its weight is low. If predictions swing on it, it's high-weight — and it needs a justification.
A minimal version of the audit, in code:
# Step 1: Find what the model actually leans on
importances = compute_feature_importance(model, X_val) # e.g. SHAP
# Step 2: Force a ruling on each heavy feature
for feature, weight in importances.top(k=10):
print(f"{feature}: weight={weight:.3f}")
print(" Q: Is this the OPERATIVE CAUSE of the outcome,")
print(" or a STAND-IN for something morally irrelevant?")
And a justification log — reviewable, written down, owned by someone:
# feature_justifications.yaml
- feature: repayment_history
weight: 0.41
ruling: operative
justification: >
Direct causal mechanism: past repayment behavior reflects
financial habits and obligations that drive future repayment.
reviewed_by: risk-committee
date: 2025-03-14
- feature: postcode
weight: 0.29
ruling: proxy
justification: >
No causal path from address to repayment. Correlation runs
through wealth segregation and historical discrimination.
Proxy for protected attributes.
action: REMOVED — does not ship
The shipping rule is simple and strict:
If you cannot name why a feature is the operative cause, it does not ship.
Notice what this rule inherits from the jurists. Not "remove it if someone complains." Not "keep it if the metrics improve." The default is exclusion, and inclusion must be earned by justification — because every person starts with equal protected standing, and every distinction your model draws is something you owe an account for.
One honest caveat: this test is simple to ask and hard to answer. Some features sit in a gray zone (is "employment length" operative, or partly a proxy for age?). That's fine. The discipline doesn't promise easy answers. It promises that the question gets asked, answered in writing, and reviewed — instead of never asked at all.
Step 7: The one-line summary of the whole idea
Most guardrails: inputs ──▶ [ model ] ──▶ outputs ──▶ ✓ check here
This discipline: inputs ──▶ [ model ] ──▶ outputs
▲
└── ✓ check HERE: the reasoning
Most guardrails check the output. This checks the reasoning.
And a question worth sitting with, for any model you're responsible for right now:
What feature is your model using that you have never actually justified?
The core lessons
- Accuracy is not innocence. A model can be accurate and profitable while being indefensible, because metrics like AUC measure predictive power — not whether the model's reasons are legitimate.
- Proxies smuggle in what you removed. Deleting protected attributes (race, ethnicity) doesn't help if correlated features (postcode, name, surname shape) carry the same information. That's how a credit model reinvents redlining.
- Ask the jurists' question of every feature. Classical Muslim jurists (the ʿilla) and Roman jurists (ratio legis) developed the same test: is this attribute the operative cause of the outcome, or merely incidental — correlated, along for the ride? An inference built on an incidental attribute is invalid, full stop.
- Postcode does not repay a loan. People do. Repayment history has a real causal path to repayment; postcode does not. When a model leans on postcode, it has extended a ruling on the wrong attribute.
- The burden of proof is on the builder. The principle of ʿiṣma — the equal inviolability of persons — means you must justify every distinction the model draws before shipping, rather than making affected people prove harm afterward.
- Make it operational. For every high-weight feature: force a ruling — operative or proxy? Log the justification. Make it reviewable. If you can't name why a feature is the operative cause, it does not ship.
- Check the reasoning, not just the output. Outcome-disparity dashboards audit what the model decided. This discipline audits why — and that's where the real failure lives.
