Skip to main content

Free tool · Educational demo

Fairness Audit Console

Accuracy is not the whole story in lending. A model can score well overall and still approve one group more often than another. This console audits a small logistic regression on the public UCI German Credit dataset at seven operating points: five decision thresholds and two Fairlearn mitigations. Every number was computed offline with a fixed seed. This page runs no model and takes no input.

Nothing leaves your browser. All data is embedded in this page.

The accuracy/fairness tradeoff

Seven operating points. Numbers 1-5 are threshold sweeps; 6-7 are Fairlearn mitigations. Click a point for the full breakdown.

0.600.650.700.750.850.900.951.001.05parity (1.0)Disparate-impact ratio (female / male approval rate)AccuracyThreshold 0.3: accuracy 0.7250, DI ratio 1.03701Threshold 0.4: accuracy 0.7450, DI ratio 1.03912Threshold 0.5: accuracy 0.7350, DI ratio 0.91743Threshold 0.6: accuracy 0.6950, DI ratio 0.91844Threshold 0.7: accuracy 0.6350, DI ratio 0.87505ThresholdOptimizer, demographic parity: accuracy 0.7500, DI ratio 1.00006ThresholdOptimizer, equalized odds: accuracy 0.7600, DI ratio 0.95837
Threshold sweepFairlearn ThresholdOptimizerDashed line: parity (1.0), where both groups are approved at the same rate.

Threshold sweep

Threshold 0.5

Accuracy

0.7350

whole test set

Disparate-impact ratio

0.9174

female / male approval rate

Equalized-odds difference

0.1300

driven by true-positive-rate gap

Approval rate, women

0.7667

n=60

Approval rate, men

0.8357

n=140

Women (n=60)

Good, approved (TP)
32
Good, denied (FN)
8
Bad, approved (FP)
14
Bad, denied (TN)
6
True-positive rate
0.8000
False-positive rate
0.7000
Group accuracy
0.6333

Men (n=140)

Good, approved (TP)
93
Good, denied (FN)
7
Bad, approved (FP)
24
Bad, denied (TN)
16
True-positive rate
0.9300
False-positive rate
0.6000
Group accuracy
0.7786

TPR gap 0.1300 · FPR gap 0.1000. Approval-rate and error-rate gaps persist even though sex is not a model input.

All operating points, test set (n=200)
Operating pointAccuracyDisparate-impact ratioEqualized-odds diffApproval: womenApproval: men
Threshold 0.30.72501.03700.12501.00000.9643
Threshold 0.40.74501.03910.15000.95000.9143
Threshold 0.50.73500.91740.13000.76670.8357
Threshold 0.60.69500.91840.17500.61670.6714
Threshold 0.70.63500.87500.13000.45000.5143
ThresholdOptimizer, demographic parity0.75001.00000.17500.80000.8000
ThresholdOptimizer, equalized odds0.76000.95830.07500.76670.8000

How this demo was built

Model and data

A logistic regression (C=1.0, standardized features, lbfgs) trained on the public UCI Statlog German Credit dataset: 1,000 rows, an 800/200 stratified split, random seed 42. The sensitive attribute is sex, derived from the personal_status codes (A91/A93/A94 map to male, A92/A95 to female, the standard mapping in the fairlearn literature). Sex was excluded from the features: a fairness-through-unawareness setup, so the gaps you see survive its removal.

Metrics and mitigations

Five decision thresholds (0.3 to 0.7) plus two Fairlearn ThresholdOptimizer runs under demographic-parity and equalized-odds constraints. Disparate-impact ratio is the female approval rate divided by the male approval rate; equalized-odds difference is the larger of the true-positive-rate and false-positive-rate gaps. The equalized-odds mitigation reaches an EO difference of 0.0750 at accuracy 0.7600, the best combined accuracy and equalized-odds result of the seven points.

Caveats, stated plainly. Only 60 women are in the test set, so per-group rates carry noise. The fairness signal is mild: at the base threshold of 0.5 the disparate-impact ratio is 0.9174, above the 0.8 rule-of-thumb flag. The four-fifths rule is a U.S. employment screening rule of thumb, not a credit-lending legal test and not Canadian law.

Built with

Offline: scikit-learn, Fairlearn, AIF360, and the UCI Statlog German Credit dataset. This page uses static rendering with embedded data. No backend, no visitor data. Supporting code:movahedi-ca/fairness-audit-demo.

Common questions

What does the disparate-impact ratio measure?

It divides the approval rate of the unprivileged group by the approval rate of the privileged group. A ratio of 1.0 means both groups are approved at the same rate. The U.S. EEOC four-fifths rule of thumb flags ratios below 0.8, but that is a screening rule of thumb from U.S. employment guidance, not a credit-lending test and not Canadian law. Treat it as a signal that deserves a closer look, never as a verdict.

What is equalized odds?

Equalized odds asks for similar true-positive rates and similar false-positive rates across groups: qualified applicants should be approved at similar rates, and unqualified applicants should be approved at similar rates too. The equalized-odds difference is the larger of the two gaps, so 0 is perfect parity. It is stricter than demographic parity because it accounts for who actually repays.

Sex was not a model input. Why do gaps remain?

Because fairness through unawareness does not guarantee fairness. Other features correlate with group membership, so removing the sensitive attribute from the inputs does not remove the pattern from the predictions. This demo excluded sex from the features on purpose, to show exactly that: the gaps at the base operating point survive the removal.

Why is the female test group only 60 rows?

That is what the stratified 80/20 split of this dataset yields: 140 men and 60 women in the test set. Rates computed on 60 rows carry noise, especially true-positive and false-positive rates on the smaller denied subgroups. The console states the sample sizes wherever the numbers appear so you read them with the right caution.

Can I use this console for a real fairness review?

No. This is an educational demo, not a compliance review and not legal advice. A real audit needs larger samples, the protected attributes that matter in the jurisdiction, and qualified review. Nothing on this page runs a model live or takes any input: every number was computed offline with a fixed seed.