How do you prove an underwriting model is fair when you do not hold the protected attribute?
Audit what the target variable stands in for first, because that is where bias usually enters. Then report decision rates and performance across the segments you can observe. Evidence the deployed configuration, re-tested each release. Document which groups you could not examine and which proxies you relied on.
In insurance and lending this is the standard situation: you are asked to demonstrate that an automated decision does not discriminate, using data that deliberately excludes the characteristic in question. Removing the attribute does not remove the risk — it removes your ability to see it.
Start with the target, not the features
The most consequential bias usually enters through what the model is trained to predict, not through its inputs. The clearest published demonstration is from Science in 2019: an algorithm used on large patient populations predicted healthcare costs as a stand-in for healthcare need. Because less was historically spent on Black patients at the same level of illness, patients at an identical risk score were considerably sicker. The authors calculated that correcting the target would raise the share of Black patients identified for additional help from 17.7% to 46.5%.
No protected attribute was in the model. The proxy did the work. So the first audit question is not “which features are risky” but “what does the thing we predict actually stand in for, and does that substitution hold equally for everyone?”
Then test the segments you can see
You can almost always construct meaningful segments without the protected attribute: geography, product, channel, tenure, language of correspondence, deprivation indices, age bands where lawfully held. Report performance and decision rates per segment rather than in aggregate. This is the same discipline that finds hidden failure generally — ACM CHIL research in 2020 found relative performance differences of over 20% on important subsets inside models whose overall numbers looked strong.
Where a segment shows a materially different outcome, the finding is not a verdict. It is a question that goes to the people who own the product, with the data behind it.
Check that equal scores mean equal behaviour
Two models can score identically on a held-out set and behave differently in deployment. Work published in the Journal of Machine Learning Research in 2022 showed that standard pipelines routinely produce such models. Fairness evidence therefore has to come from the deployed configuration, re-tested on each release, not from the version that happened to be evaluated once.
Write down what you could not measure
The strongest evidence pack states its own limits: which groups could not be examined, which proxies were used, what sampling the conclusions rest on. A reviewer trusts a document that marks its boundaries more than one that implies complete coverage.
Regulatory frameworks expect this shape of evidence. The EU AI Act requires, for high-risk systems, a documented risk management system (Article 9) and examination of training, validation and test data for possible biases (Article 10). Whether and how these apply to your product is a question for your counsel; this page describes the engineering and evidence work.
Related questions
If the model never sees the protected attribute, can it still discriminate?
Yes. A 2019 Science study found an algorithm predicting healthcare cost as a stand-in for need produced markedly unequal outcomes with no protected attribute in the model; correcting the target raised the share of Black patients identified for extra help from 17.7% to 46.5%.
What can we test without the attribute?
Observable segments: geography, product, channel, tenure, correspondence language, deprivation indices, and lawfully held age bands. Report decision rates and performance per segment rather than in aggregate.
Is one fairness assessment enough?
No. JMLR research in 2022 showed pipelines producing models that score identically in testing yet behave differently once deployed, so evidence must come from the deployed configuration and be re-tested on each release.
AI Testing & Certification
Test AI solutions for accuracy, safety and conformity, and prepare them for external certification.
Discuss your requirement