A vendor says their AI is 95% accurate. What should we check before we believe it?
Ask five things: which population it was measured on, whether anyone outside the developer has validated it, how it performs per subgroup and site, what shortcuts it may be keying on, and whether you can run it on your own data against criteria you set first. A refusal on the last point is itself an answer.
A single accuracy figure is the least informative number a vendor can give you, and the published record shows why. Five questions turn it into something you can act on.
1. Accurate on whom, and where?
Performance reported by the developer often does not survive contact with another organisation. In an external validation published in JAMA Internal Medicine in 2021, a widely deployed proprietary sepsis prediction model scored an area under the curve of 0.63 at an independent hospital system — against the 0.76 to 0.83 its developer had reported — and at the alert threshold in use it failed to identify 67% of patients who developed sepsis. Same model, different hospital, different answer.
Ask: on which population was this figure measured, and has it been validated anywhere the vendor does not control?
2. Has anyone outside the developer tested it?
Often not. A systematic review in The Lancet Digital Health in 2019 examined 82 studies comparing deep learning with clinicians on medical imaging and found that only 25 of them performed out-of-sample external validation. The authors flagged poor reporting as a limit on interpreting the accuracy claims at all.
3. Accurate for which subgroups?
An average conceals who the system fails. Work presented at ACM CHIL in 2020 found models with strong overall performance showing relative performance differences of over 20% on clinically important subsets of the data — subsets that were not labelled and therefore not noticed. In commercial gender classification, research at FAT* in 2018 measured error rates of up to 34.7% for darker-skinned women against a maximum of 0.8% for lighter-skinned men. Both systems would report a healthy average.
Ask for performance broken out by the groups you serve, by site and by device.
4. What is it really keying on?
Models find shortcuts. A study in PLOS Medicine in 2018 found a pneumonia detection model scoring 0.931 internally but 0.815 at an external hospital — and that convolutional networks could identify which hospital system an image came from for 99.95% of images in one dataset. A model that can tell where an image was taken can learn the local disease prevalence instead of the disease.
5. Would the same score survive your test set?
Ask to run the model on your data, against criteria you set before you see the results. If the vendor will not permit a local evaluation, that is itself an answer.
What to do with the answers
Agree the acceptance criteria in writing before testing, evaluate on your own population, and report per subgroup. A procurement decision made on one number is a decision made on the least reliable number available.
Related questions
Do vendor accuracy figures usually hold at another organisation?
Often not. A 2021 JAMA Internal Medicine external validation found a widely deployed proprietary sepsis model scoring 0.63 AUC at an independent hospital, against the 0.76–0.83 reported by its developer, missing 67% of sepsis patients at the threshold in use.
Why is a single accuracy number misleading?
It averages away the groups the system fails. ACM CHIL research in 2020 found over 20% relative performance differences on clinically important subsets that were never labelled, inside models with strong overall scores.
What does it mean if a model performs worse at our site?
It may have learned a local shortcut. A 2018 PLOS Medicine study found a pneumonia model dropping from 0.931 to 0.815 externally, and networks identifying the source hospital for 99.95% of images — a proxy for local disease prevalence.
AI Testing & Certification
Test AI solutions for accuracy, safety and conformity, and prepare them for external certification.
Discuss your requirement