Can AI Make Diagnosis Better?
On narrow, well-defined tasks, AI now reads medical images about as well as a specialist — and better than the least experienced. The gain is real. The mistake is assuming it survives contact with a different hospital's data.
What improves, and why
Diagnosis is, in large part, pattern recognition under time pressure — exactly the kind of task deep learning does well. In screening mammography, AI-assisted reading has raised the cancer-detection rate and, at the same time, reduced recalls: more true cancers found, fewer women called back for nothing.Source 1 A large randomised trial found AI-supported single reading was a safe alternative to standard double reading, easing a workload that keeps rising faster than the specialist supply.Source 2
The important word is assisted. The evidence is strong for AI as a second set of eyes and a triage layer; it is far weaker for AI reading alone. Even the best benchmarks place fully autonomous interpretation years away, on ethical and safety grounds as much as technical ones.Source 6
The lesson from computer-aided detection is unambiguous: a tool that shines in retrospective studies can quietly make real-world reading worse. Prospective trials, not benchmarks, decide.
Where it falls short
Dataset shift. Accuracy measured on the training population can fall 15–30% when the model meets a different scanner, protocol or demographic. A model validated in Stockholm is not validated in Sydney.Source 5
The CAD precedent. Earlier computer-aided detection looked excellent retrospectively, then degraded accuracy in routine practice. The field has made this exact mistake before.Source 4
Narrow, not general. Strong performance on one lesion type or modality says little about the next. In head-to-head benchmarking, overall radiologist accuracy still exceeded AI.Source 3
Where governance decides the outcome
A better diagnosis is only better if it holds up after go-live. That makes diagnosis a governance problem as much as a modelling one: external validation on local data before deployment, continuous monitoring for drift, and a clear line of accountability for an AI-influenced decision. Automation bias — the tendency to defer to a confident machine — is a documented failure mode, and the countermeasure is process, not a better model. Better is the headline dimension here; Safer is the one that determines whether the headline lasts.
How to think about adoption
Adopt AI as an assistive layer with the clinician firmly in the loop, not as a replacement reader. Insist on validation against your own population before trusting a vendor's benchmark, instrument for drift from day one, and measure recall and detection — not headline accuracy — as the outcomes that matter. Treat every published figure, including the ones above, as a hypothesis to be re-tested in your setting.
- 1iCAD ProFound AI validation, published in Radiology: a 9.6% increase in cancer-detection rate and a 5.6% reduction in recall rate versus radiologists reading without AI. Vendor-linked validation — read alongside the RCT caveat below.
- 2Lång K, Josefsson V, Larsson A-M, et al. AI-supported single reading versus standard double reading in the MASAI trial: a randomised, controlled clinical-safety analysis. Lancet Oncology, 2023.
- 3International non-inferiority benchmarking (PMC): AI AUROC 0.87 (95% CI 0.83–0.90) versus radiologists 0.96 (0.94–0.97); AI was noninferior only to the least-experienced radiologists for mammography.
- 4Radiology: AI — the cautionary precedent. Traditional computer-aided detection (CAD) showed promise in retrospective studies but decreased diagnostic accuracy in real-world clinical use.
- 5Misdiagnosis review (PMC, 2025): benchmark accuracies as high as ~94.5% but 15–30% real-world performance drops due to population shifts and integration barriers; automation bias compounds the risk.
- 6Regulatory / WHO position: fully autonomous interpretation remains a distant possibility; current tools operate as a computational layer whose output is presented to the radiologist for verification.
BFCS.ai does not fabricate figures. Where evidence is vendor-linked or not yet confirmed by prospective trial, we label it as such and grade our confidence accordingly.