BFCS.ai
HealthcareBetterEvidence review

Can AI Make Diagnosis Better?

On narrow, well-defined tasks, AI now reads medical images about as well as a specialist — and better than the least experienced. The gain is real. The mistake is assuming it survives contact with a different hospital's data.

The BFCS verdictEvery article scored on the same four dimensions
Better
Yes — with a clinicianHigh confidence (assisted)
Assisted reading lifts detection and cuts recallsAI-supported reading raises cancer-detection rate and reduces unnecessary recalls; unaided autonomy is a different, unproven claim.
Faster
YesHigh confidence
Triage and worklist prioritisation cut workloadPrioritising suspicious studies and pre-reading normals relieves a genuine and growing radiology workload.
Cheaper
PartlyModerate confidence
Reader-time saved, monitoring cost addedEfficiency is real, but total cost of ownership includes integration, external validation and continuous drift monitoring.
Safer
ConditionalGovernance-gated
Accuracy can fall sharply off the training distributionReal-world accuracy can drop 15–30% under population shift, and automation bias is a live hazard. Safety is a monitoring discipline, not a launch event.

What improves, and why

Diagnosis is, in large part, pattern recognition under time pressure — exactly the kind of task deep learning does well. In screening mammography, AI-assisted reading has raised the cancer-detection rate and, at the same time, reduced recalls: more true cancers found, fewer women called back for nothing.Source 1 A large randomised trial found AI-supported single reading was a safe alternative to standard double reading, easing a workload that keeps rising faster than the specialist supply.Source 2

The important word is assisted. The evidence is strong for AI as a second set of eyes and a triage layer; it is far weaker for AI reading alone. Even the best benchmarks place fully autonomous interpretation years away, on ethical and safety grounds as much as technical ones.Source 6

The evidence, in four numbers
+9.6%
cancer detection rate
AI-assisted vs unaided reading (iCAD ProFound AI validation, Radiology)
−5.6%
recall rate
fewer women called back unnecessarily — same validation
0.87 / 0.96
AUROC — AI vs radiologists
AI noninferior only to the least-experienced readers (international benchmarking)
15–30%
real-world accuracy drop
from benchmark to deployment, driven by population shift & integration

The lesson from computer-aided detection is unambiguous: a tool that shines in retrospective studies can quietly make real-world reading worse. Prospective trials, not benchmarks, decide.

Where it falls short

  • Dataset shift. Accuracy measured on the training population can fall 15–30% when the model meets a different scanner, protocol or demographic. A model validated in Stockholm is not validated in Sydney.Source 5

  • The CAD precedent. Earlier computer-aided detection looked excellent retrospectively, then degraded accuracy in routine practice. The field has made this exact mistake before.Source 4

  • Narrow, not general. Strong performance on one lesion type or modality says little about the next. In head-to-head benchmarking, overall radiologist accuracy still exceeded AI.Source 3

Where governance decides the outcome

A better diagnosis is only better if it holds up after go-live. That makes diagnosis a governance problem as much as a modelling one: external validation on local data before deployment, continuous monitoring for drift, and a clear line of accountability for an AI-influenced decision. Automation bias — the tendency to defer to a confident machine — is a documented failure mode, and the countermeasure is process, not a better model. Better is the headline dimension here; Safer is the one that determines whether the headline lasts.

How to think about adoption

Adopt AI as an assistive layer with the clinician firmly in the loop, not as a replacement reader. Insist on validation against your own population before trusting a vendor's benchmark, instrument for drift from day one, and measure recall and detection — not headline accuracy — as the outcomes that matter. Treat every published figure, including the ones above, as a hypothesis to be re-tested in your setting.

Sources
  1. 1iCAD ProFound AI validation, published in Radiology: a 9.6% increase in cancer-detection rate and a 5.6% reduction in recall rate versus radiologists reading without AI. Vendor-linked validation — read alongside the RCT caveat below.
  2. 2Lång K, Josefsson V, Larsson A-M, et al. AI-supported single reading versus standard double reading in the MASAI trial: a randomised, controlled clinical-safety analysis. Lancet Oncology, 2023.
  3. 3International non-inferiority benchmarking (PMC): AI AUROC 0.87 (95% CI 0.83–0.90) versus radiologists 0.96 (0.94–0.97); AI was noninferior only to the least-experienced radiologists for mammography.
  4. 4Radiology: AI — the cautionary precedent. Traditional computer-aided detection (CAD) showed promise in retrospective studies but decreased diagnostic accuracy in real-world clinical use.
  5. 5Misdiagnosis review (PMC, 2025): benchmark accuracies as high as ~94.5% but 15–30% real-world performance drops due to population shifts and integration barriers; automation bias compounds the risk.
  6. 6Regulatory / WHO position: fully autonomous interpretation remains a distant possibility; current tools operate as a computational layer whose output is presented to the radiologist for verification.

BFCS.ai does not fabricate figures. Where evidence is vendor-linked or not yet confirmed by prospective trial, we label it as such and grade our confidence accordingly.

Continue the thread
HEALTHCARE · FASTERCan AI Cut Clinical Documentation Time?9 min · Read →HEALTHCARE · SAFERCan AI Make Medication Safer?12 min · Read →CONSTRUCTION · SAFERCan AI Make Construction Safer?14 min · Read →