What happens when medical artificial intelligence confidently delivers the wrong diagnosis? Researchers at TU Dresden have investigated a surprisingly simple way to identify potentially unreliable answers: asking the same question more than once.
Their study, published in Nature Medicine in September 2026, suggests that diagnostic consistency could help clinicians distinguish more dependable AI-generated conclusions from those requiring closer examination.
<h3>Putting Medical AI To The Test</h3>
The research team developed an AI system that operates entirely within a medical institution's own infrastructure, keeping sensitive patient information under local control.
To evaluate its performance, the scientists created simulated consultations involving two AI agents. One acted as a doctor, asking questions and requesting medical examinations, while the other represented a patient using anonymised clinical records.
The tests included conditions such as appendicitis, pneumonia, pulmonary embolism and urinary tract infections. The strongest locally operated models achieved diagnostic accuracy of 90.04% in a seven-disease assessment and 83.8% in a separate four-disease assessment.
Physicians independently reviewed 181 selected cases, with their consensus agreeing with the automated evaluation in 92.3% of cases.
<h3>Why Repeating Answers Matters</h3>
A high overall accuracy rate does not establish whether an individual diagnosis is correct. The researchers therefore examined 551 clinical cases, comparing several indicators of reliability. They assessed the AI's internal probability scores, the certainty expressed in its explanations and the consistency of its answers across five independent runs of the same case.
Diagnostic consistency proved the strongest indicator of correctness, achieving an area under the curve of 0.860, a statistical measure of how effectively a test distinguishes between two outcomes.
When researchers selected cases with consistency scores of at least 0.90, the system achieved 98.9% accuracy. However, only 49.4% of cases met that threshold, leaving approximately half requiring additional assessment.
<h3>Confidence Can Be Misleading</h3>
The distinction became particularly important when researchers deliberately removed reliable clinical information.
Diagnostic accuracy fell from 90.6% to 70.2%, yet some of the AI's internal probability scores remained high. In contrast, consistency declined as the information became less dependable and continued to distinguish correct from incorrect diagnoses.
The findings illustrate why an AI system's apparent certainty should not be confused with genuine diagnostic reliability. An answer can sound authoritative even when the evidence supporting it is incomplete.
<h3>A Different Role For Doctors</h3>
The researchers propose using consistency to support selective autonomy, where AI could handle a carefully identified subset of cases while referring uncertain results to clinicians.
Senior researcher Jakob Nikolas Kather emphasised that the objective is to support medical professionals rather than replace their judgement. Making uncertainty visible, he explained, is essential to ensuring that responsibility for diagnosis and treatment remains with clinicians.
Running the technology within hospital-controlled infrastructure could also give healthcare organisations greater control over patient privacy, access permissions and system monitoring.
<h3>What Still Needs To Change</h3>
Despite encouraging results, consistency is not a guarantee of correctness. The system occasionally produced stable but incorrect diagnoses, and its performance was lower in simulated cases involving older patients.
The study also relied on retrospective simulations rather than routine hospital care. Repeating each assessment increases computational demands, while reliability thresholds may need adjusting for different AI models. Further clinical testing will be necessary to establish whether this approach improves patient safety in practice.
Rather than expecting medical AI to be certain about every decision, the research points towards a more practical objective: developing systems that can recognise when their conclusions deserve a second opinion.