←back to Blog

Real-Time Artificial Intelligence Diagnostic Copilot in Simulated Primary Care Consultations: Randomized Simulation Study

JMIR Form Res. 2026 Sep 22;10:e104579. doi: 10.2196/104579.

ABSTRACT

BACKGROUND: Diagnostic errors in primary care contribute substantially to avoidable morbidity. Real-time, voice-based artificial intelligence (AI) diagnostic assistants may support diagnostic reasoning during clinical encounters, but formative evidence that can be used to assess both diagnostic benefit and overreliance risk before clinical deployment is needed from interactive settings.

OBJECTIVE: This study aimed to conduct a formative, high-difficulty simulation stress test of whether a real-time, voice-based AI diagnostic assistant was associated with improved physician diagnostic accuracy and measurable overreliance during simulated primary care consultations.

METHODS: We conducted a case-level randomized, adjudicator-blinded simulation study in a web-based virtual primary care clinic. Thirteen board-certified family and community medicine physicians managed 260 simulated voice-based consultations. Cases were assigned to either an AI-assisted condition or an unassisted, resource-restricted simulation condition without any external diagnostic aids. No real patients were enrolled, and no clinical care was delivered or modified. The primary outcome was Top-3 diagnostic accuracy, defined as at least one correct diagnosis among up to 3 submitted diagnoses assessed by 3 blinded adjudicators and analyzed using a frequentist binomial generalized linear mixed model fitted by maximum likelihood with random intercepts for physician and case.

RESULTS: All 260 consultations were completed and analyzed. Adjusted Top-3 diagnostic accuracy was 62.3% (95% CI 49.1% to 75.7%) in the unassisted condition and 74.6% (95% CI 63.7% to 85.4%) in the AI-assisted condition (adjusted odds ratio [AOR] 2.68, 95% CI 1.24 to 6.12; χ21=6.4, P=.01). This corresponds to an adjusted absolute difference of +12.3 percentage points (95% CI +2.7 to +22.6) and a simulation-context number needed to treat equivalent of 8.1 (95% CI 4.4 to 37.2). The post hoc Top-2 analysis was supportive (AOR 2.59, 95% CI 1.23 to 5.77; χ21=6.3, P=.01). Top-1 point estimates favored AI assistance but were inconclusive: Adjusted accuracy was 47.8% in the unassisted condition versus 57% with AI assistance (AOR 1.83, 95% CI 0.93 to 3.80; χ21=3.1, P=.08). Leave-one-physician-out and virtual-patient simulator-exclusion sensitivity analyses were consistent with the primary finding. In risk scenarios with an incorrect operative AI suggestion, adjusted concordance was 38.8% in the unassisted condition and 55.8% in the AI-assisted condition, but the between-condition comparison was not statistically significant (AOR 2.08, 95% CI 0.89 to 5.19; χ21=2.9, P=.09). Mean consultation time increased by 10.7%, and physicians rated the system highly for usefulness and satisfaction.

CONCLUSIONS: In this randomized, high-difficulty simulation study, real-time, voice-based AI assistance was associated with higher physician Top-3 diagnostic accuracy than an unassisted, resource-restricted comparator. Some sensitivity analyses were supportive, while others were directionally favorable but inconclusive, and the safety analysis suggested possible overreliance without statistically conclusive difference between conditions. These findings support further refinement and prospective evaluation in real clinical workflows before conclusions are drawn about routine clinical effectiveness or implementation.

PMID:42771885 | DOI:10.2196/104579