←back to Blog

Impact of Responsibility Allocation Structures on Diagnostic Quality in AI-Assisted Diagnosis: Randomized Controlled Experiment

J Med Internet Res. 2026 Jul 23;28:e97588. doi: 10.2196/97588.

ABSTRACT

BACKGROUND: AI is increasingly used to support clinical diagnosis, but the appropriate allocation of responsibility between clinicians and AI remains unclear. Different responsibility structures may influence how clinicians evaluate AI recommendations and revise their diagnostic judgments.

OBJECTIVE: This study aimed to examine how different physician-AI responsibility allocation structures affect diagnostic accuracy and confidence calibration during AI-assisted diagnosis.

METHODS: This individually randomized, 4-arm, parallel-group controlled experiment was conducted in a simulated clinical environment on the Credamo platform (Beijing Yishumofa Technology Co, Ltd). A total of 105 licensed physicians were randomly assigned to the dynamic responsibility, full responsibility, equal responsibility, or control group. Nine participants who failed the prespecified attention checks were excluded from the primary analysis, resulting in an analytic sample of 96 physicians. Participants completed 10 clinical vignette-based diagnostic tasks. The only between-group difference was the responsibility allocation structure. Primary outcomes were final diagnostic accuracy and confidence calibration; secondary outcomes included agreement rates and posttask subjective evaluations.

RESULTS: Responsibility allocation structures significantly modulated diagnostic quality. Compared to the control group, the full responsibility structure yielded no significant improvement in accuracy (mean 0.596, SD 0.152 vs 0.592, SD 0.169; mean difference 0.004, 95% CI -0.089 to 0.098; P=.93) or confidence calibration (mean 0.193, SD 0.134 vs 0.150, SD 0.145; mean difference 0.043, 95% CI -0.038 to 0.124; P=.29), while the equal responsibility structure showed suggestive evidence of lower diagnostic accuracy (mean 0.496, SD 0.185 vs 0.592, SD 0.169; mean difference -0.096, 95% CI -0.199 to 0.007; P=.07) and significantly poorer confidence calibration (mean 0.383, SD 0.175 vs 0.150, SD 0.145; mean difference 0.233, 95% CI 0.139 to 0.326; P<.001). Conversely, the dynamic responsibility structure demonstrated superior performance, significantly improving diagnostic accuracy (mean 0.717, SD 0.105 vs 0.592, SD 0.169; mean difference 0.125, 95% CI 0.043 to 0.207; P=.004) and reducing confidence calibration (mean 0.040, SD 0.084 vs 0.150, SD 0.105; mean difference -0.110, 95% CI -0.179 to -0.040; P=.003).

CONCLUSIONS: The dynamic responsibility structure may enable health care organizations to use AI more fully and appropriately without compromising clinicians’ diagnostic performance, thereby improving the safety and quality of AI-assisted diagnosis.

PMID:42492064 | DOI:10.2196/97588