Statistical testing showed that GPT-4o (t(46) = 11.92, p < 0.001) and DeepSeek (t(46) = 18.99, p < 0.001) were significantly more complete than residents, while the difference between GPT-3.5 and residents did not reach statistical significance (t(46) = 1.95, p = 0.058).
← all excerpts
Comparing Artificial Intelligence and Obstetrics Residents in Answering Standardized Patient Questions Regarding Gestational Diabetes.
1
0.0580
0.0580