Barely Significant
← all excerpts

Comparing Artificial Intelligence and Obstetrics Residents in Answering Standardized Patient Questions Regarding Gestational Diabetes.

Cureus · 2025 · PMC12618048 · PMID 41246718

1
hedged sentence
0.0580
closest p · 1.2× alpha
0.0580
boldest claim

The sentences

did not reach statistical significancep = 0.058so close (0.05 < p ≤ 0.1)
Statistical testing showed that GPT-4o (t(46) = 11.92, p < 0.001) and DeepSeek (t(46) = 18.99, p < 0.001) were significantly more complete than residents, while the difference between GPT-3.5 and residents did not reach statistical significance (t(46) = 1.95, p = 0.058).

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.