Barely Significant
← all excerpts

When AI models take the exam: large language models vs medical students on multiple-choice course exams.

Med Educ Online · 2025 · PMC12667333 · PMID 41316903

1
hedged sentence
0.0960
closest p · 1.9× alpha
0.0960
boldest claim

The sentences

a nonsignificant trendp = 0.096so close (0.05 < p ≤ 0.1)
In Respiratory Medicine, OpenAI o1 performed significantly better than ChatGPT−4 ( p = 0.006), Gemini ( p = 0.016), and Copilot ( p = 0.003), while the comparison with DeepSeek showed only a nonsignificant trend ( p = 0.096); no other pairwise differences were significant.

also in 5,406 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.