Barely Significant
← all excerpts

Comparison of GPT-5 and GPT-4o in Solving the Polish Centre for Medical Examinations (CEM) Gastroenterology Examination.

Cureus · 2026 · PMC12949840 · PMID 41773137

1
hedged sentence
0.1300
closest p · 2.6× alpha
0.1300
boldest claim

The sentences

did not reach statistical significancep = 0.13not close (p > 0.1)
For GPT-4o, the point-biserial correlation between confidence and correctness was weak and did not reach statistical significance (r = 0.14, 95% CI: −0.05 to 0.31, p = 0.13).

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.