Barely Significant
← all excerpts

Comparative Performance of Seven Mainstream Large Language Models on the 2022 American College of Radiology Diagnostic Imaging In-Training Examination.

Cureus · 2026 · PMC13148569 · PMID 42099318

1
hedged sentence
0.2260
closest p · 4.5× alpha
0.2260
boldest claim

The sentences

did not reach statistical significancep=0.226not close (p > 0.1)
Although the Cochran's Q test again did not reach statistical significance (Q=5.66, df=4, p=0.226), the absolute performance gap between written and image-based questions averaged approximately 35 percentage points across models, a clinically and educationally meaningful difference that was consistent across all five architectures.

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.