Barely Significant
← all excerpts

GPT-4o and OpenAI o1 Performance on the 2024 Spanish Competitive Medical Specialty Access Examination: Cross-Sectional Quantitative Evaluation Study.

JMIR Med Educ · 2026 · PMC12795474 · PMID 41525685

1
hedged sentence
0.2200
closest p · 4.4× alpha
0.2200
boldest claim

The sentences

GPT-4o, OpenAI o1, and the AMIR consensus achieved significantly higher accuracy scores than the average student (in all cases P <.001); however, differences between these 3 arms did not reach statistical significance ( P =.22 for GPT-4o vs OpenAI o1; P =.07 for GPT-4o vs AMIR consensus; P =.75 for OpenAI o1 vs AMIR consensus).

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.