Barely Significant
← all excerpts

The performance of ChatGPT on medical image-based assessments and implications for medical education.

BMC Med Educ · 2025 · PMC12374324 · PMID 40849473

1
hedged sentence
closest p
boldest claim

The sentences

While GPT-4o demonstrated higher accuracy than GPT-4 across the image-based items, this difference did not reach statistical significance, likely due to the limited sample size.

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.