Barely Significant
← all excerpts

Comparative performance of large language models in answering cornea and cataract surgery questions for resident training.

BMC Ophthalmol · 2026 · PMC13151304 · PMID 41904507

1
hedged sentence
closest p
boldest claim

The sentences

In interpretation-type questions, ChatGPT-5 ( p = 0.031), Gemini 3.0 Pro ( p = 0.031), and Claude Sonnet 4.5 ( p = 0.031) achieved significantly higher scores than all trainees, whereas ChatGPT-4o did not reach statistical significance (Table 1 ).

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.