Barely Significant
← all excerpts

Comparative Assessment of Large Language Models in Optics and Refractive Surgery: Performance on Multiple-Choice Questions.

Vision (Basel) · 2025 · PMC12550897 · PMID 41133609

1
hedged sentence
closest p
boldest claim

The sentences

Comparisons with DeepSeek R1 and ChatGPT O3 Mini did not reach statistical significance, indicating that their performance was not substantially different from ChatGPT O1.

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.