Barely Significant
← all excerpts

Evaluating the diagnostic reasoning of large language models in complex neuro-ophthalmological cases: a comparative analysis of GPT-o1 Pro, GPT-4o, Gemini, Grok 2 and DeepSeek.

BMJ Open Ophthalmol · 2025 · PMC12684123 · PMID 41344901

1
hedged sentence
0.0010
closest p · 0.0× alpha
0.0010
boldest claim

The sentences

highly significantp<0.001actually significant
Simplicity GPT-o1 Pro, which used the fewest words, showed a highly significant difference compared with GPT-4o (p<0.001) and also with Gemini (p=0.032).

also in 132,142 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.