Barely Significant
← all excerpts

Comparative performance of GPT-4, GPT-o3, GPT-5, Gemini-3-Flash, and DeepSeek-R1 in ophthalmology question answering.

Front Cell Dev Biol · 2026 · PMC12894337 · PMID 41695391

1
hedged sentence
0.0630
closest p · 1.3× alpha
0.0630
boldest claim

The sentences

did not reach statistical significanceP = 0.063so close (0.05 < p ≤ 0.1)
However, the difference between Gemini-3-Flash and GPT-4 (75.8%) did not reach statistical significance (P = 0.063) ( Figure 6 ). 3.5 Model performance by Ophthalmic Subspecialty Analysis of model accuracy across ten subspecialties revealed distinct performance profiles.

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.