Barely Significant
← all excerpts

Benchmark evaluation of multi-modal large language models for ophthalmic diagnosis in real world.

Front Med (Lausanne) · 2026 · PMC13333443 · PMID 42440611

1
hedged sentence
0.0010
closest p · 0.0× alpha
0.0010
boldest claim

The sentences

highly significantp < 0.001actually significant
The Friedman test revealed no statistically significant difference between HAIBU-ReMUD and ChatGPT-4o ( p > 0.05), while highly significant differences ( p < 0.001) were observed compared to all other models, solidifying HAIBU-ReMUD as the only benchmark-level model consistently maintaining scores above 4.5.

also in 132,142 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.