Barely Significant
← all excerpts

ChatGPT-4o and OpenAI-o1: A Comparative Analysis of Its Accuracy in Refractive Surgery.

J Clin Med · 2025 · PMC12347465 · PMID 40806797

2
hedged sentences
0.1045
closest p · 2.1× alpha
0.1045
boldest claim

The sentences

did not reach statistical significancep = 0.1045not close (p > 0.1)
The five-point margin of OpenAI-o1 over ChatGPT-4o did not reach statistical significance ( p = 0.1045) but could represent one additional correct decision in twenty clinically relevant scenarios.

also in 111,027 other papers

practically significantno p-value reported
Our primary aim was to determine whether any observed performance gain is clinically meaningful, defined as a statistically and practically significant margin over resident performance that could justify integrating the model as a decision support aid during patient work-up and surgical counseling.

also in 1,558 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.