Barely Significant
← all excerpts

Assessing risk of bias of cohort studies with large language models.

Res Synth Methods · 2025 · PMC12657654 · PMID 41626981

1
hedged sentence
0.6400
closest p · 12.8× alpha
0.6400
boldest claim

The sentences

did not reach statistical significanceP = 0.64not close (p > 0.1)
The differences in overall correct assessment rates among the three models did not reach statistical significance (ChatGPT-4o versus Moonshot-v1-128k: RD = 2.3%, 95% CI: −0.38% to 4.98%, P = 0.64; ChatGPT-4o versus DeepSeek-V3: RD = 0.09%, 95% CI: −0.83% to 1.01%, P = 0.98).

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.