Barely Significant
← all excerpts

Comparative evaluation and performance of large language models on expert level critical care questions: a benchmark study.

Crit Care · 2025 · PMC11809097 · PMID 39930514

1
hedged sentence
0.1960
closest p · 3.9× alpha
0.1960
boldest claim

The sentences

did not reach statistical significancep = 0.196not close (p > 0.1)
However, in contrast to the other evaluated LLMs ( p < 0.001), the performance of GPT-3.5-turbo compared to human physicians did not reach statistical significance ( p = 0.196).

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.