Barely Significant
← all excerpts

LLM enabled classification of patient self-reported symptoms and needs in health systems across the USA.

NPJ Digit Med · 2025 · PMC12215686 · PMID 40595018

1
hedged sentence
0.0500
closest p · 1.0× alpha
0.0500
boldest claim

The sentences

did not reach statistical significancep < 0.05actually significant
Paired t -tests on the individual classification outcomes revealed that the differences in recall (∆ = 0.008, p = 0.53), precision (∆ = 0.0252, p = 0.047), and accuracy (∆ = 0.008, p = 0.53) between GPT-4 and the supervised NLP did not reach statistical significance ( p < 0.05 across all tests) suggesting that the overall performance of GPT-4 and the supervised NLP model is statistically comparable (when GPT-4 is provided a curated master list).

also in 61,332 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.