Barely Significant
← all excerpts

Supervised Learning and Large Language Model Benchmarks on Mental Health Datasets: Cognitive Distortions and Suicidal Risks in Chinese Social Media.

Bioengineering (Basel) · 2025 · PMC12383806 · PMID 40868395

1
hedged sentence
0.0940
closest p · 1.9× alpha
0.0940
boldest claim

The sentences

did not reach statistical significancep = 0.094so close (0.05 < p ≤ 0.1)
Notably, for GPT-4, its best prompt strategy (scene-definition) outperformed its weakest strategy (basic) by approximately 1.67 percentage points in average F1, but the difference did not reach statistical significance ( p = 0.094 ).

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.