Barely Significant
← all excerpts

Evaluation of large language models for VI-RADS reports: a comparative analysis of zero-shot and few-shot prompting.

BMC Med Imaging · 2026 · PMC13191856 · PMID 41965564

1
hedged sentence
0.1080
closest p · 2.2× alpha
0.1080
boldest claim

The sentences

did not reach statistical significancep = 0.108not close (p > 0.1)
In the ChatGPT (OpenAI; GPT-5.2) group, few-shot prompting modestly improved accuracy (from 0.700 to 0.850), macro F1 score (from 0.693 to 0.848), and kappa (from 0.400 to 0.700), although the difference did not reach statistical significance ( p = 0.108).

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.