Barely Significant
← all excerpts

Efficacy of Large Language Models in Providing Evidence-Based Patient Education for Celiac Disease: A Comparative Analysis.

Nutrients · 2025 · PMC12735979 · PMID 41470773

3
hedged sentences
0.0010
closest p · 0.0× alpha
0.7780
boldest claim

The sentences

highly significantp < 0.001actually significant
Readability and Linguistic Analysis Highly significant between-model differences were observed across all readability metrics with large effect sizes (all Friedman p < 0.001): Flesch Reading Ease: χ 2 = 19.70, p < 0.001; Flesch-Kincaid Grade Level: χ 2 = 22.55, p < 0.001; SMOG Index: χ 2 = 23.97, p < 0.001.

also in 132,142 other papers

approached significancep = 0.053so close (0.05 < p ≤ 0.1)
Gemini 2.0 comparison approached significance ( p = 0.053).

also in 8,237 other papers

did not reach statistical significancep = 0.778not close (p > 0.1)
Gemini also demonstrated lower misinformation rates (13.3%) compared to ChatGPT-4 (23.3%) and Claude 3.7 (24.2%), representing a clinically meaningful 40–45% reduction in potentially harmful content, although these differences did not reach statistical significance ( p = 0.778).

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.