Barely Significant
← all excerpts

Evaluation of validity, reliability, and readability of AI chatbots for gestational diabetes mellitus: a multi-model comparative study.

Front Public Health · 2026 · PMC12913397 · PMID 41717624

1
hedged sentence
0.0001
closest p · 0.0× alpha
0.0001
boldest claim

The sentences

highly significantp < 0.0001actually significant
ChatGPT-5 again had the highest score (71.67 ± 6.17), followed by DeepSeek-V3.2 (66.00 ± 5.07), Gemini (61.67 ± 5.88), and Claude Sonnet (59.00 ± 6.87), with a highly significant overall difference ( p < 0.0001).

also in 132,142 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.