Barely Significant
← all excerpts

Clinical Safety and Reliability of Large Language Models in Answering Hemorrhoid-Related Patient Questions: A Comparative Study of ChatGPT, Gemini, and DeepSeek.

Healthcare (Basel) · 2026 · PMC13361431 · PMID 42450919

1
hedged sentence
closest p
boldest claim

The sentences

Post hoc analysis demonstrated a significant difference only between ChatGPT and DeepSeek ( p = 0.018), whereas comparisons between ChatGPT and Gemini and between Gemini and DeepSeek did not reach statistical significance ( Table 5 ).

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.