Post hoc analysis demonstrated a significant difference only between ChatGPT and DeepSeek ( p = 0.018), whereas comparisons between ChatGPT and Gemini and between Gemini and DeepSeek did not reach statistical significance ( Table 5 ).
← all excerpts
Clinical Safety and Reliability of Large Language Models in Answering Hemorrhoid-Related Patient Questions: A Comparative Study of ChatGPT, Gemini, and DeepSeek.
1
—
—