Tukey post hoc analysis showed a significant difference between Copilot-5 and DeepSeek 3.1V for both FRES and CLI, while comparisons involving ChatGPT-5 did not reach statistical significance.
← all excerpts
Assessing large language model responses to pediatric depression FAQs: a cross-sectional study on readability, accuracy, and sentiment.
1
—
—