DeepSeek and Gemini demonstrated relatively higher median reliability scores overall; however, pairwise comparisons showed that differences between Claude and either DeepSeek or Gemini did not reach statistical significance.
← all excerpts
Large language model chatbots as sources of pediatric anesthesia health advice: An evaluation of reliability and readability.
1
—
—