For content comprehensiveness, both ChatGPT-4o and Gemini-2.5 Pro had a median score of 4.4, higher than DeepSeek-R1 (median: 4.2), though differences did not reach statistical significance (p=0.0536).
← all excerpts
Exploring and Comparing the Use of Large Language Models in Supporting Osteoporosis Health Consultations.
1
0.0536
0.0536