When comparing the total scores across all models, the results approached but did not reach statistical significance (ANOVA, p =0.215; Kruskal-Wallis, p =0.219).
← all excerpts
Performance of large language models in fluoride-related dental knowledge: a comparative evaluation study of ChatGPT-4, Claude 3.5 Sonnet, Copilot, and Grok 3.
1
0.2150
0.2150