Corroborating evidence was provided by Wilcoxon test, which did not detect any statistically significant difference overall between the scores given by the 2 evaluators for the answers provided by the 4 LLMs ( Table 2 ), except for the scores given for the answers provided by ChatGPT-4, between which a marginally statistically significant difference was found ( P =.049).
← all excerpts
Evaluation of the Performance of Generative AI Large Language Models ChatGPT, Google Bard, and Microsoft Bing Chat in Supporting Evidence-Based Dentistry: Comparative Mixed Methods Study.
1
0.0490
0.0490