DISCERN score Perplexity demonstrated the highest mean DISCERN score among the evaluated models; however, the overall Kruskal-Wallis test did not reach statistical significance (p=0.058).
← all excerpts
Assessing Artificial Intelligence (AI) in Patient Education: Evaluating Accuracy and Readability of Responses on Surgical Procedures for Patellar Tendon Rupture.
2
0.0580
0.0750
The sentences
approached significancep = 0.075
Even though Perplexity demonstrated the highest mean DISCERN scores among the evaluated AI models, no statistically significant differences in readability were observed among the four chatbots, although results approached significance (p = 0.075).