highly significantp < 0.001
In terms of DISCERN scores, a highly significant difference was detected among the groups ( p < 0.001), and post-hoc Dunn testing indicated that Google Gemini 2.5 Flash differed significantly from ChatGPT-3.5, ChatGPT-5, Deepseek-V3 and Grok 3.