highly significantp < 0.001
Results The comprehensive scores ranked from highest to lowest were DeepSeek (4.47), Qwen (4.33), Kimi (4.24), Doubao (4.13), and ChatGPT (3.41), with highly significant differences were observed among all models ( H =182.14, p < 0.001).