highly significantp < 0.001
The Friedman test revealed no statistically significant difference between HAIBU-ReMUD and ChatGPT-4o ( p > 0.05), while highly significant differences ( p < 0.001) were observed compared to all other models, solidifying HAIBU-ReMUD as the only benchmark-level model consistently maintaining scores above 4.5.