Although DeepSeek produced the highest proportion of higher-order cognitive questions and the lowest proportion of Knowledge/Comprehension items, the differences between the three models did not reach statistical significance ( p = 0.08 for higher-order levels; p = 0.06 for Knowledge/Comprehension).
← all excerpts
Evaluation of three artificial intelligence chatbots for generating clinical hematology multiple choice questions for medical students.
1
0.0800
0.0800