For the remaining LLMs, the differences in performance between the 2 sections did not reach statistical significance (minimum P =.34).
← all excerpts
Performance of the Large Language Models on the Chinese National Nurse Licensure Examination: Cross-Sectional Evaluation Study.
2
0.3400
0.3400
The sentences
(2) In the Practical Skills section, DeepSeek V3 showed an increasing trend for both unique correct answers (3 and 6) and unique incorrect answers (2 and 3) between Attempt 1 and Attempt 2.