Fisher’s exact test showed that the differences in overall accuracy among the models did not reach statistical significance (χ2 = 8.548, P = 0.070), as the value fell within the 95% confidence interval.
← all excerpts
Performance of five large language models in oral and maxillofacial surgery exam questions: a comparative study.
2
0.0700
0.0700
The sentences
The two domestic models, Qwen3 and DeepSeek-R1-0528, exhibited a slight trend towards higher accuracy on certain questions, particularly those closely associated with Chinese clinical guidelines or local epidemiology.