Although the Cochran's Q test again did not reach statistical significance (Q=5.66, df=4, p=0.226), the absolute performance gap between written and image-based questions averaged approximately 35 percentage points across models, a clinically and educationally meaningful difference that was consistent across all five architectures.
← all excerpts
Comparative Performance of Seven Mainstream Large Language Models on the 2022 American College of Radiology Diagnostic Imaging In-Training Examination.
1
0.2260
0.2260