Statistical testing across categories showed that differences in accuracy between the two models did not reach statistical significance (p > 0.05 for all comparisons).
← all excerpts
Comparative Analysis of ChatGPT-4o and Gemini Advanced Performance on Diagnostic Radiology In-Training Exams.
1
0.0500
0.0500