GPT-4o, OpenAI o1, and the AMIR consensus achieved significantly higher accuracy scores than the average student (in all cases P <.001); however, differences between these 3 arms did not reach statistical significance ( P =.22 for GPT-4o vs OpenAI o1; P =.07 for GPT-4o vs AMIR consensus; P =.75 for OpenAI o1 vs AMIR consensus).
← all excerpts
GPT-4o and OpenAI o1 Performance on the 2024 Spanish Competitive Medical Specialty Access Examination: Cross-Sectional Quantitative Evaluation Study.
1
0.2200
0.2200