In interpretation-type questions, ChatGPT-5 ( p = 0.031), Gemini 3.0 Pro ( p = 0.031), and Claude Sonnet 4.5 ( p = 0.031) achieved significantly higher scores than all trainees, whereas ChatGPT-4o did not reach statistical significance (Table 1 ).
← all excerpts
Comparative performance of large language models in answering cornea and cataract surgery questions for resident training.
1
—
—