However, the difference between Gemini-3-Flash and GPT-4 (75.8%) did not reach statistical significance (P = 0.063) ( Figure 6 ). 3.5 Model performance by Ophthalmic Subspecialty Analysis of model accuracy across ten subspecialties revealed distinct performance profiles.
← all excerpts
Comparative performance of GPT-4, GPT-o3, GPT-5, Gemini-3-Flash, and DeepSeek-R1 in ophthalmology question answering.
1
0.0630
0.0630