Gemini 2.5 significantly outperformed GPT-4.5 ( P = 0.029) and Gemini 2.0 ( P = 0.010) for all questions combined, while other pairwise comparisons did not reach statistical significance ( P > 0.05).
← all excerpts
Analysis of multimodal large language models on visually-based questions in the Japanese National Examination for Dental Hygienists: A preliminary comparative study.
1
0.0500
0.0500