Within-model comparison: While numerical scores were higher in English than in Turkish for all models, these differences did not reach statistical significance ( p > 0.05).
← all excerpts
Comparative performance of ChatGPT, Gemini, and Deepseek on endodontic exam questions in Turkish and English.
2
0.0500
0.0618
The sentences
marginally significantp = 0.0618
Only in the ChatGPT-4, the difference in providing correct explanations between Type S and Type C questions in Turkish was found to be marginally significant ( p = 0.0618).