Barely Significant
← all excerpts

Comparative performance of ChatGPT, Gemini, and Deepseek on endodontic exam questions in Turkish and English.

BMC Oral Health · 2026 · PMC12949502 · PMID 41639680

2
hedged sentences
0.0500
closest p · 1.0× alpha
0.0618
boldest claim

The sentences

did not reach statistical significancep > 0.05actually significant
Within-model comparison: While numerical scores were higher in English than in Turkish for all models, these differences did not reach statistical significance ( p > 0.05).

also in 111,027 other papers

marginally significantp = 0.0618so close (0.05 < p ≤ 0.1)
Only in the ChatGPT-4, the difference in providing correct explanations between Type S and Type C questions in Turkish was found to be marginally significant ( p = 0.0618).

also in 26,082 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.