Barely Significant
← all excerpts

Analysis of multimodal large language models on visually-based questions in the Japanese National Examination for Dental Hygienists: A preliminary comparative study.

J Dent Sci · 2026 · PMC12825506 · PMID 41585131

1
hedged sentence
0.0500
closest p · 1.0× alpha
0.0500
boldest claim

The sentences

did not reach statistical significanceP > 0.05actually significant
Gemini 2.5 significantly outperformed GPT-4.5 ( P = 0.029) and Gemini 2.0 ( P = 0.010) for all questions combined, while other pairwise comparisons did not reach statistical significance ( P > 0.05).

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.