Barely Significant
← all excerpts

Comparative Performance of State-of-the-Art LLMs on the KDLE: A 2025 Benchmark Study.

Int Dent J · 2026 · PMC12962161 · PMID 41762790

1
hedged sentence
0.0510
closest p · 1.0× alpha
0.0510
boldest claim

The sentences

borderline significantP = .051so close (0.05 < p ≤ 0.1)
When accuracy differences across the 4 subtypes (radiograph, clinical photo, schematic illustration, and mixed) were assessed using chi-square tests, subtype-level variability within each model was significant (Claude-4 Opus, P = .040; Gemini 2.5 Pro, P = .005) or borderline significant (ChatGPT-4o, P = .051), confirming that image subtype substantially influenced model performance.

also in 11,409 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.