Barely Significant
← all excerpts

The effectiveness of large language models in dental specialty questions: a comparative study in the field of prosthodontics.

BMC Med Educ · 2026 · PMC13005378 · PMID 41689029

1
hedged sentence
closest p
boldest claim

The sentences

In the present study, performance declined markedly in some models, with reduced accuracy for knowledge-based questions under sequential querying in ChatGPT-5, in contrast, the Gemini 2.5 Pro model showed a numerical decrease in performance on knowledge-based questions and a numerical increase on case-based questions; however, these changes did not reach statistical significance.

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.