In the present study, performance declined markedly in some models, with reduced accuracy for knowledge-based questions under sequential querying in ChatGPT-5, in contrast, the Gemini 2.5 Pro model showed a numerical decrease in performance on knowledge-based questions and a numerical increase on case-based questions; however, these changes did not reach statistical significance.
← all excerpts
The effectiveness of large language models in dental specialty questions: a comparative study in the field of prosthodontics.
1
—
—