Barely Significant
← all excerpts

Evaluation of Large Language Models in Infectious Disease Decision-Making: From Examination to Clinical Practice.

Infect Drug Resist · 2026 · PMC13271053 · PMID 42311928

2
hedged sentences
0.3400
closest p · 6.8× alpha
0.3400
boldest claim

The sentences

showed a trendp = 0.34not close (p > 0.1)
LLMs showed a trend toward better performance on low-order, knowledge-based questions ( p = 0.34), whereas doctors tended to perform better on simple case-based questions ( p = 0.74), particularly those requiring higher-order clinical reasoning ( p = 0.10).

also in 33,237 other papers

a numerical trendno p-value reported
Gemini 2.5 Flash showed a numerical trend toward slightly higher overall performance, with higher mean scores for accuracy and completeness (accuracy: 4.60 ± 0.56; completeness: 4.70 ± 0.47).

also in 723 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.