Barely Significant
← all excerpts

Performance of five large language models in oral and maxillofacial surgery exam questions: a comparative study.

BMC Oral Health · 2026 · PMC13077904 · PMID 41792703

2
hedged sentences
0.0700
closest p · 1.4× alpha
0.0700
boldest claim

The sentences

did not reach statistical significanceP = 0.070so close (0.05 < p ≤ 0.1)
Fisher’s exact test showed that the differences in overall accuracy among the models did not reach statistical significance (χ2 = 8.548, P = 0.070), as the value fell within the 95% confidence interval.

also in 111,027 other papers

a slight trendno p-value reported
The two domestic models, Qwen3 and DeepSeek-R1-0528, exhibited a slight trend towards higher accuracy on certain questions, particularly those closely associated with Chinese clinical guidelines or local epidemiology.

also in 4,078 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.