Barely Significant
← all excerpts

Evaluation of three artificial intelligence chatbots for generating clinical hematology multiple choice questions for medical students.

Sci Rep · 2026 · PMC12895027 · PMID 41559279

1
hedged sentence
0.0800
closest p · 1.6× alpha
0.0800
boldest claim

The sentences

did not reach statistical significancep = 0.08so close (0.05 < p ≤ 0.1)
Although DeepSeek produced the highest proportion of higher-order cognitive questions and the lowest proportion of Knowledge/Comprehension items, the differences between the three models did not reach statistical significance ( p = 0.08 for higher-order levels; p = 0.06 for Knowledge/Comprehension).

also in 111,027 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.