Barely Significant
← all excerpts

Benchmarking Large Language Models Against Psychiatry Residents Using Traditional Institutional Assessments.

Indian J Psychol Med · 2026 · PMC13065625 · PMID 41970368

3
hedged sentences
0.0730
closest p · 1.5× alpha
0.0730
boldest claim

The sentences

approached but did not reach significancep = .073so close (0.05 < p ≤ 0.1)
Kruskal–Wallis omnibus analysis revealed statistically significant differences between groups for Paper II ( H = 8.78, df = 3, p = .032), while Paper I approached but did not reach significance ( H = 6.96, df = 3, p = .073).

also in 94 other papers

While post hoc pairwise comparisons did not reach statistical significance after Bonferroni correction (threshold α = 0.0167), the large and consistent effect sizes across all three independent AI systems suggest educationally meaningful differences warranting attention from psychiatric educators. 5 Some differences reached eight standard deviations above human, a gap so substantial it challenges fundamental assumptions about medical knowledge assessment.

also in 50,127 other papers

practically significantno p-value reported
These effect sizes far exceed Cohen’s conventions for “large” effects ( d > 0.8), indicating practically significant differences.

also in 785 other papers

Quoted from the open-access full text in Europe PMC under the licence the publisher applied. The sentence is reproduced exactly as published; the emphasis is ours.