Kruskal–Wallis omnibus analysis revealed statistically significant differences between groups for Paper II ( H = 8.78, df = 3, p = .032), while Paper I approached but did not reach significance ( H = 6.96, df = 3, p = .073).
← all excerpts
Benchmarking Large Language Models Against Psychiatry Residents Using Traditional Institutional Assessments.
3
0.0730
0.0730
The sentences
While post hoc pairwise comparisons did not reach statistical significance after Bonferroni correction (threshold α = 0.0167), the large and consistent effect sizes across all three independent AI systems suggest educationally meaningful differences warranting attention from psychiatric educators. 5 Some differences reached eight standard deviations above human, a gap so substantial it challenges fundamental assumptions about medical knowledge assessment.
These effect sizes far exceed Cohen’s conventions for “large” effects ( d > 0.8), indicating practically significant differences.