The performance gap remained statistically and practically significant, with a mean improvement of 14 percentage points and a very large effect size (Cohen’s d = 1.03).
← all excerpts
ChatGPT-4.0 and Medical Students: A Recognition-Gated Comparative Evaluation on Image-Based Medical Examinations.
1
—
—