For GPT-4o, the point-biserial correlation between confidence and correctness was weak and did not reach statistical significance (r = 0.14, 95% CI: −0.05 to 0.31, p = 0.13).
← all excerpts
Comparison of GPT-5 and GPT-4o in Solving the Polish Centre for Medical Examinations (CEM) Gastroenterology Examination.
1
0.1300
0.1300