The majority of LLMs scored higher on Paper 1 compared to Paper 2, though this did not reach statistical significance ( p =0.137).
← all excerpts
Performance of large language models at the MRCS Part A: a tool for medical education?
1
0.1370
0.1370
The majority of LLMs scored higher on Paper 1 compared to Paper 2, though this did not reach statistical significance ( p =0.137).