However, in contrast to the other evaluated LLMs ( p < 0.001), the performance of GPT-3.5-turbo compared to human physicians did not reach statistical significance ( p = 0.196).
← all excerpts
Comparative evaluation and performance of large language models on expert level critical care questions: a benchmark study.
1
0.1960
0.1960