The difference in scores of the most recent ChatGPT version, ChatGPT-d, did not reach statistical significance ( P = 0.079).
← all excerpts
Chatbot responses suggest that hypothetical biology questions are harder than realistic ones.
1
0.0790
0.0790
The difference in scores of the most recent ChatGPT version, ChatGPT-d, did not reach statistical significance ( P = 0.079).