Comparisons with DeepSeek R1 and ChatGPT O3 Mini did not reach statistical significance, indicating that their performance was not substantially different from ChatGPT O1.
← all excerpts
Comparative Assessment of Large Language Models in Optics and Refractive Surgery: Performance on Multiple-Choice Questions.
1
—
—