The test results showed a trend completely opposite to that of the short-answer and multiple-choice questions, as GPT-4o’s performance was comprehensively superior to that of DeepSeek-R1’s performance.
← all excerpts
Large Language Model Evaluation in Traditional Chinese Medicine for Stroke: Quantitative Benchmarking Study.
1
—
—