In the safety dimension, DeepSeek R1 achieved a median consensus score of 4.0 (IQR 4.0‐4.5), which was lower than that of GPT-5 (median consensus score 4.25, IQR 4.0‐4.5), although this difference did not reach statistical significance ( P =.94).
← all excerpts
Performance Evaluation of GPT-5, Grok 4, and DeepSeek R1 in Interpreting Complete Blood Count Reports for Hematologic Diseases: Retrospective Comparative Study.
1
0.9400
0.9400