The sensitivity difference showed a numerical trend favoring ChatGPT-5.4 but did not reach statistical significance ( p = 0.087), and specificity was similar between the two models ( p = 0.885). 3.3.
← all excerpts
A retrospective comparative study of ChatGPT-5.4 and DeepSeek-VL2 for C-TIRADS-based risk stratification of thyroid nodules on ultrasound images.
1
—
—