In the treatment domain, cardiovascular surgeons demonstrated superior performance, achieving 96.3% accuracy by correctly answering 26 of 27 pooled responses, compared to MLLM’s 88.9% accuracy; however, this difference did not reach statistical significance, with a p -value of 0.362.
← all excerpts
Comparative Performance of Multimodal and Unimodal Large Language Models Versus Multicenter Human Clinical Experts in Aortic Dissection Management.
1
—
—