Diagnostic accuracy Of the 20 cases, Gemini Advanced (19/20; 95%; p=0.01) performed significantly better, and ChatGPT-4.0 (18/20; 90%; p=0.05) almost reached statistical significance compared to ChatGPT-3.5 (Table 4 ).
← all excerpts
Performance of Large Language Models (ChatGPT and Gemini Advanced) in Gastrointestinal Pathology and Clinical Review of Applications in Gastroenterology.
1
—
—