Gemini excelled in “pathogenesis” and “diagnosis” but received the most critical feedback in “prevention and treatment.” Although trends in performance differences were noted, they did not reach statistical significance.
← all excerpts
Evaluating the performance of large language models in sarcopenia-related patient queries: a foundational assessment for patient-centered validation.
1
—
—