While improvements were observed in GPT-4o’s primary diagnosis and Claude 3.5 Sonnet’s top three differential diagnoses upon revealing the quiz context, these did not reach statistical significance.
← all excerpts
"This Is a Quiz" Premise Input: A Key to Unlocking Higher Diagnostic Accuracy in Large Language Models.
1
—
—