When analyzed by question type, Gemini consistently showed higher stability across A1, A2, and A3/A4 questions, but these differences did not reach statistical significance compared to Claude or GPT-4o (Fig. 3 B-E).
← all excerpts
Large language models in Chinese anesthesiology residency examinations: a comparative analysis of performance, reliability and clinical reasoning.
1
—
—