Quality scores were numerically higher for Claude Opus 4.1 at 4.22 ± 0.62, followed by Gemini 2.5 Flash at 4.19 ± 0.67 and ChatGPT-5 at 4.10 ± 0.76, although this difference did not reach statistical significance.
← all excerpts
AI at the Sella Turcica: Multi-Model Large Language Model Evaluation in Pituitary Adenomas.
1
—
—