highly significantp < 0.001
Authenticity of CoT responses by reasoning-capable LLMs For reasoning authenticity, the Kruskal-Wallis H test demonstrated highly significant differences in physician ratings among the four models ( p < 0.001).
Authenticity of CoT responses by reasoning-capable LLMs For reasoning authenticity, the Kruskal-Wallis H test demonstrated highly significant differences in physician ratings among the four models ( p < 0.001).
Although ChatGPT-O3 showed numerically higher accuracy than Grok-3, this difference did not reach statistical significance after BH correction ( p = 0.053).