a nonsignificant trendp = 0.096
In Respiratory Medicine, OpenAI o1 performed significantly better than ChatGPT−4 ( p = 0.006), Gemini ( p = 0.016), and Copilot ( p = 0.003), while the comparison with DeepSeek showed only a nonsignificant trend ( p = 0.096); no other pairwise differences were significant.