highly significantp < 0.001
The Friedman test indicated highly significant differences between DeepSeek-R1 and both OpenAI o1 (p < 0.001) and Grok3 (p < 0.001), with no significant difference between OpenAI o1 and Grok3 (p = 1.000)( Figure 2a ).
The Friedman test indicated highly significant differences between DeepSeek-R1 and both OpenAI o1 (p < 0.001) and Grok3 (p < 0.001), with no significant difference between OpenAI o1 and Grok3 (p = 1.000)( Figure 2a ).