highly significantp < 0.0001
However, the performance gap between the leading model (DeepSeek) and the baseline model (ChatGPT-3.5 Turbo) was highly significant ( p < 0.0001), confirming substantial advancements in newer models.
However, the performance gap between the leading model (DeepSeek) and the baseline model (ChatGPT-3.5 Turbo) was highly significant ( p < 0.0001), confirming substantial advancements in newer models.