While the differences between the other model pairs did not reach statistical significance ( p > 0.05), this marginal p-value suggests a potential trend toward performance variation between Gemini and DeepSeek that might be further elucidated with a larger dataset.
← all excerpts
Comparative performance of artificial intelligence models in intensive care nursing questions: an evaluation of ChatGPT, DeepSeek, and Google Gemini.
1
0.0500
0.0500