The Friedman test results indicated statistically significant between model differences for ARI, CLI, OLWF, LWGLF, and FRF (all p < 0.01), whereas GFI did not reach statistical significance ( p = 0.064).
← all excerpts
Can large language models be trusted? Reliability and readability of responses to perinatal depression FAQs.
1
0.0640
0.0640