157.78), the difference did not reach statistical significance ( U = 12339.00, z = –1.594, P = 0.111).
← all excerpts
Evaluating Large Language Models in Ptosis-Related inquiries: A Cross-Lingual Study.
2
0.1110
0.1110
The sentences
showed a trendP = 0.111
Total Scores: Although GPT-4o showed a trend toward better overall performance (mean rank = 173.22 vs. 157.78), the difference did not reach statistical significance ( U = 12339.00, z = –1.594, P = 0.111).