In expert ratings, only 1 comparison failed to reach statistical significance: ChatGPT-4o versus Zhipu Qingyan ( r =0.455; P adjusted=.64); all other pairwise comparisons reached significance ( Multimedia Appendix 7 ).
← all excerpts
Performance of 5 Large Language Models in Perioperative Consultation for Pediatric Hypospadias: Cross-Sectional Comparative Study.
3
—
—
The sentences
On the expert side (N=23), professional seniority modulated discrimination most strongly: Gemini-2.5-Pro was rated monotonically higher with increasing seniority (Junior median 4/Intermediate median 5/Senior median 5; Kruskal-Wallis P =.004; P trend <.001), OpenEvidence monotonically lower (Junior median 2/Intermediate median 2/Senior median 1; Kruskal-Wallis P =.026; P trend=.007), and ChatGPT-4o showed a significant trend across seniority ( P trend=.04) despite a nonsignificant omnibus test.
Subgroup analysis showed a monotonic decline in OpenEvidence ratings as caregiver income rose ( P trend =.033), with a marginal trend in the opposite direction for Gemini-2.5-Pro; higher-income caregivers, who may have greater health literacy [ 33 , 34 ] or stronger expectations about delivery style, appear less tolerant of OpenEvidence’s academic phrasing.