The comparison between Gemini 3 Pro and DeepSeek-V3.1 approached but did not reach statistical significance after correction (adjusted P = 0.063), and the remaining pairwise comparisons were not statistically significant ( Table 4 ; Figure 5C ).
← all excerpts
Risk-centered benchmarking of large language models for AI-enabled counseling in chronic autoimmune thyroid eye disease.
1
0.0630
0.0630