Pairwise comparisons of model accuracy produced confidence intervals that included zero, and permutation testing showed higher median F1 scores for Model 3 that did not reach statistical significance (all p > 0.97).
← all excerpts
Evaluating GPT-4o for emergency disposition of complex respiratory cases with pulmonology consultation: a diagnostic accuracy study.
1
0.9700
0.9700