nominally significantP = 0.016
After correcting for multiple comparisons (Holm-Bonferroni), stepwise improvements in balanced metrics (F1 score and MCC) remained statistically robust across both internal and external cohorts, whereas the observed increases in AUC and sensitivity-focused F2 score—although nominally significant in certain cases (e.g., external cohort AUC nominal P = 0.016)—did not consistently survive multiple testing correction.