highly significantp < 0.001
To evaluate the statistical significance of improvements, as shown in Table S7, a Friedman test was used and revealed highly significant overall differences among the models ( p < 0.001), followed by pairwise Wilcoxon signed-rank tests with Benjamini–Hochberg correction for multiple comparisons.