highly significantP <.05
In contrast, for N staging, highly significant differences ( P <.05) were consistently demonstrated across both test sets, indicating statistically meaningful performance disparities.
In contrast, for N staging, highly significant differences ( P <.05) were consistently demonstrated across both test sets, indicating statistically meaningful performance disparities.
In the black-box evaluation, where the difference did not reach statistical significance ( P =.10), the medium-to-large effect size (Cohen ω=0.399) suggests a potential trend, which points to a possible performance difference in detecting distant metastasis evidence in model behavior.