highly significantp < 0.0001
Here, the baseline model demonstrated the expected strong association between internal data redundancy and model performance, with a substantial and highly significant ( p < 0.0001, bootstrap test with 10,000 replications) drop in performance as the partitioning threshold is decreased (from an AUC value of 0.67 at 99% to 0.63 at 90%)—resulting in a lower similarity between the training and test data sets.