highly significantp = 5.95 × 10 −145
Furthermore, Wilcoxon signed‐rank tests comparing the absolute prediction errors of RINAMI and the baseline model showed highly significant performance differences across the three Mega‐scale test subdatasets, with p ‐values below 0.01 in all splits ( p = 5.95 × 10 −145 , p < 1.00 × 10 −300 , and p = 1.39 × 10 −25 ).