highly significantp = 1.6e-11
Statistical comparison with models trained with baseline, or week 2 data, show highly significant differences (p = 1.6e-11, p = 2.4e-13, respectively) although only an overall modest mean performance improvement of 3.1 and 3.4% respectively. 3.4.