highly significantp ≤ 0.001
This performance difference remained highly significant in the randomized k-fold evaluation, where significant differences were again observed for both the low [χ 2 (3) = 135.05, p ≤ 0.001) and high [χ 2 (3) = 162.84, p ≤ 0.001] class-separability contrasts. 4 Discussion The objective of the current study was to investigate the extent to which pBCI model evaluation metrics may be biased when temporal dependencies between train and test samples are not considered in cross-validation.