However, we found no strong evidence for that, since neither the number nor the proportion of common amino acid sequences in our and the used tools’ data sets showed significant association with F 1 scores, although the proportion of common sequences showed a weakly significant association with MCC (see Supplementary Material Figs.
2
—
—
The sentences
Unexpectedly, there was a weak trend that when a query taxon was present in the training set of a tool, then F 1 and MCC scores tended to be slightly smaller (see Supplementary Material Fig.