Repeating this Infernal search with this kind of model (built from a new “seed” alignment created by simply concatenating two copies of the original Rfam “seed”) finds highly significant (E < 10 −10 ) tandem glycine structures that completely cover all 11 RNA-PATTERN hits missed by the original Rfam model.) Can we trust that the statistically significant matches to the CM are really homologs, and that increased numbers of predictions really reflect increased detection sensitivity?
← all excerpts
Computational identification of functional RNA homologs in metagenomic data.
2
—
—
The sentences
This increase in resolution doesn’t matter much if a sequence is already readily detected by primary sequence comparison (improving an already significant E-value of 10 −30 to 10 −33 , for example), but it becomes important when lifting a marginally insignificant E-value to significance (0.1 to 10 −4 , for example).