CRISPR sequences are sometimes erroneously translated and can contaminate public databases with spurious proteins containing spaced repeats.

Alejandro Rubio, Pablo Mier, Miguel A Andrade-Navarro, Andrés Garzón, Juan Jiménez, Antonio J Pérez-Pulido

Database : the journal of biological databases and curation 2020 Jan 01

filter terms:

The genomics era is resulting in the generation of a plethora of biological sequences that are usually stored in public databases. There are many computational tools that facilitate the annotation of these sequences, but sometimes they produce mistakes that enter the databases and can be propagated when erroneous data are used for secondary analyses, such as gene prediction or homology searching. While developing a computational gene finder based on protein-coding sequences, we discovered that the reference UniProtKB protein database is contaminated with some spurious sequences translated from DNA containing clustered regularly interspaced short palindromic repeats. We therefore encourage developers of prokaryotic computational gene finders and protein database curators to consider this source of error. © The Author(s) 2020. Published by Oxford University Press.

Citation

Alejandro Rubio, Pablo Mier, Miguel A Andrade-Navarro, Andrés Garzón, Juan Jiménez, Antonio J Pérez-Pulido. CRISPR sequences are sometimes erroneously translated and can contaminate public databases with spurious proteins containing spaced repeats. Database : the journal of biological databases and curation. 2020 Jan 01;2020

Mesh Tags

Substances

PMID: 33206958

View Full Text

FAQ

CRISPR sequences are sometimes erroneously translated and can contaminate public databases with spurious proteins containing spaced repeats.

filter terms:

Citation

var meshTagsSectionCollapsed = true; Mesh Tags

var substancesSectionCollapsed = true; Substances

Mesh Tags

Substances