µgreen-db: a reference database for the 23S rRNA gene of eukaryotic plastids and cyanobacteria.
Christophe DjemielDamien PlassardSébastien TerratOlivier CrouzetJoana SauzeSamuel MondyVirginie NowakLisa WingateJerome OgeePierre-Alain MaronPublished in: Scientific reports (2020)
Studying the ecology of photosynthetic microeukaryotes and prokaryotic cyanobacterial communities requires molecular tools to complement morphological observations. These tools rely on specific genetic markers and require the development of specialised databases to achieve taxonomic assignment. We set up a reference database, called µgreen-db, for the 23S rRNA gene. The sequences were retrieved from generalist (NCBI, SILVA) or Comparative RNA Web (CRW) databases, in addition to a more original approach involving recursive BLAST searches to obtain the best possible sequence recovery. At present, µgreen-db includes 2,326 23S rRNA sequences belonging to both eukaryotes and prokaryotes encompassing 442 unique genera and 736 species of photosynthetic microeukaryotes, cyanobacteria and non-vascular land plants based on the NCBI and AlgaeBase taxonomy. When PR2/SILVA taxonomy is used instead, µgreen-db contains 2,217 sequences (399 unique genera and 696 unique species). Using µgreen-db, we were able to assign 96% of the sequences of the V domain of the 23S rRNA gene obtained by metabarcoding after amplification from soil DNA at the genus level, highlighting good coverage of the database. µgreen-db is accessible at http://microgreen-23sdatabase.ea.inra.fr.