Login / Signup

Pan4Draft: A Computational Tool to Improve the Accuracy of Pan-Genomic Analysis Using Draft Genomes.

Allan VerasFabricio AraujoKenny PinheiroLuis GuimarãesVasco AzevedoSiomar de Castro SoaresArtur Luiz da Costa da SilvaRommel Ramos
Published in: Scientific reports (2018)
High-throughput sequencing technologies are a milestone in molecular biology for facilitating great advances in genomics by enabling the deposit of large volumes of biological data to public databases. The availability of such data has made possible the comparative genomic analysis through pipelines, using the entire gene repertoire of genomes. However, a large number of unfinished genomes exist in public databases; their number is approximately 16-fold higher than the number of complete genomes, which creates bias during comparative analyses. Therefore, the present work proposes a new tool called Pan4Drafts, an automated pipeline for pan-genomic analysis of draft prokaryotic genomes to maximize the representation and accuracy of the gene repertoire of unfinished genomes by using reads from sequencing data. Pan4Draft allows to perform comparative analyses using different methodologies such as combining complete and draft genomes, using only draft genomes or only complete genomes. Pan4Draft is available at http://www.computationalbiology.ufpa.br/pan4drafts and the test dataset is available at https://sourceforge.net/projects/pan4drafts .
Keyphrases
  • high throughput sequencing
  • big data
  • healthcare
  • copy number
  • electronic health record
  • mental health
  • emergency department
  • single cell
  • gene expression
  • dna methylation
  • deep learning
  • artificial intelligence