De novo bacterial genome sequencing: Millions of very short reads assembled on a desktop computer
Top Cited Papers
- 10 March 2008
- journal article
- research article
- Published by Cold Spring Harbor Laboratory in Genome Research
- Vol. 18 (5) , 802-809
- https://doi.org/10.1101/gr.072033.107
Abstract
Novel high-throughput DNA sequencing technologies allow researchers to characterize a bacterial genome during a single experiment and at a moderate cost. However, the increase in sequencing throughput that is allowed by using such platforms is obtained at the expense of individual sequence read length, which must be assembled into longer contigs to be exploitable. This study focuses on the Illumina sequencing platform that produces millions of very short sequences that are 35 bases in length. We propose a de novo assembler software that is dedicated to process such data. Based on a classical overlap graph representation and on the detection of potentially spurious reads, our software generates a set of accurate contigs of several kilobases that cover most of the bacterial genome. The assembly results were validated by comparing data sets that were obtained experimentally for Staphylococcus aureus strain MW2 and Helicobacter acinonychis strain Sheeba with that of their published genomes acquired by conventional sequencing of 1.5- to 3.0-kb fragments. We also provide indications that the broad coverage achieved by high-throughput sequencing might allow for the detection of clonal polymorphisms in the set of DNA molecules being sequenced.Keywords
This publication has 31 references indexed in Scilit:
- SHARCGS, a fast and highly accurate short-read assembly algorithm for de novo genomic sequencingGenome Research, 2007
- Genome Analysis of Minibacterium massiliensis Highlights the Convergent Evolution of Water-Living BacteriaPLoS Genetics, 2007
- Whole-Genome Sequencing and Assembly with High-Throughput, Short-Read TechnologiesPLOS ONE, 2007
- Tracking the in vivo evolution of multidrug resistance in Staphylococcus aureus by whole-genome sequencingProceedings of the National Academy of Sciences, 2007
- Environmental Shotgun Sequencing: Its Potential and Challenges for Studying the Hidden World of MicrobesPLoS Biology, 2007
- Minimus: a fast, lightweight genome assemblerBMC Bioinformatics, 2007
- Genetic adaptation by Pseudomonas aeruginosa to the airways of cystic fibrosis patientsProceedings of the National Academy of Sciences, 2006
- Comparative Genomics of Multidrug Resistance in Acinetobacter baumanniiPLoS Genetics, 2006
- Combinatorial algorithms for DNA sequence assemblyAlgorithmica, 1995
- A New Algorithm for DNA Sequence AssemblyJournal of Computational Biology, 1995