Use of simulated data sets to evaluate the fidelity of metagenomic processing methods
- 29 April 2007
- journal article
- research article
- Published by Springer Nature in Nature Methods
- Vol. 4 (6) , 495-500
- https://doi.org/10.1038/nmeth1043
Abstract
Metagenomics is a rapidly emerging field of research for studying microbial communities. To evaluate methods presently used to process metagenomic sequences, we constructed three simulated data sets of varying complexity by combining sequencing reads randomly selected from 113 isolate genomes. These data sets were designed to model real metagenomes in terms of complexity and phylogenetic composition. We assembled sampled reads using three commonly used genome assemblers (Phrap, Arachne and JAZZ), and predicted genes using two popular gene-finding pipelines (fgenesb and CRITICA/GLIMMER). The phylogenetic origins of the assembled contigs were predicted using one sequence similarity–based (blast hit distribution) and two sequence composition–based (PhyloPythia, oligonucleotide frequencies) binning methods. We explored the effects of the simulated community structure and method combinations on the fidelity of each processing step by comparison to the corresponding isolate genomes. The simulated data sets are available online to facilitate standardized benchmarking of tools for metagenomic analysis. Please visit methagora to view and post comments on this articleKeywords
This publication has 24 references indexed in Scilit:
- Accurate phylogenetic classification of variable-length DNA fragmentsNature Methods, 2006
- Genomic analysis of the uncultivated marine crenarchaeote Cenarchaeum symbiosumProceedings of the National Academy of Sciences, 2006
- An experimental metagenome data management and analysis systemBioinformatics, 2006
- Deciphering the evolution and metabolism of an anammox bacterium from a community genomeNature, 2006
- Metagenomics: DNA sequencing of environmental samplesNature Reviews Genetics, 2005
- Environmental Genome Shotgun Sequencing of the Sargasso SeaScience, 2004
- Community structure and metabolism through reconstruction of microbial genomes from the environmentNature, 2004
- Improved microbial gene identification with GLIMMERNucleic Acids Research, 1999
- Gapped BLAST and PSI-BLAST: a new generation of protein database search programsNucleic Acids Research, 1997
- Dinucleotide relative abundance extremes: a genomic signatureTrends in Genetics, 1995