Genomic Classification Using an Information-Based Similarity Index: Application to the SARS Coronavirus
- 1 October 2005
- journal article
- research article
- Published by Mary Ann Liebert Inc in Journal of Computational Biology
- Vol. 12 (8) , 1103-1116
- https://doi.org/10.1089/cmb.2005.12.1103
Abstract
Measures of genetic distance based on alignment methods are confined to studying sequences that are conserved and identifiable in all organisms under study. A number of alignment-free techniques based on either statistical linguistics or information theory have been developed to overcome the limitations of alignment methods. We present a novel alignment-free approach to measuring the similarity among genetic sequences that incorporates elements from both word rank order-frequency statistics and information theory. We first validate this method on the human influenza A viral genomes as well as on the human mitochondrial DNA database. We then apply the method to study the origin of the SARS coronavirus. We find that the majority of the SARS genome is most closely related to group 1 coronaviruses, with smaller regions of matches to sequences from groups 2 and 3. The information based similarity index provides a new tool to measure the similarity between datasets based on their information content and may have a wide range of applications in the large-scale analysis of genomic databases.Keywords
This publication has 35 references indexed in Scilit:
- Phylogeny of the SARS CoronavirusScience, 2003
- Search for SARS Origins StallsScience, 2003
- Clues to the Animal Origins of SARSScience, 2003
- Identification of a Novel Coronavirus in Patients with Severe Acute Respiratory SyndromeNew England Journal of Medicine, 2003
- SWORDS: A statistical tool for analysing large DNA sequencesJournal of Biosciences, 2002
- Genome signature comparisons among prokaryote, plasmid, and mitochondrial DNAProceedings of the National Academy of Sciences, 1999
- Mitochondrial DNA and human evolutionNature, 1987
- Evolution of Human Influenza A Viruses Over 50 Years: Rapid, Uniform Rate of Change in NS GeneScience, 1986
- CONFIDENCE LIMITS ON PHYLOGENIES: AN APPROACH USING THE BOOTSTRAPEvolution, 1985
- Construction of Phylogenetic TreesScience, 1967