Genomic functional annotation using co-evolution profiles of gene clusters
Open Access
- 10 October 2002
- journal article
- research article
- Published by Springer Nature in Genome Biology
Abstract
The current speed of sequencing already exceeds the capability of annotation, creating a potential bottleneck. A large proportion of the genes in microbial genomes remains uncharacterized. Here we propose a new method for functional annotation using the conservation patterns of gene clusters. If several gene clusters show the same coevolution pattern across different genomes it is reasonable to infer they are functionally related. The gene cluster phylogenetic profile integrates chromosomal proximity information and phylogenetic profile information and allows us to infer functional dependences between the gene clusters even at great distance on the chromosome. As a proof of concept, we applied our method to the genome of Escherichia coli K12 strain. Our method establishes functional relationships among 176 gene clusters, comprising 738 E. coli genes. The accuracy of pair phylogenetic profiles was compared with the single-gene phylogenetic profile and was shown to be higher. As a result, we are able to suggest functional roles for several previously unknown genes or unknown genomic regions in E. coli. We also examined the robustness of coevolution signals across a larger set of genomes and suggest a possible upper limit of accuracy for the phylogenetic profile methods. The higher-order phylogenetic profiles, such as the gene-pair phylogenetic profiles, can detect functional dependences that are missed by using conventional single-gene phylogenetic profile or the chromosomal proximity method only. We show that the gene-pair phylogenetic profile is more accurate than the single-gene phylogenetic profiles.Keywords
This publication has 24 references indexed in Scilit:
- Intrinsic Lipid Preferences and Kinetic Mechanism of Escherichia coli MurGBiochemistry, 2002
- Pattern and Timing of Gene Duplication in Animal GenomesGenome Research, 2001
- The Pleiotropic Two-Component Regulatory System PhoP-PhoQJournal of Bacteriology, 2001
- Interim Report on Genomics of Escherichia ColiAnnual Review of Microbiology, 2000
- Detecting Protein Function and Protein-Protein Interactions from Genome SequencesScience, 1999
- Conservation of gene order: a fingerprint of proteins that physically interactPublished by Elsevier ,1998
- Constructing Multigenome Views of Whole Microbial GenomesMicrobial & Comparative Genomics, 1998
- A Genomic Perspective on Protein FamiliesScience, 1997
- Gapped BLAST and PSI-BLAST: a new generation of protein database search programsNucleic Acids Research, 1997
- Comparison of the phenotypes of thelpxAandlpxDmutants ofEscherichia coliFEMS Microbiology Letters, 1995