The global trace graph, a novel paradigm for searching protein sequence databases
Open Access
- 15 September 2007
- journal article
- research article
- Published by Oxford University Press (OUP) in Bioinformatics
- Vol. 23 (18) , 2361-2367
- https://doi.org/10.1093/bioinformatics/btm358
Abstract
Motivation: Propagating functional annotations to sequence-similar, presumably homologous proteins lies at the heart of the bioinformatics industry. Correct propagation is crucially dependent on the accurate identification of subtle sequence motifs that are conserved in evolution. The evolutionary signal can be difficult to detect because functional sites may consist of non-contiguous residues while segments in-between may be mutated without affecting fold or function. Results: Here, we report a novel graph clustering algorithm in which all known protein sequences simultaneously self-organize into hypothetical multiple sequence alignments. This eliminates noise so that non-contiguous sequence motifs can be tracked down between extremely distant homologues. The novel data structure enables fast sequence database searching methods which are superior to profile-profile comparison at recognizing distant homologues. This study will boost the leverage of structural and functional genomics and opens up new avenues for data mining a complete set of functional signature motifs. Availability:http://www.bioinfo.biocenter.helsinki.fi/gtg Contact:liisa.holm@helsinki.fi Supplementary information: Supplementary data are available at Bioinformatics online.Keywords
This publication has 35 references indexed in Scilit:
- From sequences to a functional unitPhysiological Genomics, 2006
- Predicting protein function from sequence and structural dataPublished by Elsevier ,2005
- Amino acid substitution matrices from an information theoretic perspectivePublished by Elsevier ,2005
- Patterns and clusters within the PSM column in , 1992?2004Trends in Biochemical Sciences, 2004
- Single‐body residue‐level knowledge‐based energy score combined with sequence‐profile and secondary structure information for fold recognitionProteins-Structure Function and Bioinformatics, 2004
- The Pfam protein families databaseNucleic Acids Research, 2004
- FUGUE: sequence-structure homology recognition using environment-specific substitution tables and structure-dependent gap penalties11Edited by B. HonigJournal of Molecular Biology, 2001
- T-coffee: a novel method for fast and accurate multiple sequence alignment 1 1Edited by J. ThorntonJournal of Molecular Biology, 2000
- Identification of related proteins on family, superfamily and fold level 1 1Edited by F. C. CohenJournal of Molecular Biology, 2000
- Sequence comparisons using multiple sequences detect three times as many remote homologues as pairwise methodsJournal of Molecular Biology, 1998