Statistical significance of hierarchical multi‐body potentials based on Delaunay tessellation and their application in sequence‐structure alignment
Open Access
- 1 July 1997
- journal article
- research article
- Published by Wiley in Protein Science
- Vol. 6 (7) , 1467-1481
- https://doi.org/10.1002/pro.5560060711
Abstract
Statistical potentials based on pairwise interactions between Cα atoms are commonly used in protein threading/fold‐recognition attempts. Inclusion of higher order interaction is a possible means of improving the specificity of these potentials. Delaunay tessellation of the Cα‐atom representation of protein structure has been suggested as a means of defining multi‐body interactions.A large number of parameters are required to define all four‐body interactions of 20 amino acid types (204 = 160,000). Assuming that residue order within a four‐body contact is irrelevant reduces this to a manageable 8,855 parameters, using a nonredundant dataset of 608 protein structures.Three lines of evidence support the significance and utility of the four‐body potential for sequence‐structure matching. First, compared to the four‐body model, all lower‐order interaction models (three‐body, two‐body, one‐body) are found statistically inadequate to explain the frequency distribution of residue contacts.Second, coherent patterns of interaction are seen in a graphic presentation of the four‐body potential. Many patterns have plausible biophysical explanations and are consistent across sets of residues sharing certain properties (e.g., size, hydrophobicity, or charge).Third, the utility of the multi‐body potential is tested on a test set of 12 same‐length pairs of proteins of known structure for two protocols: Sequence‐recognizes‐structure, where a query sequence is threaded (without gap) through the native and a non‐native structure; and structure‐recognizes‐sequence, where a query structure is threaded by its native and another non‐native sequence. Using cross‐validated training, protein sequences correctly recognized their native structure in all 24 cases. Conversely, structures recognized the native sequence in 23 of 24 cases. Further, the score differences between correct and decoy structures increased significantly using the three‐ or four‐body potential compared to potentials of lower order.Keywords
This publication has 18 references indexed in Scilit:
- The quickhull algorithm for convex hullsACM Transactions on Mathematical Software, 1996
- Delaunay Tessellation of Proteins: Four Body Nearest-Neighbor Propensities of Amino Acid ResiduesJournal of Computational Biology, 1996
- Knowledge-based potentials for proteinsCurrent Opinion in Structural Biology, 1995
- Enlarged representative set of protein structuresProtein Science, 1994
- Estimation of the maximum change in stability of globular proteins upon mutation of a hydrophobic residue to another of smaller sizeProtein Science, 1993
- Generating and testing protein foldsCurrent Opinion in Structural Biology, 1993
- Topology fingerprint approach to the inverse protein folding problemJournal of Molecular Biology, 1992
- Evaluation of protein models by atomic solvation preferenceJournal of Molecular Biology, 1992
- The role of internal packing interactions in determining the structure and stability of a proteinJournal of Molecular Biology, 1991
- Estimation of effective interresidue contact energies from protein crystal structures: quasi-chemical approximationMacromolecules, 1985