Conrad: Gene prediction using conditional random fields
Open Access
- 9 August 2007
- journal article
- research article
- Published by Cold Spring Harbor Laboratory in Genome Research
- Vol. 17 (9) , 1389-1398
- https://doi.org/10.1101/gr.6558107
Abstract
We present Conrad, the first comparative gene predictor based on semi-Markov conditional random fields (SMCRFs). Unlike the best standalone gene predictors, which are based on generalized hidden Markov models (GHMMs) and trained by maximum likelihood, Conrad is discriminatively trained to maximize annotation accuracy. In addition, unlike the best annotation pipelines, which rely on heuristic and ad hoc decision rules to combine standalone gene predictors with additional information such as ESTs and protein homology, Conrad encodes all sources of information as features and treats all features equally in the training and inference algorithms. Conrad outperforms the best standalone gene predictors in cross-validation and whole chromosome testing on two fungi with vastly different gene structures. The performance improvement arises from the SMCRF’s discriminative training methods and their ability to easily incorporate diverse types of information by encoding them as feature functions. OnCryptococcus neoformans, configuring Conrad to reproduce the predictions of a two-species phylo-GHMM closely matches the performance of Twinscan. Enabling discriminative training increases performance, and adding new feature functions further increases performance, achieving a level of accuracy that is unprecedented for this organism. Similar results are obtained onAspergillus nidulanscomparing Conrad versus Fgenesh. SMCRFs are a promising framework for gene prediction because of their highly modular nature, simplifying the process of designing and testing potential indicators of gene structure. Conrad’s implementation of SMCRFs advances the state of the art in gene prediction in fungi and provides a robust platform for both current application and future research.Keywords
This publication has 31 references indexed in Scilit:
- Global Discriminative Learning for Higher-Accuracy Computational Gene PredictionPLoS Computational Biology, 2007
- Using Multiple Alignments to Improve Gene PredictionJournal of Computational Biology, 2006
- ExonHunter: a comprehensive approach to gene findingBioinformatics, 2005
- Gene prediction and verification in a compact genome with numerous small intronsGenome Research, 2004
- Multiple-sequence functional annotation and the generalized hidden Markov phylogenyBioinformatics, 2004
- Computational Gene Prediction Using Multiple Sources of EvidenceGenome Research, 2004
- Sequencing and comparison of yeast species to identify genes and regulatory elementsNature, 2003
- GAZE: A Generic Framework for the Integration of Gene-Prediction Data by Dynamic ProgrammingGenome Research, 2002
- Computational Inference of Homologous Gene Structures in the Human GenomeGenome Research, 2001
- Prediction of complete gene structures in human genomic DNAJournal of Molecular Biology, 1997