Computational identification of promoters and first exons in the human genome
- 26 November 2001
- journal article
- research article
- Published by Springer Nature in Nature Genetics
- Vol. 29 (4) , 412-417
- https://doi.org/10.1038/ng780
Abstract
The identification of promoters and first exons has been one of the most difficult problems in gene-finding. We present a set of discriminant functions that can recognize structural and compositional features such as CpG islands, promoter regions and first splice-donor sites. We explain the implementation of the discriminant functions into a decision tree that constitutes a new program called FirstEF. By using different models to predict CpG-related and non-CpG-related first exons, we showed by cross-validation that the program could predict 86% of the first exons with 17% false positives. We also demonstrated the prediction accuracy of FirstEF at the genome level by applying it to the finished sequences of human chromosomes 21 and 22 as well as by comparing the predictions with the locations of the experimentally verified first exons. Finally, we present the analysis of the predicted first exons for all of the 24 chromosomes of the human genome.Keywords
This publication has 23 references indexed in Scilit:
- The Sequence of the Human GenomeScience, 2001
- Making Sense of the SequenceScience, 2001
- Initial sequencing and analysis of the human genomeNature, 2001
- CART Classification of Human 5' UTR SequencesGenome Research, 2000
- Highly specific localization of promoter regions in large genomic sequences by PromoterInspector: a novel context analysis approachJournal of Molecular Biology, 2000
- Computational methods for the identification of genes in vertebrate genomic sequencesHuman Molecular Genetics, 1997
- Prediction of complete gene structures in human genomic DNAJournal of Molecular Biology, 1997
- Identification of protein coding regions in the human genome by quadratic discriminant analysisProceedings of the National Academy of Sciences, 1997
- The New Genomics: Global Views of BiologyScience, 1996
- Predicting internal exons by oligonucleotide composition and discriminant analysis of spliceable open reading framesNucleic Acids Research, 1994