Promoter features related to tissue specificity as measured by Shannon entropy
Top Cited Papers
Open Access
- 29 March 2005
- journal article
- research article
- Published by Springer Nature in Genome Biology
- Vol. 6 (4) , 1-24
- https://doi.org/10.1186/gb-2005-6-4-r33
Abstract
Background: The regulatory mechanisms underlying tissue specificity are a crucial part of the development and maintenance of multicellular organisms. A genome-wide analysis of promoters in the context of gene-expression patterns in tissue surveys provides a means of identifying the general principles for these mechanisms. Results: We introduce a definition of tissue specificity based on Shannon entropy to rank human genes according to their overall tissue specificity and by their specificity to particular tissues. We apply our definition to microarray-based and expressed sequence tag (EST)-based expression data for human genes and use similar data for mouse genes to validate our results. We show that most genes show statistically significant tissue-dependent variations in expression level. We find that the most tissue-specific genes typically have a TATA box, no CpG island, and often code for extracellular proteins. As expected, CpG islands are found in most of the least tissue-specific genes, which often code for proteins located in the nucleus or mitochondrion. The class of genes with no CpG island or TATA box are the most common mid-specificity genes and commonly code for proteins located in a membrane. Sp1 was found to be a weak indicator of less-specific expression. YY1 binding sites, either as initiators or as downstream sites, were strongly associated with the least-specific genes. Conclusions: We have begun to understand the components of promoters that distinguish tissue-specific from ubiquitous genes, to identify associations that can predict the broad class of gene expression from sequence data alone.Keywords
This publication has 64 references indexed in Scilit:
- Applied bioinformatics for the identification of regulatory elementsNature Reviews Genetics, 2004
- The UCSC Table Browser data retrieval toolNucleic Acids Research, 2004
- Impact of Alternative Initiation, Splicing, and Termination on the Diversity of the mRNA Transcripts Encoded by the Mouse TranscriptomeGenome Research, 2003
- Initial sequencing and comparative analysis of the mouse genomeNature, 2002
- Splice Variation in Mouse Full-Length cDNAs Identified by Mapping to the Mouse GenomeGenome Research, 2002
- The Human Genome Browser at UCSCGenome Research, 2002
- Identification of regulatory regions which confer muscle-specific gene expressionJournal of Molecular Biology, 1998
- dbEST — database for “expressed sequence tags”Nature Genetics, 1993
- Sequence logos: a new way to display consensus sequencesNucleic Acids Research, 1990
- The “initiator” as a transcription control elementCell, 1989