Insights into corn genes derived from large-scale cDNA sequencing

Open Access

21 October 2008

journal article
research article
Published by Springer Nature in Plant Molecular Biology

Vol. 69 (1-2) , 179-194
https://doi.org/10.1007/s11103-008-9415-4

Abstract

We present a large portion of the transcriptome of Zea mays, including ESTs representing 484,032 cDNA clones from 53 libraries and 36,565 fully sequenced cDNA clones, out of which 31,552 clones are non-redundant. These and other previously sequenced transcripts have been aligned with available genome sequences and have provided new insights into the characteristics of gene structures and promoters within this major crop species. We found that although the average number of introns per gene is about the same in corn and Arabidopsis, corn genes have more alternatively spliced isoforms. Examination of the nucleotide composition of coding regions reveals that corn genes, as well as genes of other Poaceae (Grass family), can be divided into two classes according to the GC content at the third position in the amino acid encoding codons. Many of the transcripts that have lower GC content at the third position have dicot homologs but the high GC content transcripts tend to be more specific to the grasses. The high GC content class is also enriched with intronless genes. Together this suggests that an identifiable class of genes in plants is associated with the Poaceae divergence. Furthermore, because many of these genes appear to be derived from ancestral genes that do not contain introns, this evolutionary divergence may be the result of horizontal gene transfer from species not only with different codon usage but possibly that did not have introns, perhaps outside of the plant kingdom. By comparing the cDNAs described herein with the non-redundant set of corn mRNAs in GenBank, we estimate that there are about 50,000 different protein coding genes in Zea. All of the sequence data from this study have been submitted to DDBJ/GenBank/EMBL under accession numbers EU940701–EU977132 (FLI cDNA) and FK944382-FL482108 (EST).

Keywords

This publication has 38 references indexed in Scilit:

The Arabidopsis Information Resource (TAIR): gene structure and function annotation
Nucleic Acids Research, 2007
The Chlamydomonas Genome Reveals the Evolution of Key Animal and Plant Functions
Science, 2007
F-Box Proteins in Rice. Genome-Wide Analysis, Classification, Temporal and Spatial Gene Expression during Panicle and Seed Development, and Regulation by Light and Abiotic Stress
Plant Physiology, 2007
Rapid divergence of codon usage patterns within the rice genome
BMC Ecology and Evolution, 2007
Genome-wide transcriptional analysis of salinity stressed japonica and indica rice genotypes during panicle initiation stage
Plant Molecular Biology, 2006
The TIGR Rice Genome Annotation Resource: improvements and new features
Nucleic Acids Research, 2006
A genetic signature of interspecies variations in gene expression
Nature Genetics, 2006
Genomewide comparative analysis of alternative splicing in plants
Proceedings of the National Academy of Sciences, 2006
WebLogo: A Sequence Logo Generator: Figure 1
Genome Research, 2004
Simple cDNA normalization using kamchatka crab duplex-specific nuclease
Nucleic Acids Research, 2004