Fast algorithms for large-scale genome alignment and comparison
Top Cited Papers
- 1 June 2002
- journal article
- Published by Oxford University Press (OUP) in Nucleic Acids Research
- Vol. 30 (11) , 2478-2483
- https://doi.org/10.1093/nar/30.11.2478
Abstract
We describe a suffix-tree algorithm that can align the entire genome sequences of eukaryotic and prokaryotic organisms with minimal use of computer time and memory. The new system, MUMmer 2, runs three times faster while using one-third as much memory as the original MUMmer system. It has been used successfully to align the entire human and mouse genomes to each other, and to align numerous smaller eukaryotic and prokaryotic genomes. A new module permits the alignment of multiple DNA sequence fragments, which has proven valuable in the comparison of incomplete genome sequences. We also describe a method to align more distantly related genomes by detecting protein sequence homology. This extension to MUMmer aligns two genomes after translating the sequence in all six reading frames, extracts all matching protein sequences and then clusters together matches. This method has been applied to both incomplete and complete genome sequences in order to detect regions of conserved synteny, in which multiple proteins from one organism are found in the same order and orientation in another. The system code is being made freely available by the authors.Keywords
This publication has 16 references indexed in Scilit:
- SSAHA: A Fast Search Method for Large DNA DatabasesGenome Research, 2001
- The Sequence of the Human GenomeScience, 2001
- Genome sequence of enterohaemorrhagic Escherichia coli O157:H7Nature, 2001
- Analysis of the genome sequence of the flowering plant Arabidopsis thalianaNature, 2000
- Evidence for symmetric chromosomal inversions around the replication origin in bacteriaGenome Biology, 2000
- Blocks-based methods for detecting protein homologyElectrophoresis, 2000
- PipMaker—A Web Server for Aligning Two Genomic DNA SequencesGenome Research, 2000
- Sequence and analysis of chromosome 2 of the plant Arabidopsis thalianaNature, 1999
- CAP3: A DNA Sequence Assembly ProgramGenome Research, 1999
- Alignment of whole genomesNucleic Acids Research, 1999