An Algorithm for the DNA Sequence Generation from k-Tuple Word Contents of the Minimal Number of Random Fragments

1 April 1991

journal article
research article
Published by Taylor & Francis in Journal of Biomolecular Structure and Dynamics

Vol. 8 (5) , 1085-1102
https://doi.org/10.1080/07391102.1991.10507867

Abstract

An algorithm is described for generation of the long sequence written in a four letter alphabet from the constituent k-tuple words in the minimal number of separate, randomly defined fragments of the starting sequence. It is primarily intended for use in sequencing by hybridization (SBH) process- a potential method for sequencing human genome DNA (Drmanac et al., Genomics 4, pp. 114–128, 1989). The algorithm is based on the formerly defined rules and informative entities of the linear sequence. The algorithm requires neither knowledge on the number of appearances of a given k-tuple in sequence fragments, nor the information on which k-tuple words are on the ends of a fragment. It operates with the mixed content of k-tuples of the various lengths. The concept of the algorithm enables operations with the k-tuple sets containing false positive and false negative k-tuples. The content of the false k-tuples primarily affects the completeness of the generated sequence, and its correctness in the specific cases only. The algorithm can be used for the optimization of SBH parameters in the simulation experiments, as well as for the sequence generation in the real SBH experiments on the genomic DNA.

Keywords

This publication has 9 references indexed in Scilit:

Reliable Hybridization of Oligonucleotides as Short as Six Nucleotides
DNA and Cell Biology, 1990
Automated DNA sequencing of the human HPRT locus
Genomics, 1990
An oligonucleotide hybridization approach to DNA sequencing
FEBS Letters, 1989
l-Tuple DNA Sequencing: Computer Analysis
Journal of Biomolecular Structure and Dynamics, 1989
Sequencing of megabase plus DNA by hybridization: Theory of the method
Genomics, 1989
A novel method for nucleic acid sequence determination
Journal of Theoretical Biology, 1988
Molecular Approaches to Mammalian Genetics
Cold Spring Harbor Symposia on Quantitative Biology, 1986
DNA sequencing with chain-terminating inhibitors
Proceedings of the National Academy of Sciences, 1977
A new method for sequencing DNA.
Proceedings of the National Academy of Sciences, 1977