A technique for isolating differences between files

1 April 1978

journal article
Published by Association for Computing Machinery (ACM) in Communications of the ACM

Vol. 21 (4) , 264-268
https://doi.org/10.1145/359460.359467

Abstract

A simple algorithm is described for isolating the differences between two files. One application is the comparing of two versions of a source program or other file in order to display all differences. The algorithm isolates differences in a way that corresponds closely to our intuitive notion of difference, is easy to implement, and is computationally efficient, with time linear in the file length. For most applications the algorithm isolates differences similar to those isolated by the longest common subsequence. Another application of this algorithm merges files containing independently generated changes into a single file. The algorithm can also be used to generate efficient encodings of a file in the form of the differences between itself and a given “datum” file, permitting reconstruction of the original file from the diference and datum files.

Keywords

This publication has 5 references indexed in Scilit:

Bounds on the Complexity of the Longest Common Subsequence Problem
Journal of the ACM, 1976
A linear space algorithm for computing maximal common subsequences
Communications of the ACM, 1975
The String-to-String Correction Problem
Journal of the ACM, 1974
WYLBUR
Communications of the ACM, 1973
An online editor
Communications of the ACM, 1967