Skip to main navigation Skip to search Skip to main content

Detection of duplicates in document image databases

  • University of Maryland, College Park

Research output: Contribution to conferencePaperpeer-review

25 Scopus citations

Abstract

In this paper we propose and implement a method for detecting duplicate documents in very large image databases. The method is based on a robust `signature' extracted from each document image which is used to index into a table of previously processed documents. The approach has a number of advantages over OCR or other recognition based methods including speed and robustness to imaging distortions. To justify the approach and test the scalability, we have developed a simulator which allows us to change parameters of the system and examine performance for millions of document signatures. A complete system is implemented and tested on a test collection of technical articles and memos.

Original languageEnglish
Pages314-318
Number of pages5
StatePublished - 1997
EventProceedings of the 1997 4th International Conference on Document Analysis and Recognition, ICDAR'97. Part 1 (of 2) - Ulm, Ger
Duration: Aug 18 1997Aug 20 1997

Conference

ConferenceProceedings of the 1997 4th International Conference on Document Analysis and Recognition, ICDAR'97. Part 1 (of 2)
CityUlm, Ger
Period08/18/9708/20/97

Fingerprint

Dive into the research topics of 'Detection of duplicates in document image databases'. Together they form a unique fingerprint.

Cite this