Abstract
In this paper we propose and implement a method for detecting duplicate documents in very large image databases. The method is based on a robust `signature' extracted from each document image which is used to index into a table of previously processed documents. The approach has a number of advantages over OCR or other recognition based methods including speed and robustness to imaging distortions. To justify the approach and test the scalability, we have developed a simulator which allows us to change parameters of the system and examine performance for millions of document signatures. A complete system is implemented and tested on a test collection of technical articles and memos.
| Original language | English |
|---|---|
| Pages | 314-318 |
| Number of pages | 5 |
| State | Published - 1997 |
| Event | Proceedings of the 1997 4th International Conference on Document Analysis and Recognition, ICDAR'97. Part 1 (of 2) - Ulm, Ger Duration: Aug 18 1997 → Aug 20 1997 |
Conference
| Conference | Proceedings of the 1997 4th International Conference on Document Analysis and Recognition, ICDAR'97. Part 1 (of 2) |
|---|---|
| City | Ulm, Ger |
| Period | 08/18/97 → 08/20/97 |
Fingerprint
Dive into the research topics of 'Detection of duplicates in document image databases'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver