Abstract
In this paper, we present a learning based approach for computing structural similarities among document images for unsupervised exploration in large document collections. The approach is based on multiple levels of content and structure. At a local level, a bag-of-visual words based on SURF features provides an effective way of computing content similarity. The document is then recursively partitioned and a histogram of codewords is computed for each partition. Structural similarity is computed using a random forest classifier trained with these histogram features. We experiment with three diverse datasets of document images varying in size, degree of structural similarity, and types of document images. Our results demonstrate that the proposed approach provides an effective general framework for grouping structurally similar document images.
| Original language | English |
|---|---|
| Article number | 6628809 |
| Pages (from-to) | 1225-1229 |
| Number of pages | 5 |
| Journal | Proceedings of the International Conference on Document Analysis and Recognition, ICDAR |
| DOIs | |
| State | Published - 2013 |
| Event | 12th International Conference on Document Analysis and Recognition, ICDAR 2013 - Washington, DC, United States Duration: Aug 25 2013 → Aug 28 2013 |
Keywords
- Clustering
- Random forest
- Structural similarity
- Unsupervised classification
Fingerprint
Dive into the research topics of 'Unsupervised classification of structurally similar document images'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver