Abstract
This paper presents a novel approach to defining document image structural similarity for the applications of classification and retrieval. We first build a codebook of SURF descriptors extracted from a set of representative training images. We then encode each document and model the spatial relationships between them by recursively partitioning the image and computing histograms of codewords in each partition. A random forest classifier is trained with the resulting features, and used for classification and retrieval. We demonstrate the effectiveness of our approach on table and tax form retrieval, and show that the proposed method outperforms previous approaches even when the training data is limited.
| Original language | English |
|---|---|
| Pages (from-to) | 119-126 |
| Number of pages | 8 |
| Journal | Pattern Recognition Letters |
| Volume | 43 |
| Issue number | 1 |
| DOIs | |
| State | Published - Jul 1 2014 |
Keywords
- Classification
- Random forest
- Retrieval
- Structural similarity
Fingerprint
Dive into the research topics of 'Structural similarity for document image classification and retrieval'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver