Skip to main navigation Skip to search Skip to main content

Unsupervised classification of structurally similar document images

  • University of Maryland, College Park

Research output: Contribution to journalConference articlepeer-review

38 Scopus citations

Abstract

In this paper, we present a learning based approach for computing structural similarities among document images for unsupervised exploration in large document collections. The approach is based on multiple levels of content and structure. At a local level, a bag-of-visual words based on SURF features provides an effective way of computing content similarity. The document is then recursively partitioned and a histogram of codewords is computed for each partition. Structural similarity is computed using a random forest classifier trained with these histogram features. We experiment with three diverse datasets of document images varying in size, degree of structural similarity, and types of document images. Our results demonstrate that the proposed approach provides an effective general framework for grouping structurally similar document images.

Original languageEnglish
Article number6628809
Pages (from-to)1225-1229
Number of pages5
JournalProceedings of the International Conference on Document Analysis and Recognition, ICDAR
DOIs
StatePublished - 2013
Event12th International Conference on Document Analysis and Recognition, ICDAR 2013 - Washington, DC, United States
Duration: Aug 25 2013Aug 28 2013

Keywords

  • Clustering
  • Random forest
  • Structural similarity
  • Unsupervised classification

Fingerprint

Dive into the research topics of 'Unsupervised classification of structurally similar document images'. Together they form a unique fingerprint.

Cite this