Abstract
Searching for documents by their type or genre is a natural way to enhance the effectiveness of document retrieval. The layout of a document contains a significant amount of information that can be used to classify a document's type in the absence of domain specific models. A document type or genre can be defined by the user based primarily on layout structure. Our classification approach is based on `visual similarity' of the layout structure by building a supervised classifier, given examples of the class. We use image features, such as the percentages of text and non-text (graphics, image, table, and ruling) content regions, column structures, variations in the point sizes of fonts, the density of content area, and various statistics on features of connected components which can be derived from class samples without class knowledge. In order to obtain class labels for training samples, we conducted a user relevance test where subjects ranked UW-I document images with respect to the 12 representative images. We implemented our classification scheme using the OC1, a decision tree classifier, and report our findings.
| Original language | English |
|---|---|
| Pages (from-to) | 182-190 |
| Number of pages | 9 |
| Journal | Proceedings of SPIE - The International Society for Optical Engineering |
| Volume | 3967 |
| State | Published - 2000 |
| Event | Proceedings of the 2000 Document Recognition and Retrieval VII - San Jose, CA, USA Duration: Jan 26 2000 → Jan 27 2000 |
Fingerprint
Dive into the research topics of 'Classification of document page images based on visual similarity of layout structures'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver