Skip to main navigation Skip to search Skip to main content

Form classification

  • SUNY Buffalo

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

10 Scopus citations

Abstract

The problem of form classification is to assign a single-page form image to one of a set of predefined form types or classes. We classify the form images using low level pixel density information from the binary images of the documents. In this paper, we solve the form classification problem with a classifier based on the k-means algorithm, supported by adaptive boosting. Our classification method is tested on the NIST scanned tax forms data bases (special forms databases 2 and 6) which include machine-typed and handwritten documents. Our method improves the performance over published results on the same databases, while still using a simple set of image features.

Original languageEnglish
Title of host publicationDocument Recognition and Retrieval XV
DOIs
StatePublished - 2008
EventDocument Recognition and Retrieval XV - San Jose, CA, United States
Duration: Jan 29 2008Jan 31 2008

Publication series

NameProceedings of SPIE - The International Society for Optical Engineering
Volume6815
ISSN (Print)0277-786X

Conference

ConferenceDocument Recognition and Retrieval XV
Country/TerritoryUnited States
CitySan Jose, CA
Period01/29/0801/31/08

Keywords

  • AdaBoost
  • Document image classification
  • Form classification
  • Image-level features
  • NIST tax forms datasets

Fingerprint

Dive into the research topics of 'Form classification'. Together they form a unique fingerprint.

Cite this