Skip to main navigation Skip to search Skip to main content

Word level script identification for scanned document images

  • University of Maryland, College Park

Research output: Contribution to journalConference articlepeer-review

16 Scopus citations

Abstract

In this paper, we compare the performance of three classifiers used to identify the script of words in scanned document images. In both training and testing, a Gabor filter is applied and 16 channels of features are extracted. Three classifiers (Support Vector Machines (SVM), Gaussian Mixture Model (GMM) and k-Nearest-Neighbor (k-NN)) are used to identify different scripts at the word level (glyphs separated by white space). These three classifiers are applied to a variety of bilingual dictionaries and their performance is compared. Experimental results show the capability of Gabor filter to capture script features and the effectiveness of these three classifiers for script identification at the word level.

Original languageEnglish
Pages (from-to)124-135
Number of pages12
JournalProceedings of SPIE - The International Society for Optical Engineering
Volume5296
DOIs
StatePublished - 2004
EventDocument Recognition and Retrieval XI - San Jose, CA, United States
Duration: Jan 21 2004Jan 22 2004

Keywords

  • Gabor Filter
  • Gaussian Mixture Model (GMM)
  • K-Nearest-Neighbor (k-NN)
  • Script Identification
  • Support Vector Machines (SVM)

Fingerprint

Dive into the research topics of 'Word level script identification for scanned document images'. Together they form a unique fingerprint.

Cite this