Skip to main navigation Skip to search Skip to main content

Adaptive word style classification using a Gaussian Mixture Model

  • University of Maryland, College Park

Research output: Contribution to journalConference articlepeer-review

5 Scopus citations

Abstract

In this paper, we present a new approach to detect bold and italic words in scanned documents. Under the assumption that OCR results are available, features used for classification are selected automatically using feature selection. For each scanned page, a Gaussian Mixture Model is constructed for characters with the same character code, and word styles are determined using a weighted majority vote. We applied this method to a variety of documents and compared the results with current commercial OCR software that provides style information. The experimental results show that our method performs better.

Original languageEnglish
Pages (from-to)606-609
Number of pages4
JournalProceedings - International Conference on Pattern Recognition
Volume2
StatePublished - 2004
EventProceedings of the 17th International Conference on Pattern Recognition, ICPR 2004 - Cambridge, United Kingdom
Duration: Aug 23 2004Aug 26 2004

Fingerprint

Dive into the research topics of 'Adaptive word style classification using a Gaussian Mixture Model'. Together they form a unique fingerprint.

Cite this