Skip to main navigation Skip to search Skip to main content

Classification of crowdsourced text correction

  • Indraprastha Institute of Information Technology Delhi
  • University of California at Riverside
  • Columbia University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Optical Character Recognition (OCR) is a commonly used technique for digitizing printed material enabling them to be displayed online, searched and used in text mining applications. The text generated from OCR devices is often garbled due to variations in quality of the input paper, size and style of the font and column layout. This adversely affects retrieval effectiveness and hence techniques for cleaning the garbled text need to be improvised.This prototype system is expected to be deployed on historical newspaper archives that make extensive use of user text corrections.

Original languageEnglish
Title of host publicationProceedings of the 2nd ACM IKDD Conference on Data Sciences, CoDS 2015
PublisherAssociation for Computing Machinery
Pages142-143
Number of pages2
ISBN (Electronic)9781450334365
DOIs
StatePublished - Mar 18 2015
Event2nd ACM IKDD Conference on Data Sciences, CoDS 2015 - Bangalore, India
Duration: Mar 18 2015Mar 21 2015

Publication series

NameACM International Conference Proceeding Series
Volume18-21-March-2015

Conference

Conference2nd ACM IKDD Conference on Data Sciences, CoDS 2015
Country/TerritoryIndia
CityBangalore
Period03/18/1503/21/15

Fingerprint

Dive into the research topics of 'Classification of crowdsourced text correction'. Together they form a unique fingerprint.

Cite this