TY - GEN
T1 - Classification of crowdsourced text correction
AU - Gupta, Megha
AU - Dutta, Haimonti
AU - Geiger, Brian
N1 - Publisher Copyright:
Copyright 2015 ACM.
PY - 2015/3/18
Y1 - 2015/3/18
N2 - Optical Character Recognition (OCR) is a commonly used technique for digitizing printed material enabling them to be displayed online, searched and used in text mining applications. The text generated from OCR devices is often garbled due to variations in quality of the input paper, size and style of the font and column layout. This adversely affects retrieval effectiveness and hence techniques for cleaning the garbled text need to be improvised.This prototype system is expected to be deployed on historical newspaper archives that make extensive use of user text corrections.
AB - Optical Character Recognition (OCR) is a commonly used technique for digitizing printed material enabling them to be displayed online, searched and used in text mining applications. The text generated from OCR devices is often garbled due to variations in quality of the input paper, size and style of the font and column layout. This adversely affects retrieval effectiveness and hence techniques for cleaning the garbled text need to be improvised.This prototype system is expected to be deployed on historical newspaper archives that make extensive use of user text corrections.
UR - https://www.scopus.com/pages/publications/84958531177
U2 - 10.1145/2732587.2732619
DO - 10.1145/2732587.2732619
M3 - Conference contribution
AN - SCOPUS:84958531177
T3 - ACM International Conference Proceeding Series
SP - 142
EP - 143
BT - Proceedings of the 2nd ACM IKDD Conference on Data Sciences, CoDS 2015
PB - Association for Computing Machinery
T2 - 2nd ACM IKDD Conference on Data Sciences, CoDS 2015
Y2 - 18 March 2015 through 21 March 2015
ER -