TY - GEN
T1 - A lexicon reduction strategy in the context of handwritten medical forms
AU - Milewski, Robert
AU - Setlur, Srirangaraj
AU - Govindaraju, Venu
PY - 2005
Y1 - 2005
N2 - Traditional handwriting recognition algorithms rely heavily on small lexicons and clean word images. Unfortunately, emergency medical documents do not satisfy either of these conditions. This is a significant road-block that is hampering efforts to rapidly convert valuable offline healthcare handwriting data into digital content that can be efficiently mined for information. This paper describes a strategy whereby given an image representing a noisy handwritten word from a medical document, and a large lexicon consisting of English, medical and pharmacological words, symbols, abbreviations and acronyms, significantly reduces the size of the lexicon while keeping the unknown desired entry within the lexicon. The approach combines geometric interpretations of the word image along with contextual inference of concepts to reduce lexicons for word recognition. The data extracted can then be efficiently and securely disseminated for epidemiological and outbreak detection/analysis. Experimental results on NY State PCR forms are reported.
AB - Traditional handwriting recognition algorithms rely heavily on small lexicons and clean word images. Unfortunately, emergency medical documents do not satisfy either of these conditions. This is a significant road-block that is hampering efforts to rapidly convert valuable offline healthcare handwriting data into digital content that can be efficiently mined for information. This paper describes a strategy whereby given an image representing a noisy handwritten word from a medical document, and a large lexicon consisting of English, medical and pharmacological words, symbols, abbreviations and acronyms, significantly reduces the size of the lexicon while keeping the unknown desired entry within the lexicon. The approach combines geometric interpretations of the word image along with contextual inference of concepts to reduce lexicons for word recognition. The data extracted can then be efficiently and securely disseminated for epidemiological and outbreak detection/analysis. Experimental results on NY State PCR forms are reported.
UR - https://www.scopus.com/pages/publications/33947406476
U2 - 10.1109/ICDAR.2005.20
DO - 10.1109/ICDAR.2005.20
M3 - Conference contribution
AN - SCOPUS:33947406476
SN - 0769524206
SN - 9780769524207
T3 - Proceedings of the International Conference on Document Analysis and Recognition, ICDAR
SP - 1146
EP - 1150
BT - Proceedings of the Eighth International Conference on Document Analysis and Recognition
T2 - 8th International Conference on Document Analysis and Recognition
Y2 - 31 August 2005 through 1 September 2005
ER -