Abstract
Extracting and reading handwritten data from medical forms is an important task in medical informatics as it paves the way for efficient archival, indexing, and retrieval. This paper addresses two important challenges: (i) extraction of handwritten text data from images of carbon copies, and (ii) intelligent use of context to reduce lexicons to make the task of handwriting recognition tractable. We have developed a smart binarization algorithm targeted to carbon copy images that outperforms methods reported in the literature. The lexicon reduction method is based on learning the medical concept, and hence the probable medical terms to be encountered in the narrative part that describes the chief complaint of the patient by training on OCR output. In our experiments, we have worked with about 600 medical forms, 20 medical concepts, and a lexicon size of 4,700. We have observed that if the concept is one of top 3 choices, the lexicon can be reduced by two-thirds on an unseen form.
| Original language | English |
|---|---|
| Pages | 73-74 |
| Number of pages | 2 |
| DOIs | |
| State | Published - 2006 |
| Event | 7th Annual International Conference on Digital Government Research, Dg.o 2006 - San Diego, CA, United States Duration: May 21 2006 → May 24 2006 |
Conference
| Conference | 7th Annual International Conference on Digital Government Research, Dg.o 2006 |
|---|---|
| Country/Territory | United States |
| City | San Diego, CA |
| Period | 05/21/06 → 05/24/06 |
Keywords
- Medical Forms Processing
- OCR
Fingerprint
Dive into the research topics of 'Indexing and searching handwritten medical forms'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver