Abstract
In this paper the Damerau-Levenshtein string difference metric is generalized in two ways to more accurately compensate for the types of errors that are present in the script recognition domain. First, the basic dynamic programming method for computing such a measure is extended to allow for merges, splits and two-letter substitutions. Second, edit operations are refined into categories according to the effect they have on the visual "appearance" of words. A set of recognizer-independent constraints is developed to reflect the severity of the information lost due to each operation. These constraints are solved to assign specific costs to the operations. Experimental results on 2335 corrupted strings and a lexicon of 21,299 words show higher correcting rates than with the original form.
| Original language | English |
|---|---|
| Pages (from-to) | 405-414 |
| Number of pages | 10 |
| Journal | Pattern Recognition |
| Volume | 29 |
| Issue number | 3 |
| DOIs | |
| State | Published - Mar 1996 |
Keywords
- Post-processing
- Script recognition
- Spelling error correction
- String distance
- String matching
- Text editing
- Word recognition and correction
Fingerprint
Dive into the research topics of 'Generalizing edit distance to incorporate domain information: Handwritten text recognation as a case study'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver