Skip to main navigation Skip to search Skip to main content

Morphology induction from limited noisy data using approximate string matching

  • University of Maryland, College Park

Research output: Contribution to conferencePaperpeer-review

3 Scopus citations

Abstract

For a language with limited resources, a dictionary may be one of the few available electronic resources. To make effective use of the dictionary for translation, however, users must be able to access it using the root form of morphologically deformed variant found in the text. Stemming and data driven methods, however, are not suitable when data is sparse. We present algorithms for discovering morphemes from limited, noisy data obtained by scanning a hard copy dictionary. Our approach is based on the novel application of the longest common substring and string edit distance metrics. Results show that these algorithms can in fact segment words into roots and affixes from the limited data contained in a dictionary, and extract affixes. This in turn allows non native speakers to perform multilingual tasks for applications where response must be rapid, and their knowledge is limited. In addition, this analysis can feed other NLP tools requiring lexicons.

Original languageEnglish
Pages60-68
Number of pages9
DOIs
StatePublished - 2006
Event8th Meeting of the ACL Special Interest Group on Computational Phonology, SIGPHON 2006, collocated with the HLT-NAACL 2006 - New York City, United States
Duration: Jun 8 2006 → …

Conference

Conference8th Meeting of the ACL Special Interest Group on Computational Phonology, SIGPHON 2006, collocated with the HLT-NAACL 2006
Country/TerritoryUnited States
CityNew York City
Period06/8/06 → …

Fingerprint

Dive into the research topics of 'Morphology induction from limited noisy data using approximate string matching'. Together they form a unique fingerprint.

Cite this