Skip to main navigation Skip to search Skip to main content

NE Tagging for Urdu based on Bootstrap POS Learning

  • SUNY Buffalo

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

16 Scopus citations

Abstract

Part of Speech (POS) tagging and Named Entity (NE) tagging have become important components of effective text analysis. In this paper, we propose a bootstrapped model that involves four levels of text processing for Urdu. We show that increasing the training data for POS learning by applying bootstrapping techniques improves NE tagging results. Our model overcomes the limitation imposed by the availability of limited ground truth data required for training a learning model. Both our POS tagging and NE tagging models are based on the Conditional Random Field (CRF) learning approach. To further enhance the performance, grammar rules and lexicon lookups are applied on the final output to correct any spurious tag assignments. We also propose a model for word boundary segmentation where a bigram HMM model is trained for character transitions among all positions in each word. The generated words are further processed using a probabilistic language model. All models use a hybrid approach that combines statistical models with hand crafted grammar rules.

Original languageEnglish
Title of host publicationNAACL HLT 2009 - 3rd International Workshop on Cross Lingual Information Access
Subtitle of host publicationAddressing the Information Need of Multilingual Societies, CLIAWS3 2009 - Proceedings of the Workshop
EditorsSivaji Bandyopadhyay, Pushpak Bhattacharyya, Vasudeva Varma, Sudeshna Sarkar, A Kumaran, Raghavendra Udupa
PublisherAssociation for Computational Linguistics (ACL)
Pages61-69
Number of pages9
ISBN (Electronic)9781932432336
StatePublished - 2009
Event3rd International Workshop on Cross Lingual Information Access: Addressing the Information Need of Multilingual Societies, CLIAWS3 2009 - Boulder, United States
Duration: Jun 4 2009 → …

Publication series

NameNAACL HLT 2009 - 3rd International Workshop on Cross Lingual Information Access: Addressing the Information Need of Multilingual Societies, CLIAWS3 2009 - Proceedings of the Workshop

Conference

Conference3rd International Workshop on Cross Lingual Information Access: Addressing the Information Need of Multilingual Societies, CLIAWS3 2009
Country/TerritoryUnited States
CityBoulder
Period06/4/09 → …

Fingerprint

Dive into the research topics of 'NE Tagging for Urdu based on Bootstrap POS Learning'. Together they form a unique fingerprint.

Cite this