TY - GEN
T1 - NE Tagging for Urdu based on Bootstrap POS Learning
AU - Mukund, Smruthi
AU - Srihari, Rohini K.
N1 - Publisher Copyright:
© 2009 Association for Computational Linguistics.
PY - 2009
Y1 - 2009
N2 - Part of Speech (POS) tagging and Named Entity (NE) tagging have become important components of effective text analysis. In this paper, we propose a bootstrapped model that involves four levels of text processing for Urdu. We show that increasing the training data for POS learning by applying bootstrapping techniques improves NE tagging results. Our model overcomes the limitation imposed by the availability of limited ground truth data required for training a learning model. Both our POS tagging and NE tagging models are based on the Conditional Random Field (CRF) learning approach. To further enhance the performance, grammar rules and lexicon lookups are applied on the final output to correct any spurious tag assignments. We also propose a model for word boundary segmentation where a bigram HMM model is trained for character transitions among all positions in each word. The generated words are further processed using a probabilistic language model. All models use a hybrid approach that combines statistical models with hand crafted grammar rules.
AB - Part of Speech (POS) tagging and Named Entity (NE) tagging have become important components of effective text analysis. In this paper, we propose a bootstrapped model that involves four levels of text processing for Urdu. We show that increasing the training data for POS learning by applying bootstrapping techniques improves NE tagging results. Our model overcomes the limitation imposed by the availability of limited ground truth data required for training a learning model. Both our POS tagging and NE tagging models are based on the Conditional Random Field (CRF) learning approach. To further enhance the performance, grammar rules and lexicon lookups are applied on the final output to correct any spurious tag assignments. We also propose a model for word boundary segmentation where a bigram HMM model is trained for character transitions among all positions in each word. The generated words are further processed using a probabilistic language model. All models use a hybrid approach that combines statistical models with hand crafted grammar rules.
UR - https://www.scopus.com/pages/publications/85160214399
M3 - Conference contribution
AN - SCOPUS:85160214399
T3 - NAACL HLT 2009 - 3rd International Workshop on Cross Lingual Information Access: Addressing the Information Need of Multilingual Societies, CLIAWS3 2009 - Proceedings of the Workshop
SP - 61
EP - 69
BT - NAACL HLT 2009 - 3rd International Workshop on Cross Lingual Information Access
A2 - Bandyopadhyay, Sivaji
A2 - Bhattacharyya, Pushpak
A2 - Varma, Vasudeva
A2 - Sarkar, Sudeshna
A2 - Kumaran, A
A2 - Udupa, Raghavendra
PB - Association for Computational Linguistics (ACL)
T2 - 3rd International Workshop on Cross Lingual Information Access: Addressing the Information Need of Multilingual Societies, CLIAWS3 2009
Y2 - 4 June 2009
ER -