TY - GEN
T1 - Using sequence kernels to identify opinion entities in Urdu
AU - Mukund, Smruthi
AU - Ghosh, Debanjan
AU - Rohini, Rohini K.
PY - 2011
Y1 - 2011
N2 - Automatic extraction of opinion holders and targets (together referred to as opinion entities) is an important subtask of sentiment analysis. In this work, we attempt to accurately extract opinion entities from Urdu newswire. Due to the lack of resources required for training role labelers and dependency parsers (as in English) for Urdu, a more robust approach based on (i) generating candidate word sequences corresponding to opinion entities, and (ii) subsequently disambiguating these sequences as opinion holders or targets is presented. Detecting the boundaries of such candidate sequences in Urdu is very different than in English since in Urdu, grammatical categories such as tense, gender and case are captured in word inflections. In this work, we exploit the morphological inflections associated with nouns and verbs to correctly identify sequence boundaries. Different levels of information that capture context are encoded to train standard linear and sequence kernels. To this end the best performance obtained for opinion entity detection for Urdu sentiment analysis is 58.06% F-Score using sequence kernels and 61.55% F-Score using a combination of sequence and linear kernels.
AB - Automatic extraction of opinion holders and targets (together referred to as opinion entities) is an important subtask of sentiment analysis. In this work, we attempt to accurately extract opinion entities from Urdu newswire. Due to the lack of resources required for training role labelers and dependency parsers (as in English) for Urdu, a more robust approach based on (i) generating candidate word sequences corresponding to opinion entities, and (ii) subsequently disambiguating these sequences as opinion holders or targets is presented. Detecting the boundaries of such candidate sequences in Urdu is very different than in English since in Urdu, grammatical categories such as tense, gender and case are captured in word inflections. In this work, we exploit the morphological inflections associated with nouns and verbs to correctly identify sequence boundaries. Different levels of information that capture context are encoded to train standard linear and sequence kernels. To this end the best performance obtained for opinion entity detection for Urdu sentiment analysis is 58.06% F-Score using sequence kernels and 61.55% F-Score using a combination of sequence and linear kernels.
UR - https://www.scopus.com/pages/publications/84862286904
M3 - Conference contribution
AN - SCOPUS:84862286904
SN - 9781932432923
T3 - CoNLL 2011 - Fifteenth Conference on Computational Natural Language Learning, Proceedings of the Conference
SP - 58
EP - 67
BT - CoNLL 2011 - Fifteenth Conference on Computational Natural Language Learning, Proceedings of the Conference
PB - Association for Computational Linguistics (ACL)
T2 - 15th Conference on Computational Natural Language Learning, CoNLL 2011
Y2 - 23 June 2011 through 24 June 2011
ER -