Skip to main navigation Skip to search Skip to main content

Using sequence kernels to identify opinion entities in Urdu

  • SUNY Buffalo
  • Thomson Reuters

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

7 Scopus citations

Abstract

Automatic extraction of opinion holders and targets (together referred to as opinion entities) is an important subtask of sentiment analysis. In this work, we attempt to accurately extract opinion entities from Urdu newswire. Due to the lack of resources required for training role labelers and dependency parsers (as in English) for Urdu, a more robust approach based on (i) generating candidate word sequences corresponding to opinion entities, and (ii) subsequently disambiguating these sequences as opinion holders or targets is presented. Detecting the boundaries of such candidate sequences in Urdu is very different than in English since in Urdu, grammatical categories such as tense, gender and case are captured in word inflections. In this work, we exploit the morphological inflections associated with nouns and verbs to correctly identify sequence boundaries. Different levels of information that capture context are encoded to train standard linear and sequence kernels. To this end the best performance obtained for opinion entity detection for Urdu sentiment analysis is 58.06% F-Score using sequence kernels and 61.55% F-Score using a combination of sequence and linear kernels.

Original languageEnglish
Title of host publicationCoNLL 2011 - Fifteenth Conference on Computational Natural Language Learning, Proceedings of the Conference
PublisherAssociation for Computational Linguistics (ACL)
Pages58-67
Number of pages10
ISBN (Print)9781932432923
StatePublished - 2011
Event15th Conference on Computational Natural Language Learning, CoNLL 2011 - Portland, OR, United States
Duration: Jun 23 2011Jun 24 2011

Publication series

NameCoNLL 2011 - Fifteenth Conference on Computational Natural Language Learning, Proceedings of the Conference

Conference

Conference15th Conference on Computational Natural Language Learning, CoNLL 2011
Country/TerritoryUnited States
CityPortland, OR
Period06/23/1106/24/11

Fingerprint

Dive into the research topics of 'Using sequence kernels to identify opinion entities in Urdu'. Together they form a unique fingerprint.

Cite this