Skip to main navigation Skip to search Skip to main content

Information extraction for multi-participant, task-oriented, synchronous, computer-mediated communication: A corpus study of chat data

  • Janya Inc.

Research output: Contribution to conferencePaperpeer-review

3 Scopus citations

Abstract

Applications for synchronous computer-mediated communication, i.e. chat, like instant messaging and chatroom channels, are playing an everincreasing role in task-oriented (vs. recreational) domains. Information extraction (IE) from chat could provide great value to any process where acting on a real-time information stream is important. However, no previous work exists on performing IE on chat data, and only limited research has been done on the linguistic differences between chat, spoken dialog, and written text. Chat likely poses several challenges for standard IE methods developed for heavily-edited written text, including: (i) surface-form noise, e.g. non-standard usage of punctuation; (ii) discourse-level noise, e.g. complex discourse structures that make resolution of the high frequency of context-dependent and anaphoric linguistic forms even more difficult. This paper describes an annotated corpus of taskoriented chat logs created in order to assess how the noise in chat data will affect the development of high accuracy IE technology for chat.

Original languageEnglish
Pages131-138
Number of pages8
StatePublished - 2007
EventIJCAI 2007 Workshop on Analytics for Noisy Unstructured Text Data, AND 2007 - Hyderabad, India
Duration: Jan 8 2007Jan 8 2007

Conference

ConferenceIJCAI 2007 Workshop on Analytics for Noisy Unstructured Text Data, AND 2007
Country/TerritoryIndia
CityHyderabad
Period01/8/0701/8/07

Fingerprint

Dive into the research topics of 'Information extraction for multi-participant, task-oriented, synchronous, computer-mediated communication: A corpus study of chat data'. Together they form a unique fingerprint.

Cite this