Abstract
Applications for synchronous computer-mediated communication, i.e. chat, like instant messaging and chatroom channels, are playing an everincreasing role in task-oriented (vs. recreational) domains. Information extraction (IE) from chat could provide great value to any process where acting on a real-time information stream is important. However, no previous work exists on performing IE on chat data, and only limited research has been done on the linguistic differences between chat, spoken dialog, and written text. Chat likely poses several challenges for standard IE methods developed for heavily-edited written text, including: (i) surface-form noise, e.g. non-standard usage of punctuation; (ii) discourse-level noise, e.g. complex discourse structures that make resolution of the high frequency of context-dependent and anaphoric linguistic forms even more difficult. This paper describes an annotated corpus of taskoriented chat logs created in order to assess how the noise in chat data will affect the development of high accuracy IE technology for chat.
| Original language | English |
|---|---|
| Pages | 131-138 |
| Number of pages | 8 |
| State | Published - 2007 |
| Event | IJCAI 2007 Workshop on Analytics for Noisy Unstructured Text Data, AND 2007 - Hyderabad, India Duration: Jan 8 2007 → Jan 8 2007 |
Conference
| Conference | IJCAI 2007 Workshop on Analytics for Noisy Unstructured Text Data, AND 2007 |
|---|---|
| Country/Territory | India |
| City | Hyderabad |
| Period | 01/8/07 → 01/8/07 |
Fingerprint
Dive into the research topics of 'Information extraction for multi-participant, task-oriented, synchronous, computer-mediated communication: A corpus study of chat data'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver