Skip to main navigation Skip to search Skip to main content

A bootstrapping approach to information extraction domain porting

  • Cymfony Inc

Research output: Contribution to conferencePaperpeer-review

Abstract

This paper presents a seed-driven, bootstrapping approach to domain porting that could be used to customize a generic information extraction (IE) capability for a specific domain. The approach taken is based on the existence of a robust, domain-independent IE engine that can continue to be enhanced, independent of any particular domain. This approach combines the strengths of parsing-based symbolic rule learning and the high performance linear string-based Hidden Markov Model (HMM) to automatically derive a customized IE system with balanced precision and recall. The key idea is to apply precision-oriented symbolic rules learned in the first stage to a large corpus in order to construct an automatically tagged training corpus. This training corpus is then used to train an HMM to boost the recall. The experiments conducted in named entity (NE) tagging and relationship extraction show a performance close to the performance of supervised learning systems.

Original languageEnglish
Pages62-67
Number of pages6
StatePublished - 2004
Event19th National Conference on Artificial Intelligence - San Jose, CA, United States
Duration: Jul 25 2004Jul 26 2004

Conference

Conference19th National Conference on Artificial Intelligence
Country/TerritoryUnited States
CitySan Jose, CA
Period07/25/0407/26/04

Fingerprint

Dive into the research topics of 'A bootstrapping approach to information extraction domain porting'. Together they form a unique fingerprint.

Cite this