Abstract
This paper presents a seed-driven, bootstrapping approach to domain porting that could be used to customize a generic information extraction (IE) capability for a specific domain. The approach taken is based on the existence of a robust, domain-independent IE engine that can continue to be enhanced, independent of any particular domain. This approach combines the strengths of parsing-based symbolic rule learning and the high performance linear string-based Hidden Markov Model (HMM) to automatically derive a customized IE system with balanced precision and recall. The key idea is to apply precision-oriented symbolic rules learned in the first stage to a large corpus in order to construct an automatically tagged training corpus. This training corpus is then used to train an HMM to boost the recall. The experiments conducted in named entity (NE) tagging and relationship extraction show a performance close to the performance of supervised learning systems.
| Original language | English |
|---|---|
| Pages | 62-67 |
| Number of pages | 6 |
| State | Published - 2004 |
| Event | 19th National Conference on Artificial Intelligence - San Jose, CA, United States Duration: Jul 25 2004 → Jul 26 2004 |
Conference
| Conference | 19th National Conference on Artificial Intelligence |
|---|---|
| Country/Territory | United States |
| City | San Jose, CA |
| Period | 07/25/04 → 07/26/04 |
Fingerprint
Dive into the research topics of 'A bootstrapping approach to information extraction domain porting'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver