Abstract
This research explores the interaction of textual and photographic information in document understanding. The problem of performing general-purpose vision without a priori knowledge is difficult at best. The use of collateral information in scene understanding has been explored in computer vision systems that use scene context in the task of object identification. The work described here extends this notion by defining visual semantics, a theory of systematically extracting picture-specific information from text accompanying a photograph. Specifically, this paper discusses the multi-stage processing of textual captions with the following objectives: (i) predicting with objects (implicitly or explicitly mentioned in the caption) are present in the picture and (ii) generating constraints useful in locating/identifying these objects. The implementation and use of a lexicon specifically designed for the integration of linguistic and visual information is discussed. Finally, the research described here has been successfully incorporated into PICTION, a caption-based face identification system.
| Original language | English |
|---|---|
| Pages | 793-798 |
| Number of pages | 6 |
| State | Published - 1994 |
| Event | Proceedings of the 12th National Conference on Artificial Intelligence. Part 1 (of 2) - Seattle, WA, USA Duration: Jul 31 1994 → Aug 4 1994 |
Conference
| Conference | Proceedings of the 12th National Conference on Artificial Intelligence. Part 1 (of 2) |
|---|---|
| City | Seattle, WA, USA |
| Period | 07/31/94 → 08/4/94 |
Fingerprint
Dive into the research topics of 'Visual semantics: extracting visual information from text accompanying pictures'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver