Skip to main navigation Skip to search Skip to main content

Use of collateral text in understanding photos in documents

Research output: Contribution to journalConference articlepeer-review

1 Scopus citations

Abstract

This research explores the interaction of textual and photographic information in document understanding. Specifically, it presents a computational model whereby textual captions are used as collateral information in the interpretation of the corresponding photographs. The final understanding of the picture and caption reflects a consolidation of the information obtained from each of the two sources and can thus be used in intelligent information retrieval tasks. The problem of performing general-purpose vision without apriori knowledge is very difficult at best. The concept of using collateral information in scene understanding has been explored in systems that use general scene context in the task of object identification. The work described here extends this notion by incorporating picture specific information. A multi-stage system FICTION which uses captions to identify humans in an accompanying photograph is described. This provides a computationally less expensive alternative to traditional methods of face recognition. It does not require a pre-stored database of face models for all people to be identified. A key component of the system is the utilisation of spatial and characteristic constraints (derived from the caption )in labeling face candidates (generated by a face locator).

Original languageEnglish
Pages (from-to)186-199
Number of pages14
JournalProceedings of SPIE - The International Society for Optical Engineering
Volume2103
DOIs
StatePublished - Feb 25 1994
Event22nd Applied Imagery Pattern Recognition Workshop: Interdisciplinary Computer Vision: Applications and Changing Needs 1993 - Washington, United States
Duration: Oct 13 1993Oct 15 1993

Fingerprint

Dive into the research topics of 'Use of collateral text in understanding photos in documents'. Together they form a unique fingerprint.

Cite this