TY - GEN
T1 - Automatic indexing and content-based retrieval of captioned photographs
AU - Srihari, R. K.
N1 - Publisher Copyright:
© 1995 IEEE.
PY - 1995
Y1 - 1995
N2 - This research explores the interaction of textual and photographic information in an integrated text/image database environment. Specifically, We present a content-based retrieval system for captioned group photographs of people (i.e., human faces) where groups can consist of one or more members. By understanding the caption accompanying a picture, we are able to extract information useful in (i) retrieving the picture and (ii) directing an image interpretation system identify relevant objects (in this case, faces) in the picture. For the latter, we incorporate techniques from our ongoing research on photo understanding using accompanying text. Current image-based techniques have limitations; for example, similarity techniques used for retrieving faces will not perform well in group photographs where the locations of faces is not known a priori or where face sizes are small. By exploiting caption information, we assist a face locator in detecting human faces in a photograph and subsequently labelling them. Text-based similarity algorithms have principally relied on statistical techniques to index and classify documents (e.g., vector models). It is necessary to employ natural language processing techniques in order to derive deeper semantics from captions which contain far fewer words than documents. Our approach is unique since it goes beyond a superficial combination of existing text-based and image-based approaches to information retrieval.
AB - This research explores the interaction of textual and photographic information in an integrated text/image database environment. Specifically, We present a content-based retrieval system for captioned group photographs of people (i.e., human faces) where groups can consist of one or more members. By understanding the caption accompanying a picture, we are able to extract information useful in (i) retrieving the picture and (ii) directing an image interpretation system identify relevant objects (in this case, faces) in the picture. For the latter, we incorporate techniques from our ongoing research on photo understanding using accompanying text. Current image-based techniques have limitations; for example, similarity techniques used for retrieving faces will not perform well in group photographs where the locations of faces is not known a priori or where face sizes are small. By exploiting caption information, we assist a face locator in detecting human faces in a photograph and subsequently labelling them. Text-based similarity algorithms have principally relied on statistical techniques to index and classify documents (e.g., vector models). It is necessary to employ natural language processing techniques in order to derive deeper semantics from captions which contain far fewer words than documents. Our approach is unique since it goes beyond a superficial combination of existing text-based and image-based approaches to information retrieval.
UR - https://www.scopus.com/pages/publications/0347885746
U2 - 10.1109/ICDAR.1995.602129
DO - 10.1109/ICDAR.1995.602129
M3 - Conference contribution
AN - SCOPUS:0347885746
T3 - Proceedings of the International Conference on Document Analysis and Recognition, ICDAR
SP - 1165
EP - 1168
BT - Proceedings of the 3rd International Conference on Document Analysis and Recognition, ICDAR 1995
PB - IEEE Computer Society
T2 - 3rd International Conference on Document Analysis and Recognition, ICDAR 1995
Y2 - 14 August 1995 through 16 August 1995
ER -