Skip to main navigation Skip to search Skip to main content

Semantic Enhanced Sketch Based Image Retrieval with Incomplete Multimodal Query

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Sketch Based Image Retrieval (SBIR) is a challenging problem mainly due to a significant cross-domain gap between hand-drawn sketches and natural images. While extra semantic information (such as attribute details) can facilitate query-adaptive search, we still have to face two challenges: (1) an incomplete multimodal query; (2) lack of sketch-image paired training data. Toward this end, many existing multimodal sketch retrieval frameworks utilize text-based label information to augment the limited sketch query with more semantic details. However, that single word-level category information may not always reveal sufficient characteristics on the object specific fine-grained attributes. In this work, we propose a multimodal SBIR system that allows both sketch and text level attribute description for query. In order to bridge the cross-modal gaps among sketch, image, and texts, given a semantic in consideration, two mode-specific semantic networks provide layer-wise regularizer parameters to dynamically adopt the underlying semantic within the learned sketch feature representation and thereby transform the initial generic sketch into a more comprehensive Semantic Enhanced Joint Embedding (SEJE). Also the availability of the multimodal paired samples may not be always feasible; neither during training nor during test phases. Therefore, instead of relying on strict one-one cross modal correspondence, the learning of the joint sketch embedding SEJE relies on capturing the semantic relevant cross-modal correspondences between an averaged mode-specific semantic features (image and text) and the sketch feature, which facilitates SEJE's improved generalization ability. Evaluation on two benchmark datasets: Sketchy and TU-Berlin clearly validates the superiority of the proposed method, compared to the state-of-the-art methods. In fact, by allowing a user to add text attributes to make a complete multimodal query, the proposed method improves the mAP scores by 5-7% on the challenging sketch-attribute composition test scenario.

Original languageEnglish
Title of host publicationProceedings - 2020 IEEE 6th International Conference on Multimedia Big Data, BigMM 2020
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages86-93
Number of pages8
ISBN (Electronic)9781728193250
DOIs
StatePublished - Sep 2020
Event6th IEEE International Conference on Multimedia Big Data, BigMM 2020 - New Delhi, India
Duration: Sep 24 2020Sep 26 2020

Publication series

NameProceedings - 2020 IEEE 6th International Conference on Multimedia Big Data, BigMM 2020

Conference

Conference6th IEEE International Conference on Multimedia Big Data, BigMM 2020
Country/TerritoryIndia
CityNew Delhi
Period09/24/2009/26/20

Keywords

  • multimodal search
  • semantic search
  • sketch based image retrieval
  • triplet loss

Fingerprint

Dive into the research topics of 'Semantic Enhanced Sketch Based Image Retrieval with Incomplete Multimodal Query'. Together they form a unique fingerprint.

Cite this