Skip to main navigation Skip to search Skip to main content

Cross-modal video moment retrieval with spatial and language-temporal attention

  • Hunan University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

90 Scopus citations

Abstract

Given an untrimmed video and a description query, temporal moment retrieval aims to localize the temporal segment within the video that best describes the textual query. Existing studies predominantly employ coarse frame-level features as the visual representation, obfuscating the specific details which may provide critical cues for localizing the desired moment. We propose a SLTA (short for "Spatial and Language-Temporal Attention") method to address the detail missing issue. Specifically, the SLTA method takes advantage of object-level local features and attends to the most relevant local features (e.g., the local features "girl", "cup") by spatial attention. Then we encode the sequence of local features on consecutive frames to capture the interaction information among these objects (e.g., the interaction "pour" involving these two objects). Meanwhile, a language-temporal attention is utilized to emphasize the keywords based on moment context information. Therefore, our proposed two attention sub-networks can recognize the most relevant objects and interactions in the video, and simultaneously highlight the keywords in the query. Extensive experiments on TACOS, Charades-STA and DiDeMo datasets demonstrate the effectiveness of our model as compared to state-of-the-art methods.

Original languageEnglish
Title of host publicationICMR 2019 - Proceedings of the 2019 ACM International Conference on Multimedia Retrieval
PublisherAssociation for Computing Machinery, Inc
Pages217-225
Number of pages9
ISBN (Electronic)9781450367653
DOIs
StatePublished - Jun 5 2019
Event2019 ACM International Conference on Multimedia Retrieval, ICMR 2019 - Ottawa, Canada
Duration: Jun 10 2019Jun 13 2019

Publication series

NameICMR 2019 - Proceedings of the 2019 ACM International Conference on Multimedia Retrieval

Conference

Conference2019 ACM International Conference on Multimedia Retrieval, ICMR 2019
Country/TerritoryCanada
CityOttawa
Period06/10/1906/13/19

Keywords

  • Cross-modal video retrieval
  • Language-temporal attention
  • Moment localization
  • Spatial attention

Fingerprint

Dive into the research topics of 'Cross-modal video moment retrieval with spatial and language-temporal attention'. Together they form a unique fingerprint.

Cite this