TY - GEN
T1 - Cross-modal video moment retrieval with spatial and language-temporal attention
AU - Jiang, Bin
AU - Huang, Xin
AU - Yang, Chao
AU - Yuan, Junsong
N1 - Publisher Copyright:
© 2019 Association for Computing Machinery.
PY - 2019/6/5
Y1 - 2019/6/5
N2 - Given an untrimmed video and a description query, temporal moment retrieval aims to localize the temporal segment within the video that best describes the textual query. Existing studies predominantly employ coarse frame-level features as the visual representation, obfuscating the specific details which may provide critical cues for localizing the desired moment. We propose a SLTA (short for "Spatial and Language-Temporal Attention") method to address the detail missing issue. Specifically, the SLTA method takes advantage of object-level local features and attends to the most relevant local features (e.g., the local features "girl", "cup") by spatial attention. Then we encode the sequence of local features on consecutive frames to capture the interaction information among these objects (e.g., the interaction "pour" involving these two objects). Meanwhile, a language-temporal attention is utilized to emphasize the keywords based on moment context information. Therefore, our proposed two attention sub-networks can recognize the most relevant objects and interactions in the video, and simultaneously highlight the keywords in the query. Extensive experiments on TACOS, Charades-STA and DiDeMo datasets demonstrate the effectiveness of our model as compared to state-of-the-art methods.
AB - Given an untrimmed video and a description query, temporal moment retrieval aims to localize the temporal segment within the video that best describes the textual query. Existing studies predominantly employ coarse frame-level features as the visual representation, obfuscating the specific details which may provide critical cues for localizing the desired moment. We propose a SLTA (short for "Spatial and Language-Temporal Attention") method to address the detail missing issue. Specifically, the SLTA method takes advantage of object-level local features and attends to the most relevant local features (e.g., the local features "girl", "cup") by spatial attention. Then we encode the sequence of local features on consecutive frames to capture the interaction information among these objects (e.g., the interaction "pour" involving these two objects). Meanwhile, a language-temporal attention is utilized to emphasize the keywords based on moment context information. Therefore, our proposed two attention sub-networks can recognize the most relevant objects and interactions in the video, and simultaneously highlight the keywords in the query. Extensive experiments on TACOS, Charades-STA and DiDeMo datasets demonstrate the effectiveness of our model as compared to state-of-the-art methods.
KW - Cross-modal video retrieval
KW - Language-temporal attention
KW - Moment localization
KW - Spatial attention
UR - https://www.scopus.com/pages/publications/85068029978
U2 - 10.1145/3323873.3325019
DO - 10.1145/3323873.3325019
M3 - Conference contribution
AN - SCOPUS:85068029978
T3 - ICMR 2019 - Proceedings of the 2019 ACM International Conference on Multimedia Retrieval
SP - 217
EP - 225
BT - ICMR 2019 - Proceedings of the 2019 ACM International Conference on Multimedia Retrieval
PB - Association for Computing Machinery, Inc
T2 - 2019 ACM International Conference on Multimedia Retrieval, ICMR 2019
Y2 - 10 June 2019 through 13 June 2019
ER -