TY - GEN
T1 - Efficient search of Top-K video subvolumes for multi-instance action detection
AU - Goussies, Norberto A.
AU - Liu, Zicheng
AU - Yuan, Junsong
PY - 2010
Y1 - 2010
N2 - Action detection was formulated as a subvolume mutual information maximization problem in [8], where each subvolume identifies where and when the action occurs in the video. Despite the fact that the proposed branch-and-bound algorithm can find the best subvolume efficiently for low resolution videos, it is still not efficient enough to perform multiinstance detection in videos of high spatial resolution. In this paper we develop an algorithm that further speeds up the subvolume search and targets on real-time multi-instance action detection for high resolution videos (e.g. 320 × 240 or higher). Unlike the previous branch-and-bound search technique which restarts a new search for each action instance, we find the Top-K subvolumes simultaneously with a single round of search. To handle the larger spatial resolution, we downsample the volume of videos for a more efficient upperbound estimation. To validate our algorithm, we perform experiments on a challenging dataset of 54 video sequences where each video consists of several actions performed by different people in a crowded environment. The experiments show that our method is not only efficient, but also capable of handling action variations caused by performing speed and style changes, spatial scale changes, as well as cluttered and moving background.
AB - Action detection was formulated as a subvolume mutual information maximization problem in [8], where each subvolume identifies where and when the action occurs in the video. Despite the fact that the proposed branch-and-bound algorithm can find the best subvolume efficiently for low resolution videos, it is still not efficient enough to perform multiinstance detection in videos of high spatial resolution. In this paper we develop an algorithm that further speeds up the subvolume search and targets on real-time multi-instance action detection for high resolution videos (e.g. 320 × 240 or higher). Unlike the previous branch-and-bound search technique which restarts a new search for each action instance, we find the Top-K subvolumes simultaneously with a single round of search. To handle the larger spatial resolution, we downsample the volume of videos for a more efficient upperbound estimation. To validate our algorithm, we perform experiments on a challenging dataset of 54 video sequences where each video consists of several actions performed by different people in a crowded environment. The experiments show that our method is not only efficient, but also capable of handling action variations caused by performing speed and style changes, spatial scale changes, as well as cluttered and moving background.
KW - Action recognition
KW - Branch-and-bound
UR - https://www.scopus.com/pages/publications/78349239606
U2 - 10.1109/ICME.2010.5583547
DO - 10.1109/ICME.2010.5583547
M3 - Conference contribution
AN - SCOPUS:78349239606
SN - 9781424474912
T3 - 2010 IEEE International Conference on Multimedia and Expo, ICME 2010
SP - 328
EP - 333
BT - 2010 IEEE International Conference on Multimedia and Expo, ICME 2010
T2 - 2010 IEEE International Conference on Multimedia and Expo, ICME 2010
Y2 - 19 July 2010 through 23 July 2010
ER -