TY - GEN
T1 - Summarizing lecture videos by key handwritten content regions
AU - Kota, Bhargava Urala
AU - Ahmed, Saleem
AU - Stone, Alexander
AU - Davila, Kenny
AU - Setlur, Srirangaraj
AU - Govindaraju, Venu
N1 - Publisher Copyright:
©2019 IEEE
PY - 2019
Y1 - 2019
N2 - We introduce a novel method for summarization of whiteboard lecture videos using key handwritten content regions. A deep neural network is used for detecting bounding boxes that contain semantically meaningful groups of handwritten content. A neural network embedding is learnt, under triplet loss, from the detected regions in order to discriminate between unique handwritten content. The detected regions along with embeddings at every frame of the lecture video are used to extract unique handwritten content across the video which are presented as the video summary. Additionally, a spatiotemporal index is constructed from the video which records the time and location of each individual summary region in the video which can potentially be used for content-based search and navigation. We train and test our methods on the publicly available AccessMath dataset. We use the DetEval scheme to benchmark our summarization by recall of unique ground truth objects (92.09%) and average number of summary regions (128) compared to the ground truth (88).
AB - We introduce a novel method for summarization of whiteboard lecture videos using key handwritten content regions. A deep neural network is used for detecting bounding boxes that contain semantically meaningful groups of handwritten content. A neural network embedding is learnt, under triplet loss, from the detected regions in order to discriminate between unique handwritten content. The detected regions along with embeddings at every frame of the lecture video are used to extract unique handwritten content across the video which are presented as the video summary. Additionally, a spatiotemporal index is constructed from the video which records the time and location of each individual summary region in the video which can potentially be used for content-based search and navigation. We train and test our methods on the publicly available AccessMath dataset. We use the DetEval scheme to benchmark our summarization by recall of unique ground truth objects (92.09%) and average number of summary regions (128) compared to the ground truth (88).
UR - https://www.scopus.com/pages/publications/85110521180
U2 - 10.1109/ICDARW.2019.30058
DO - 10.1109/ICDARW.2019.30058
M3 - Conference contribution
AN - SCOPUS:85110521180
T3 - 2019 International Conference on Document Analysis and Recognition Workshops, ICDARW 2019
SP - 13
EP - 18
BT - Proceedings - 8th International Workshop on Camera-Based Document Analysis and Recognition, CBDAR 2019 - ICDAR 2019 Workshop
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2019 International Conference on Document Analysis and Recognition Workshops, ICDARW 2019
Y2 - 22 September 2019 through 22 September 2019
ER -