TY - GEN
T1 - Identifying extremism in social media with multi-view context-aware subset optimization
AU - Das Bhattacharjee, Sreyasee
AU - Balantrapu, Bala Venkatram
AU - Tolone, William
AU - Talukder, Ashit
N1 - Publisher Copyright:
© 2017 IEEE.
PY - 2017/7/1
Y1 - 2017/7/1
N2 - Cyber threat intelligence and security informatics play critical roles in the identification of the principal influencers and threats associated with criminal activity and extremism in web-based communities. This paper presents an effective, dynamic learning framework utilizing domain specific context information to design a novel classification model that can robustly identify malicious social media posts (e.g. enrollment propaganda for extremist groups) with expressions of extremism or criminal intent. Research towards automated identification of extreme online posts and their associated key suspects and threats faces numerous challenges, 1) Online data, particularly social media data, originated from numerous independent and heterogeneous sources are largely unstructured. 2) The tactics, techniques, and procedures (TTPs) of criminal activity and extremism are constantly evolving. 3) There are limited ground truth data to support the development of effective classification technologies. In this paper, we present a human-machine collaborative, semi-supervised learning system that can efficiently and effectively identify malicious social media posts in presence of these challenges. Our system and framework develops an initial classifier from limited annotated data and in an interactive manner evolves dynamically into a sophisticated model using shortlisted relevant samples, identified via a graph-based optimization method, solvable by maximum flow algorithm. This same method also may be used to refine the classifier as TTPs evolve. Under this framework, the classifier performance converges faster using roughly 1-2 orders of magnitude fewer annotated samples, as compared to fully supervised solutions, resulting in a reasonably acceptable accuracy of nearly 80%. We validate our framework using a large collection of English and non-English flagged words extracted from three web-based forums and manually verified by multiple independent annotators.
AB - Cyber threat intelligence and security informatics play critical roles in the identification of the principal influencers and threats associated with criminal activity and extremism in web-based communities. This paper presents an effective, dynamic learning framework utilizing domain specific context information to design a novel classification model that can robustly identify malicious social media posts (e.g. enrollment propaganda for extremist groups) with expressions of extremism or criminal intent. Research towards automated identification of extreme online posts and their associated key suspects and threats faces numerous challenges, 1) Online data, particularly social media data, originated from numerous independent and heterogeneous sources are largely unstructured. 2) The tactics, techniques, and procedures (TTPs) of criminal activity and extremism are constantly evolving. 3) There are limited ground truth data to support the development of effective classification technologies. In this paper, we present a human-machine collaborative, semi-supervised learning system that can efficiently and effectively identify malicious social media posts in presence of these challenges. Our system and framework develops an initial classifier from limited annotated data and in an interactive manner evolves dynamically into a sophisticated model using shortlisted relevant samples, identified via a graph-based optimization method, solvable by maximum flow algorithm. This same method also may be used to refine the classifier as TTPs evolve. Under this framework, the classifier performance converges faster using roughly 1-2 orders of magnitude fewer annotated samples, as compared to fully supervised solutions, resulting in a reasonably acceptable accuracy of nearly 80%. We validate our framework using a large collection of English and non-English flagged words extracted from three web-based forums and manually verified by multiple independent annotators.
KW - Active Learning
KW - Classification
KW - Darkweb
KW - Graph Search
KW - Multi-View Classification
KW - Semi Supervised Learning
UR - https://www.scopus.com/pages/publications/85047827077
U2 - 10.1109/BigData.2017.8258358
DO - 10.1109/BigData.2017.8258358
M3 - Conference contribution
AN - SCOPUS:85047827077
T3 - Proceedings - 2017 IEEE International Conference on Big Data, Big Data 2017
SP - 3638
EP - 3647
BT - Proceedings - 2017 IEEE International Conference on Big Data, Big Data 2017
A2 - Nie, Jian-Yun
A2 - Obradovic, Zoran
A2 - Suzumura, Toyotaro
A2 - Ghosh, Rumi
A2 - Nambiar, Raghunath
A2 - Wang, Chonggang
A2 - Zang, Hui
A2 - Baeza-Yates, Ricardo
A2 - Baeza-Yates, Ricardo
A2 - Hu, Xiaohua
A2 - Kepner, Jeremy
A2 - Cuzzocrea, Alfredo
A2 - Tang, Jian
A2 - Toyoda, Masashi
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 5th IEEE International Conference on Big Data, Big Data 2017
Y2 - 11 December 2017 through 14 December 2017
ER -