TY - GEN
T1 - Cyberbullying detection with a pronunciation based convolutional neural network
AU - Zhang, Xiang
AU - Tong, Jonathan
AU - Vishwamitra, Nishant
AU - Whittaker, Elizabeth
AU - Mazer, Joseph P.
AU - Kowalski, Robin
AU - Hu, Hongxin
AU - Luo, Feng
AU - Macbeth, Jamie
AU - Dillon, Edward
N1 - Publisher Copyright:
© 2016 IEEE.
PY - 2017/1/31
Y1 - 2017/1/31
N2 - Cyberbullying can have a deep and long lasting impact on its victims, who are often adolescents. Accurately detecting cyberbullying helps prevent it. However, the noise and errors in social media posts and messages make detecting cyberbullying very challenging. In this paper, we propose a novel pronunciation based convolutional neural network (PCNN) to address this challenge. Upon observing that the pronunciation of misspelled words in informal online conversations is often unchanged, we used the phoneme codes of the text as the features for a convolutional neural network. This procedure corrects spelling errors that did not alter the pronunciation, thereby alleviating the problem of noise and bullying data sparsity. To overcome class imbalance, a common problem in cyberbullying datasets, we implement three techniques that include thresholdmoving, cost function adjusting, and a hybrid solution in our model. We evaluate the performance of our models using two cyberbullying datasets collected from Twitter and Formspring.me. The results of our experiment show that PCNN can achieve improved recall and precision compared to baseline convolutional neural networks.
AB - Cyberbullying can have a deep and long lasting impact on its victims, who are often adolescents. Accurately detecting cyberbullying helps prevent it. However, the noise and errors in social media posts and messages make detecting cyberbullying very challenging. In this paper, we propose a novel pronunciation based convolutional neural network (PCNN) to address this challenge. Upon observing that the pronunciation of misspelled words in informal online conversations is often unchanged, we used the phoneme codes of the text as the features for a convolutional neural network. This procedure corrects spelling errors that did not alter the pronunciation, thereby alleviating the problem of noise and bullying data sparsity. To overcome class imbalance, a common problem in cyberbullying datasets, we implement three techniques that include thresholdmoving, cost function adjusting, and a hybrid solution in our model. We evaluate the performance of our models using two cyberbullying datasets collected from Twitter and Formspring.me. The results of our experiment show that PCNN can achieve improved recall and precision compared to baseline convolutional neural networks.
UR - https://www.scopus.com/pages/publications/85015359738
U2 - 10.1109/ICMLA.2016.0132
DO - 10.1109/ICMLA.2016.0132
M3 - Conference contribution
AN - SCOPUS:85015359738
T3 - Proceedings - 2016 15th IEEE International Conference on Machine Learning and Applications, ICMLA 2016
SP - 740
EP - 745
BT - Proceedings - 2016 15th IEEE International Conference on Machine Learning and Applications, ICMLA 2016
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 15th IEEE International Conference on Machine Learning and Applications, ICMLA 2016
Y2 - 18 December 2016 through 20 December 2016
ER -