TY - GEN
T1 - A Bayesian Approach to Script Independent Multilingual Keyword Spotting
AU - Kumar, Gaurav
AU - Govindaraju, Venu
N1 - Publisher Copyright:
© 2014 IEEE.
PY - 2014/12/9
Y1 - 2014/12/9
N2 - We propose a script independent bayesian framework for keyword spotting in multilingual handwritten documents. The approach relies on local character level score and global word level hypothesis scores and learns a bayesian logistic regression classifier to distinguish between keywords and non-keywords. In a bayesian formulation of logistic regression, the integral over weights becomes intractable. Variational approximation is used for inference. In order to learn a robust classifier with minimal number of samples, we apply bayesian active learning framework to request labels for those word images which provide maximum information gain in improving the classifier. We evaluate our system on multilingual datasets, publicly available IAM dataset for English, AMA for Arabic and LAW dataset for Devanagiri. The system is also evaluated on a synthetic multilingual dataset prepared by combining samples from IAM, AMA and LAW datasets. The results are comparable with the state of art multilingual keyword spotting framework.
AB - We propose a script independent bayesian framework for keyword spotting in multilingual handwritten documents. The approach relies on local character level score and global word level hypothesis scores and learns a bayesian logistic regression classifier to distinguish between keywords and non-keywords. In a bayesian formulation of logistic regression, the integral over weights becomes intractable. Variational approximation is used for inference. In order to learn a robust classifier with minimal number of samples, we apply bayesian active learning framework to request labels for those word images which provide maximum information gain in improving the classifier. We evaluate our system on multilingual datasets, publicly available IAM dataset for English, AMA for Arabic and LAW dataset for Devanagiri. The system is also evaluated on a synthetic multilingual dataset prepared by combining samples from IAM, AMA and LAW datasets. The results are comparable with the state of art multilingual keyword spotting framework.
KW - Bayesian Active Learning
KW - Handwritten Multilingual Documents
KW - Script Independent
KW - Spotting
UR - https://www.scopus.com/pages/publications/84942234069
U2 - 10.1109/ICFHR.2014.66
DO - 10.1109/ICFHR.2014.66
M3 - Conference contribution
AN - SCOPUS:84942234069
T3 - Proceedings of International Conference on Frontiers in Handwriting Recognition, ICFHR
SP - 357
EP - 362
BT - Proceedings - 14th International Conference on Frontiers in Handwriting Recognition, ICFHR 2014
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 14th International Conference on Frontiers in Handwriting Recognition, ICFHR 2014
Y2 - 1 September 2014 through 4 September 2014
ER -