TY - GEN
T1 - Language identification in historical Afghan Manuscripts
AU - Farooq, Faisal
AU - Govindaraju, Venu
PY - 2007
Y1 - 2007
N2 - Automatic language identification is an important step prior to optical character recognition (OCR). In this paper we present a system to discriminate between Arabic and Persian in historical Afghan Manuscripts. The classification is performed at a sub-sentence level. We propose a feature extraction algorithm for a sub-sentence based on Gabor filters followed by classification using a Support Vector Machine (SVM). An overall precision of 96.72% and 94.90% is obtained for Persian and Arabic respectively .
AB - Automatic language identification is an important step prior to optical character recognition (OCR). In this paper we present a system to discriminate between Arabic and Persian in historical Afghan Manuscripts. The classification is performed at a sub-sentence level. We propose a feature extraction algorithm for a sub-sentence based on Gabor filters followed by classification using a Support Vector Machine (SVM). An overall precision of 96.72% and 94.90% is obtained for Persian and Arabic respectively .
UR - https://www.scopus.com/pages/publications/51549109030
U2 - 10.1109/ISSPA.2007.4555588
DO - 10.1109/ISSPA.2007.4555588
M3 - Conference contribution
AN - SCOPUS:51549109030
SN - 1424407796
SN - 9781424407798
T3 - 2007 9th International Symposium on Signal Processing and its Applications, ISSPA 2007, Proceedings
BT - 2007 9th International Symposium on Signal Processing and its Applications, ISSPA 2007, Proceedings
T2 - 2007 9th International Symposium on Signal Processing and its Applications, ISSPA 2007
Y2 - 12 February 2007 through 15 February 2007
ER -