Skip to main navigation Skip to search Skip to main content

Language identification in historical Afghan Manuscripts

  • SUNY Buffalo

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Automatic language identification is an important step prior to optical character recognition (OCR). In this paper we present a system to discriminate between Arabic and Persian in historical Afghan Manuscripts. The classification is performed at a sub-sentence level. We propose a feature extraction algorithm for a sub-sentence based on Gabor filters followed by classification using a Support Vector Machine (SVM). An overall precision of 96.72% and 94.90% is obtained for Persian and Arabic respectively .

Original languageEnglish
Title of host publication2007 9th International Symposium on Signal Processing and its Applications, ISSPA 2007, Proceedings
DOIs
StatePublished - 2007
Event2007 9th International Symposium on Signal Processing and its Applications, ISSPA 2007 - Sharjah, United Arab Emirates
Duration: Feb 12 2007Feb 15 2007

Publication series

Name2007 9th International Symposium on Signal Processing and its Applications, ISSPA 2007, Proceedings

Conference

Conference2007 9th International Symposium on Signal Processing and its Applications, ISSPA 2007
Country/TerritoryUnited Arab Emirates
CitySharjah
Period02/12/0702/15/07

Fingerprint

Dive into the research topics of 'Language identification in historical Afghan Manuscripts'. Together they form a unique fingerprint.

Cite this