Skip to main navigation Skip to search Skip to main content

Asian character recognition

    • SUNY Buffalo

    Research output: Chapter in Book/Report/Conference proceedingChapterpeer-review

    Abstract

    This chapter deals with the automated recognition of Asian scripts and focuses primarily on text belonging to two script families, viz., oriental scripts such as Chinese, Japanese, and Korean (CJK) and scripts from the Indian subcontinent (Indic) such as Devanagari, Bangla, and the scripts of South India. Since these scripts have attracted the greatest interest from the document analysis community and cover most of the issues potentially encountered in the recognition of Asian scripts, application to other Asian scripts should primarily be a matter of implementation. Specific challenges encountered in the OCR of CJK and Indic scripts are due to the large number of character classes and the resultant high probability of confusions between similar character shapes for a machine reading system. This has led to a greater role being played by language models and other post-processing techniques in the development of successful OCR systems for Asian scripts.

    Original languageEnglish
    Title of host publicationHandbook of Document Image Processing and Recognition
    PublisherSpringer London
    Pages459-486
    Number of pages28
    ISBN (Electronic)9780857298591
    ISBN (Print)9780857298584
    DOIs
    StatePublished - Jan 1 2014

    Keywords

    • Character recognition
    • Chinese
    • CJK
    • Classifiers
    • Devanagari
    • Features
    • Hindi
    • HMM
    • Indic
    • Japanese
    • Korean
    • Languagemodels
    • Neural networks
    • OCR
    • Post-processing

    Fingerprint

    Dive into the research topics of 'Asian character recognition'. Together they form a unique fingerprint.

    Cite this