Skip to main navigation Skip to search Skip to main content

An OCR based approach for word spotting in Devanagari documents

  • SUNY Buffalo

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

6 Scopus citations

Abstract

This paper describes an OCR-based technique for word spotting in Devanagari printed documents. The system accepts a Devanagari word as input and returns a sequence of word images that are ranked according to their similarity with the input query. The methodology involves line and word separation, pre-processing document words, word recognition using OCR and similarity matching. We demonstrate a Block Adjacency Graph (BAG) based document cleanup in the pre-processing phase. During word recognition, multiple recognition hypotheses are generated for each document word using a font-independent Devanagari OCR. The similarity matching phase uses a cost based model to match the word input by a user and the OCR results. Experiments are conducted on document images from the publicly available ILT and Million Book Project dataset. The technique achieves an average precision of 80% for 10 queries and 67% for 20 queries for a set of 64 documents containing 5780 word images. The paper also presents a comparison of our method with template-based word spotting techniques.

Original languageEnglish
Title of host publicationDocument Recognition and Retrieval XV
DOIs
StatePublished - 2008
EventDocument Recognition and Retrieval XV - San Jose, CA, United States
Duration: Jan 29 2008Jan 31 2008

Publication series

NameProceedings of SPIE - The International Society for Optical Engineering
Volume6815
ISSN (Print)0277-786X

Conference

ConferenceDocument Recognition and Retrieval XV
Country/TerritoryUnited States
CitySan Jose, CA
Period01/29/0801/31/08

Fingerprint

Dive into the research topics of 'An OCR based approach for word spotting in Devanagari documents'. Together they form a unique fingerprint.

Cite this