TY - GEN
T1 - An OCR based approach for word spotting in Devanagari documents
AU - Bhardwaj, Anurag
AU - Kompalli, Suryaprakash
AU - Setlur, Srirangaraj
AU - Govindaraju, Venu
PY - 2008
Y1 - 2008
N2 - This paper describes an OCR-based technique for word spotting in Devanagari printed documents. The system accepts a Devanagari word as input and returns a sequence of word images that are ranked according to their similarity with the input query. The methodology involves line and word separation, pre-processing document words, word recognition using OCR and similarity matching. We demonstrate a Block Adjacency Graph (BAG) based document cleanup in the pre-processing phase. During word recognition, multiple recognition hypotheses are generated for each document word using a font-independent Devanagari OCR. The similarity matching phase uses a cost based model to match the word input by a user and the OCR results. Experiments are conducted on document images from the publicly available ILT and Million Book Project dataset. The technique achieves an average precision of 80% for 10 queries and 67% for 20 queries for a set of 64 documents containing 5780 word images. The paper also presents a comparison of our method with template-based word spotting techniques.
AB - This paper describes an OCR-based technique for word spotting in Devanagari printed documents. The system accepts a Devanagari word as input and returns a sequence of word images that are ranked according to their similarity with the input query. The methodology involves line and word separation, pre-processing document words, word recognition using OCR and similarity matching. We demonstrate a Block Adjacency Graph (BAG) based document cleanup in the pre-processing phase. During word recognition, multiple recognition hypotheses are generated for each document word using a font-independent Devanagari OCR. The similarity matching phase uses a cost based model to match the word input by a user and the OCR results. Experiments are conducted on document images from the publicly available ILT and Million Book Project dataset. The technique achieves an average precision of 80% for 10 queries and 67% for 20 queries for a set of 64 documents containing 5780 word images. The paper also presents a comparison of our method with template-based word spotting techniques.
UR - https://www.scopus.com/pages/publications/41149154360
U2 - 10.1117/12.767289
DO - 10.1117/12.767289
M3 - Conference contribution
AN - SCOPUS:41149154360
SN - 9780819469878
T3 - Proceedings of SPIE - The International Society for Optical Engineering
BT - Document Recognition and Retrieval XV
T2 - Document Recognition and Retrieval XV
Y2 - 29 January 2008 through 31 January 2008
ER -