Skip to main navigation Skip to search Skip to main content

Script independent word spotting in multilingual documents

  • SUNY Buffalo

Research output: Contribution to conferencePaperpeer-review

32 Scopus citations

Abstract

This paper describes a method for script independent word spotting in multilingual handwritten and machine printed documents. The system accepts a query in the form of text from the user and returns a ranked list of word images from document image corpus based on similarity with the query word. The system is divided into two main components. The first component known as Indexer, performs indexing of all word images present in the document image corpus. This is achieved by extracting Moment Based features from word images and storing them as index. A template is generated for keyword spotting which stores the mapping of a keyword string to its corresponding word image which is used for generating query feature vector. The second component, Similarity Matcher, returns a ranked list of word images which are most similar to the query based on a cosine similarity metric. A manual Relevance feedback is applied based on Rocchio's formula, which re-formulates the query vector to return an improved ranked listing of word images. The performance of the system is seen to be superior on printed text than on handwritten text. Experiments are reported on documents of three different languages: English, Hindi and Sanskrit. For handwritten English, an average precision of 67% was obtained for 30 query words. For machine printed Hindi, an average precision of 71% was obtained for 75 query words and for Sanskrit, an average precision of 87% with 100 queries was obtained.

Conference

Conference2nd International Workshop on Cross Lingual Information Access: Addressing the Information Need of Multilingual Societies, CLIA 2008 - held in conjunction with the 3rd International Joint Conference on Natural Language Processing, IJCNLP 2008
Country/TerritoryIndia
CityHyderabad
Period01/11/08 → …

Fingerprint

Dive into the research topics of 'Script independent word spotting in multilingual documents'. Together they form a unique fingerprint.

Cite this