Skip to main navigation Skip to search Skip to main content

Maryland at FIRE 2011: Retrieval of OCR'd Bengali

  • Indian Statistical Institute
  • University of Maryland, College Park

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Scopus citations

Abstract

In this year's Forum for Information Retrieval Evaluation (FIRE), the University of Maryland participated in the Retrieval of Indic Script OCRed Text (RISOT) task to experiment with the retrieval of Bengali script OCR'd documents. The experiments focused on evaluating a retrieval strategy motivated by recent work on Cross-Language Information Retrieval (CLIR), but which makes use of OCR error modeling rather than parallel text alignment. The approach obtains a probability distribution over substitutions for the actual query terms that possibly correspond to terms in the document representation. The results reported indicate that this is a promising way of using OCR error modeling to improve CLIR.

Original languageEnglish
Title of host publicationMultilingual Information Access in South Asian Languages - Second International Workshop, FIRE 2010 and Third International Workshop, FIRE 2011, Revised Selected Papers
Pages205-213
Number of pages9
DOIs
StatePublished - 2013
Event3rd International Workshop on Multilingual Information Access in South Asian Languages, FIRE 2011 - Bombay, India
Duration: Dec 2 2011Dec 4 2011

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume7536 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference3rd International Workshop on Multilingual Information Access in South Asian Languages, FIRE 2011
Country/TerritoryIndia
CityBombay
Period12/2/1112/4/11

Fingerprint

Dive into the research topics of 'Maryland at FIRE 2011: Retrieval of OCR'd Bengali'. Together they form a unique fingerprint.

Cite this