Skip to main navigation Skip to search Skip to main content

TRINS: Towards Multimodal Language Models that Can Read

  • Ruiyi Zhang
  • , Yanzhe Zhang
  • , Jian Chen
  • , Yufan Zhou
  • , Jiuxiang Gu
  • , Changyou Chen
  • , Tong Sun
  • Adobe Systems Incorporated
  • Georgia Institute of Technology
  • SUNY Buffalo

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

3 Scopus citations

Abstract

Large multimodal language models have shown remarkable proficiency in understanding and editing images. However, a majority of these visually-tuned models struggle to comprehend the textual content embedded in images, primar-ily due to the limitation of training data. In this work, we introduce TRINS: a Text-Rich imageIn this work, we use the phrase 'text-rich images' to describe images with rich textual information, such as posters and book covers. INStruction dataset, with the objective of enhancing the reading ability of the multimodal large language model. TRINS is built upon LAION22Work done during Q3 2023. using hybrid data annotation strategies that include machine-assisted and human-assisted annotation process. It contains 39,153 text-rich images, captions, and 102,437 questions. Specifically, we show that the number of words per annotation in TRINS is significantly longer than that of related datasets, providing new challenges. Furthermore, we introduce a simple and effective architecture, called a Language-Vision Reading Assistant (LaRA), which is good at understanding textual content within images. LaRA outperforms existing state-of-the-art multimodal large language models on the TRINS dataset as well as other classical benchmarks. Lastly, we conducted a comprehensive evaluation with TRINS on various text-rich image understanding and generation tasks, demonstrating its effectiveness.

Original languageEnglish
Title of host publicationProceedings - 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024
PublisherIEEE Computer Society
Pages22584-22594
Number of pages11
ISBN (Electronic)9798350353006
DOIs
StatePublished - 2024
Event2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 - Seattle, United States
Duration: Jun 16 2024Jun 22 2024

Publication series

NameProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
ISSN (Print)1063-6919

Conference

Conference2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024
Country/TerritoryUnited States
CitySeattle
Period06/16/2406/22/24

Fingerprint

Dive into the research topics of 'TRINS: Towards Multimodal Language Models that Can Read'. Together they form a unique fingerprint.

Cite this