Skip to main navigation Skip to search Skip to main content

SV-RAG: LORA-CONTEXTUALIZING ADAPTATION OF MLLMS FOR LONG DOCUMENT UNDERSTANDING

  • Jian Chen
  • , Ruiyi Zhang
  • , Yufan Zhou
  • , Tong Yu
  • , Franck Dernoncourt
  • , Jiuxiang Gu
  • , Ryan Rossi
  • , Changyou Chen
  • , Tong Sun
  • SUNY Buffalo
  • Adobe Systems Incorporated

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

2 Scopus citations

Abstract

Multimodal large language models (MLLMs) have recently shown great progress in text-rich image understanding, yet they still struggle with complex, multi-page visually-rich documents. Traditional methods using document parsers for retrieval-augmented generation suffer from performance and efficiency limitations, while directly presenting all pages to MLLMs leads to inefficiencies, especially with lengthy ones. In this work, we present a novel framework named Self-Visual Retrieval-Augmented Generation (SV-RAG), which can broaden horizons of any MLLM to support long-document understanding. We demonstrate that MLLMs themselves can be an effective multimodal retriever to fetch relevant pages and then answer user questions based on these pages. SV-RAG is implemented with two specific MLLM adapters, one for evidence page retrieval and the other for question answering. Empirical results show state-of-the-art performance on public benchmarks, demonstrating the effectiveness of SV-RAG.

Original languageEnglish
Title of host publication13th International Conference on Learning Representations, ICLR 2025
PublisherInternational Conference on Learning Representations, ICLR
Pages34867-34882
Number of pages16
ISBN (Electronic)9798331320850
StatePublished - 2025
Event13th International Conference on Learning Representations, ICLR 2025 - Singapore, Singapore
Duration: Apr 24 2025Apr 28 2025

Publication series

Name13th International Conference on Learning Representations, ICLR 2025

Conference

Conference13th International Conference on Learning Representations, ICLR 2025
Country/TerritorySingapore
CitySingapore
Period04/24/2504/28/25

Fingerprint

Dive into the research topics of 'SV-RAG: LORA-CONTEXTUALIZING ADAPTATION OF MLLMS FOR LONG DOCUMENT UNDERSTANDING'. Together they form a unique fingerprint.

Cite this