Skip to main navigation Skip to search Skip to main content

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation

  • Shijie Zhou
  • , Ruiyi Zhang
  • , Huaisheng Zhu
  • , Branislav Kveton
  • , Yufan Zhou
  • , Jiuxiang Gu
  • , Jian Chen
  • , Changyou Chen
  • SUNY Buffalo
  • Adobe Systems Incorporated
  • Pennsylvania State University
  • Luma Ai

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Scopus citations

Abstract

We introduce LLaVA-Reward11Project page: https://github.com/sjz5202/LLaVAReward., an efficient reward model designed to automatically evaluate text-to-image (T2I) generations across multiple perspectives, leveraging pretrained multimodal large language models (MLLMs). Existing MLLM-based approaches require instruction-following data for supervised fine-tuning and evaluate generation quality on analyzing text response, which is time-consuming and difficult to train. To address this problem, we propose LLaVA-Reward, which directly utilizes the hidden states of MLLMs given text-image pairs. To enhance the bidirectional interaction between visual and textual representations in decoder-only MLLMs, we further propose adding a Skip-connection Cross Attention (SkipCA) module. This design enhances text-image correlation reasoning by connecting early-layer visual features with later-layer hidden representations. In addition, LLaVA-Reward supports different types of preference data for efficient fine-tuning, including paired preference data and unpaired data. We train LLaVA-Reward on four evaluation perspectives: textimage alignment, fidelity/artifact, safety, and overall ranking. Empirical results demonstrate that LLaVA-Reward outperforms conventional and MLLM-based methods in generating human-aligned scores for automatic evaluations and inference-time scaling in text-to-image generations.

Original languageEnglish
Title of host publicationProceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages19638-19648
Number of pages11
ISBN (Electronic)9798331587758
DOIs
StatePublished - 2025
Event2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025 - Honolulu, United States
Duration: Oct 19 2025Oct 23 2025

Publication series

NameProceedings of the IEEE International Conference on Computer Vision
ISSN (Print)1550-5499
ISSN (Electronic)2380-7504

Conference

Conference2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
Country/TerritoryUnited States
CityHonolulu
Period10/19/2510/23/25

Fingerprint

Dive into the research topics of 'Multimodal LLMs as Customized Reward Models for Text-to-Image Generation'. Together they form a unique fingerprint.

Cite this