Skip to main navigation Skip to search Skip to main content

MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs

  • Wenqian Ye
  • , Bohan Liu
  • , Guangtao Zheng
  • , Di Wang
  • , Xu Cao
  • , Yunsheng Ma
  • , Bolin Lai
  • , James M. Rehg
  • , Aidong Zhang
  • University of Virginia
  • University of Illinois at Urbana-Champaign
  • Purdue University
  • Georgia Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Spurious bias, a tendency to exploit spurious correlations between superficial input attributes and prediction targets, has revealed a severe robustness pitfall in classical machine learning problems. Multimodal Large Language Models (MLLMs), which leverage pretrained vision and language models, have recently demonstrated strong capability in joint vision-language understanding. However, both the presence and severity of spurious biases in MLLMs remain poorly understood. In this work, we address this gap by analyzing the spurious biases in the multimodal setting and uncovering the specific inference-time data patterns that can manifest this problem. To support this analysis, we introduce MM-SpuBench, a comprehensive, human-verified benchmark dataset consisting of image-class pairs annotated with core and spurious attributes, grounded in our taxonomy of nine distinct types of spurious correlations. The benchmark is constructed using human-interpretable attribute information to capture a wide range of spurious patterns reflective of real-world knowledge. Leveraging this benchmark, we conduct a comprehensive evaluation of the state-of-the-art open-source and proprietary MLLMs with both standard accuracy and the proposed Conditional Generation Likelihood Advantage (CGLA). Our findings highlight the persistence of reliance on spurious correlations and the difficulty of mitigation on our benchmark. We hope this work can inspire new technical strides to mitigate these biases. Our benchmark is publicly available at https://huggingface.co/datasets/mmbench/MM-SpuBench.

Original languageEnglish
Title of host publicationKDD 2026 - Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1
PublisherAssociation for Computing Machinery
Pages2854-2865
Number of pages12
ISBN (Electronic)9798400722585
DOIs
StatePublished - Apr 20 2026
Event32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, KDD 2026 - Jeju Island, Korea, Republic of
Duration: Aug 9 2026Aug 13 2026

Publication series

NameProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
Volume1-A
ISSN (Print)2154-817X

Conference

Conference32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, KDD 2026
Country/TerritoryKorea, Republic of
CityJeju Island
Period08/9/2608/13/26

Keywords

  • dataset and benchmark
  • multimodal llms
  • spurious correlations

Fingerprint

Dive into the research topics of 'MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs'. Together they form a unique fingerprint.

Cite this