TY - JOUR
T1 - The Clever Hans Mirage
T2 - A Comprehensive Survey on Spurious Correlations in Machine Learning
AU - Ye, Wenqian
AU - Jiang, Luyang
AU - Xie, Eric
AU - Zheng, Guangtao
AU - Ma, Yunsheng
AU - Cao, Xu
AU - Guo, Dongliang
AU - Qi, Daiqing
AU - He, Zeyu
AU - Tian, Yijun
AU - Coffee, Megan
AU - Zeng, Zhe
AU - Li, Sheng
AU - Huang, Ting Hao ‘Kenneth’
AU - Wang, Ziran
AU - Rehg, James M.
AU - Kautz, Henry
AU - Zhang, Aidong
N1 - Publisher Copyright:
© 2026, Transactions on Machine Learning Research. All rights reserved.
PY - 2026
Y1 - 2026
N2 - Back in the early 20th century, a famous horse named Hans appeared to perform arithmetic and other intellectual tasks during exhibitions in Germany, which has drew widespread attention. However, later studies showed that Hans relied solely on subtle, involuntary cues in the trainer’s body language. Modern machine learning models are no different. These models are known to be sensitive to spurious correlations between non-essential features of the inputs (e.g., background, texture, and secondary objects) and the corresponding prediction labels. Such features and their correlations with the labels are known as “spurious” because they tend to change with shifts in real-world data distributions, which can negatively impact the model’s generalization and robustness. In this survey, we provide a comprehensive survey of this emerging issue, along with a fine-grained taxonomy of existing state-of-the-art methods for addressing spurious correlations in machine learning models. Additionally, we summarize existing datasets, benchmarks, and metrics to facilitate future research. The paper concludes with a discussion of the broader impacts, the recent advancements, and future challenges in the era of Generative Artificial Intelligence (GenAI), aiming to provide valuable insights for researchers in the related domains of the machine learning community.
AB - Back in the early 20th century, a famous horse named Hans appeared to perform arithmetic and other intellectual tasks during exhibitions in Germany, which has drew widespread attention. However, later studies showed that Hans relied solely on subtle, involuntary cues in the trainer’s body language. Modern machine learning models are no different. These models are known to be sensitive to spurious correlations between non-essential features of the inputs (e.g., background, texture, and secondary objects) and the corresponding prediction labels. Such features and their correlations with the labels are known as “spurious” because they tend to change with shifts in real-world data distributions, which can negatively impact the model’s generalization and robustness. In this survey, we provide a comprehensive survey of this emerging issue, along with a fine-grained taxonomy of existing state-of-the-art methods for addressing spurious correlations in machine learning models. Additionally, we summarize existing datasets, benchmarks, and metrics to facilitate future research. The paper concludes with a discussion of the broader impacts, the recent advancements, and future challenges in the era of Generative Artificial Intelligence (GenAI), aiming to provide valuable insights for researchers in the related domains of the machine learning community.
UR - https://www.scopus.com/pages/publications/105031242654
M3 - Article
AN - SCOPUS:105031242654
SN - 2835-8856
VL - 2026-February
JO - Transactions on Machine Learning Research
JF - Transactions on Machine Learning Research
ER -