Skip to main navigation Skip to search Skip to main content

ALPHA: Action-Based Learning for Pluralistic Human Alignment in Large Language Models

  • Aanisha Bhattacharyya
  • , Susmit Agrawal
  • , Yaman Kumar Singla
  • , Tarun Ram Menta
  • , Nikitha Sr
  • , Rajiv Ratn Shah
  • , Changyou Chen
  • , Balaji Krishnamurthy
  • Adobe Media and Data Science Research (MDSR)
  • Indraprastha Institute of Information Technology Delhi
  • SUNY Buffalo

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Large language models are widely used, yet aligning them with societal values remains challenging. Current approaches often rely on human annotations, which are hard to scale, or synthetic data produced by models that may themselves be misaligned, making it difficult to capture genuine public opinion. This limits scalability and introduces demographic biases that reduce the representativeness and fairness of model behavior. We introduce a novel approach to pluralistic alignment through behavioral learning, grounded in the psychological principle that observed actions exhibit strong consistency with underlying opinions. Specifically, we present ALPHA50M, a dataset of over 50 million samples derived from 1.5 million real-world advertisements and incorporating rich behavioral signals inferred from demographic engagement patterns. Models trained on this data achieve state-of-the-art zero-shot performance on diverse alignment benchmarks spanning cultural reasoning, political views, and social values. We also propose two new benchmarks. OpinionQA-XL aggregates large-scale survey questions covering over 100 societal topics, while GSS evaluates models’ ability to capture temporal shifts in societal opinions across decades. Our results demonstrate that learning from behavioral signals enables models to align with diverse societal values across demographic groups, capture underlying social and cultural norms, and generalize to unseen surveys, topics, and time periods beyond the training distribution. This behavioral learning paradigm offers a scalable and demographically broad alternative to existing alignment techniques.

Original languageEnglish
Title of host publicationProceedings of the AAAI Conference on Artificial Intelligence
EditorsSven Koenig, Chad Jenkins, Matthew E. Taylor
PublisherAssociation for the Advancement of Artificial Intelligence
Pages37249-37258
Number of pages10
Edition44
ISBN (Print)9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067, 9781577359067
DOIs
StatePublished - 2026
Event40th AAAI Conference on Artificial Intelligence, AAAI 2026 - Singapore, Singapore
Duration: Jan 20 2026Jan 27 2026

Publication series

NameProceedings of the AAAI Conference on Artificial Intelligence
Number44
Volume40
ISSN (Print)2159-5399
ISSN (Electronic)2374-3468

Conference

Conference40th AAAI Conference on Artificial Intelligence, AAAI 2026
Country/TerritorySingapore
CitySingapore
Period01/20/2601/27/26

Fingerprint

Dive into the research topics of 'ALPHA: Action-Based Learning for Pluralistic Human Alignment in Large Language Models'. Together they form a unique fingerprint.

Cite this