Skip to main navigation Skip to search Skip to main content

Improving Dialog Safety using Socially Aware Contrastive Learning

  • SUNY Buffalo

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

State-of-the-art conversational AI systems raise concerns due to their potential risks of generating unsafe, toxic, unethical, or dangerous content. Previous works have developed datasets to teach conversational agents the appropriate social paradigms to respond effectively to specifically designed hazardous content. However, models trained on these adversarial datasets still struggle to recognize subtle unsafe situations that appear naturally in conversations or introduce an inappropriate response in a casual context. To understand the extent of this problem, we study prosociality in both adversarial and casual dialog contexts and audit the response quality of general-purpose language models in terms of propensity to produce unsafe content. We propose a dual-step fine-tuning process to address these issues using a socially aware n-pair contrastive loss. Subsequently, we train a base model that integrates prosocial behavior by leveraging datasets like Moral Integrity Corpus (MIC) and PROSOCIALDIALOG. Experimental results on several dialog datasets demonstrate the effectiveness of our approach in generating socially appropriate responses.

Original languageEnglish
Title of host publicationSCI-CHAT 2024 - Workshop on Simulating Conversational Intelligence in Chat, Proceedings of the Workshop
EditorsYvette Graham, Qun Liu, Gerasimos Lampouras, Ignacio Iacobacci, Sinead Madden, Haider Khalid, Rameez Qureshi
PublisherAssociation for Computational Linguistics (ACL)
Pages4-18
Number of pages15
ISBN (Electronic)9798891760820
StatePublished - 2024
Event1st Workshop on Simulating Conversational Intelligence in Chat, SCI-CHAT 2024 - St. Julian's, Malta
Duration: Mar 21 2024 → …

Publication series

NameSCI-CHAT 2024 - Workshop on Simulating Conversational Intelligence in Chat, Proceedings of the Workshop

Conference

Conference1st Workshop on Simulating Conversational Intelligence in Chat, SCI-CHAT 2024
Country/TerritoryMalta
CitySt. Julian's
Period03/21/24 → …

Fingerprint

Dive into the research topics of 'Improving Dialog Safety using Socially Aware Contrastive Learning'. Together they form a unique fingerprint.

Cite this