Skip to main navigation Skip to search Skip to main content

The Viability of Best-worst Scaling and Categorical Data Label Annotation Tasks in Detecting Implicit Bias

  • Workhuman Inc.
  • Brandeis University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

2 Scopus citations

Abstract

Annotating workplace bias in text is a noisy and subjective task. In encoding the inherently continuous nature of bias, aggregated binary classifications do not suffice. Best-worst scaling (BWS) (Louviere and Woodworth, 1991) offers a framework to obtain real-valued scores through a series of comparative evaluations, but it may be impractical to deploy to traditional annotation pipelines within industry. We present analyses of a small-scale bias dataset, jointly annotated with categorical annotations and BWS annotations and show that there is a strong correlation between observed agreement and BWS score (Spearman’s r=0.72). We identify several shortcomings of BWS relative to traditional categorical annotation: (1) When compared to categorical annotation, we estimate BWS takes approximately 4.5x longer to complete; (2) BWS does not scale well to large annotation tasks with sparse target phenomena; (3) The high correlation between BWS and the traditional task shows that the benefits of BWS can be recovered from a simple categorically annotated, non-aggregated dataset.

Original languageEnglish
Title of host publication1st Workshop on Perspectivist Approaches to Disagreement in NLP, NLPerspectives 2022 as part of Language Resources and Evaluation Conference, LREC 2022 Workshop
EditorsGavin Abercrombie, Valerio Basile, Sara Tonelli, Verena Rieser, Alexandra Uma
PublisherEuropean Language Resources Association (ELRA)
Pages32-36
Number of pages5
ISBN (Electronic)9791095546986
StatePublished - 2022
Event1st Workshop on Perspectivist Approaches to Disagreement in NLP, NLPerspectives 2022 as part of Language Resources and Evaluation Conference, LREC 2022 Workshop - Marseille, France
Duration: Jun 20 2022Jun 20 2022

Publication series

Name1st Workshop on Perspectivist Approaches to Disagreement in NLP, NLPerspectives 2022 as part of Language Resources and Evaluation Conference, LREC 2022 Workshop

Conference

Conference1st Workshop on Perspectivist Approaches to Disagreement in NLP, NLPerspectives 2022 as part of Language Resources and Evaluation Conference, LREC 2022 Workshop
Country/TerritoryFrance
CityMarseille
Period06/20/2206/20/22

Keywords

  • best-worst scaling
  • categorical annotation
  • scalability

Fingerprint

Dive into the research topics of 'The Viability of Best-worst Scaling and Categorical Data Label Annotation Tasks in Detecting Implicit Bias'. Together they form a unique fingerprint.

Cite this