TY - GEN
T1 - The Viability of Best-worst Scaling and Categorical Data Label Annotation Tasks in Detecting Implicit Bias
AU - Glenn, Parker
AU - Jacobs, Cassandra L.
AU - Thielk, Marvin
AU - Chu, Yi
N1 - Publisher Copyright:
© European Language Resources Association (ELRA), licensed under CC-BY-NC-4.0.
PY - 2022
Y1 - 2022
N2 - Annotating workplace bias in text is a noisy and subjective task. In encoding the inherently continuous nature of bias, aggregated binary classifications do not suffice. Best-worst scaling (BWS) (Louviere and Woodworth, 1991) offers a framework to obtain real-valued scores through a series of comparative evaluations, but it may be impractical to deploy to traditional annotation pipelines within industry. We present analyses of a small-scale bias dataset, jointly annotated with categorical annotations and BWS annotations and show that there is a strong correlation between observed agreement and BWS score (Spearman’s r=0.72). We identify several shortcomings of BWS relative to traditional categorical annotation: (1) When compared to categorical annotation, we estimate BWS takes approximately 4.5x longer to complete; (2) BWS does not scale well to large annotation tasks with sparse target phenomena; (3) The high correlation between BWS and the traditional task shows that the benefits of BWS can be recovered from a simple categorically annotated, non-aggregated dataset.
AB - Annotating workplace bias in text is a noisy and subjective task. In encoding the inherently continuous nature of bias, aggregated binary classifications do not suffice. Best-worst scaling (BWS) (Louviere and Woodworth, 1991) offers a framework to obtain real-valued scores through a series of comparative evaluations, but it may be impractical to deploy to traditional annotation pipelines within industry. We present analyses of a small-scale bias dataset, jointly annotated with categorical annotations and BWS annotations and show that there is a strong correlation between observed agreement and BWS score (Spearman’s r=0.72). We identify several shortcomings of BWS relative to traditional categorical annotation: (1) When compared to categorical annotation, we estimate BWS takes approximately 4.5x longer to complete; (2) BWS does not scale well to large annotation tasks with sparse target phenomena; (3) The high correlation between BWS and the traditional task shows that the benefits of BWS can be recovered from a simple categorically annotated, non-aggregated dataset.
KW - best-worst scaling
KW - categorical annotation
KW - scalability
UR - https://www.scopus.com/pages/publications/85145880263
M3 - Conference contribution
AN - SCOPUS:85145880263
T3 - 1st Workshop on Perspectivist Approaches to Disagreement in NLP, NLPerspectives 2022 as part of Language Resources and Evaluation Conference, LREC 2022 Workshop
SP - 32
EP - 36
BT - 1st Workshop on Perspectivist Approaches to Disagreement in NLP, NLPerspectives 2022 as part of Language Resources and Evaluation Conference, LREC 2022 Workshop
A2 - Abercrombie, Gavin
A2 - Basile, Valerio
A2 - Tonelli, Sara
A2 - Rieser, Verena
A2 - Uma, Alexandra
PB - European Language Resources Association (ELRA)
T2 - 1st Workshop on Perspectivist Approaches to Disagreement in NLP, NLPerspectives 2022 as part of Language Resources and Evaluation Conference, LREC 2022 Workshop
Y2 - 20 June 2022 through 20 June 2022
ER -