ORCID

Abstract

The effectiveness of supervised machine learning models is heavily influenced by the quality of training data, which is often shaped by human annotators. Subjective NLP tasks such as hate speech detection, toxicity identification, and sexism classification frequently exhibit annotator disagreement due to differences in individual perspectives. This study investigates annotator disagreement in sexism detection using English tweets from the EXIST 2023 competition. To systematically analyse disagreement, tweets are categorised based on annotator consensus levels, examining how annotator demographics and linguistic features contribute to labelling inconsistencies. We interpret disagreement patterns using Shapley Additive Explanations (SHAP) and assess the consistency of SHAP-derived feature importance rankings via Spearman Rank Correlation. Our findings demonstrate that both annotator demographics and tweet characteristics significantly shape disagreement, reinforcing the need for perspectivist approaches in NLP by showing that annotator disagreement is not just noise but a meaningful signal that should be incorporated into dataset construction.

Keywords

Annotator Disagreement, Disagreement-Aware Learning, Perspectivist NLP, Sexism Detection, SHAP, Subjective NLP, XAI

Publication Date

2026-01-01

Event

3rd World Conference on Explainable Artificial Intelligence, xAI 2025

Publication Title

Explainable Artificial Intelligence - 3rd World Conference, xAI 2025, Proceedings

Publisher

Springer Science and Business Media Deutschland GmbH

ISBN

9783032083326

ISSN

1865-0929

First Page

201

Last Page

224

Deposit Date

2026-01-20

Creative Commons License

Creative Commons Attribution 4.0 International License
This work is licensed under a Creative Commons Attribution 4.0 International License.


Share

COinS