ORCID

Abstract

The growing use of automated hate speech detection systems has sparked significant concern across technical, ethical, and sociopolitical dimensions. While recent advances in natural language processing and machine learning have improved classification performance, they often do so at the expense of transparency, fairness, and user trust. Explainable artificial intelligence (XAI) offers promising tools to bridge this gap, but current methods remain fragmented in scope and limited in stakeholder alignment. This review synthesizes the state of the art in explainability for hate speech detection by integrating machine learning, human-computer interaction, and critical social science perspectives. We survey major XAI approaches, including ante-hoc, post hoc, local, global, counterfactual, and rationale-based methods, and evaluate their applicability across different stages of the ML pipeline and their relevance to key stakeholder groups such as developers, content moderators, policymakers, and affected communities. To guide this synthesis, we propose a conceptual framework that maps the intersection of explanation strategies, pipeline stages, and stakeholder needs. Using this model, we identify persistent gaps in dataset transparency, cultural and linguistic robustness, explanation evaluation practices, and participatory design processes. We argue that achieving meaningful explainability in content moderation goes beyond technical optimization; it demands sociotechnical alignment, contextual sensitivity, and accountability mechanisms that reflect the lived realities of those impacted by algorithmic decisions. The article concludes with interdisciplinary recommendations focused on dataset development, hybrid evaluation benchmarks, inclusive design, and integration. These strategies aim to foster hate speech detection systems that are not only more explainable but also more just, inclusive, and socially grounded. This article is categorized under: Commercial, Legal, and Ethical Issues > Social Considerations Fundamental Concepts of Data and Knowledge > Explainable AI Technologies > Machine Learning.

Keywords

bias mitigation, content moderation, interdisciplinary studies, interpretability, natural language processing, social media analytics, transparency

Publication Date

2026-01-01

Publication Title

Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery

Volume

16

Issue

1

ISSN

1942-4787

Deposit Date

2026-05-11

Funding

This publication has emanated from research conducted with the finan-cial support of Science Foundation Ireland (SFI), now called ResearchIreland (RI), under Grant number 18/CRT/6183. For Open Access, theauthor has applied a CC BY public copyright license to any AuthorAccepted Manuscript version arising from this submission

Creative Commons License

Creative Commons Attribution 4.0 International License
This work is licensed under a Creative Commons Attribution 4.0 International License.


Share

COinS