A

DeepHateExplainer: Explainable Hate Speech Detection in Under-resourced Bengali Language

IEEE International Conference on Data Science and Advanced Analytics (DSAA), pp. 1–10

Abstract

In this paper, we propose an explainable approach for hate speech detection from the under-resourced Bengali language, which we called DeepHateExplainer. In our approach, Bengali texts are first comprehensively preprocessed, before classifying them into political, personal, geopolitical, and religious hates using a neural ensemble method of transformer-based neural architectures (i.e., monolingual Bangla BERT-base, multilingual BERT-cased/uncased, and XLM-RoBERTa). Subsequently, important (most and least) terms are identified using sensitivity analysis and layer-wise relevance propagation (LRP), before providing human-interpretable explanations11To foster reproducible research, we make available the data, source codes, models, and notebooks: https://github.com/rezacsedu/DeepHateExplainer. Finally, we compute comprehensiveness and sufficiency scores to measure the quality of explanations w.r.t faithfulness. Evaluations against machine learning (linear and tree-based models) and neural networks (i.e., CNN, Bi-LSTM, and Conv-LSTM with word embeddings) baselines yield F1-scores of 78%, 91%, 89%, and 84%, for political, personal, geopolitical, and religious hates, respectively, outperforming both ML and DNN baselines22Read an extended version of this paper: https://arxiv.org/abs/2012.14353.

Authors 2

  1. Fraunhofer Institute for Applied Information Technology

    Affiliation as printed

    Fraunhofer Institute for Applied Information Technology FIT,Germany

  2. Noakhali Science and Technology University

    Affiliation as printed

    Noakhali Science and Technology University,Noakhali,Bangladesh

Cited by 108 stored of 108

No patents citing this paper on Lens.org (checked 2026-10-06).

References 31