Boosting Accuracy and Interpretability in Multilingual Hate Speech Detection Through Layer Freezing and Explainable AI
By: Meysam Shirdel Bilehsavar , Negin Mahmoudi , Mohammad Jalili Torkamani and more
Potential Business Impact:
Helps computers spot mean online words better.
Sentiment analysis focuses on identifying the emotional polarity expressed in textual data, typically categorized as positive, negative, or neutral. Hate speech detection, on the other hand, aims to recognize content that incites violence, discrimination, or hostility toward individuals or groups based on attributes such as race, gender, sexual orientation, or religion. Both tasks play a critical role in online content moderation by enabling the detection and mitigation of harmful or offensive material, thereby contributing to safer digital environments. In this study, we examine the performance of three transformer-based models: BERT-base-multilingual-cased, RoBERTa-base, and XLM-RoBERTa-base with the first eight layers frozen, for multilingual sentiment analysis and hate speech detection. The evaluation is conducted across five languages: English, Korean, Japanese, Chinese, and French. The models are compared using standard performance metrics, including accuracy, precision, recall, and F1-score. To enhance model interpretability and provide deeper insight into prediction behavior, we integrate the Local Interpretable Model-agnostic Explanations (LIME) framework, which highlights the contribution of individual words to the models decisions. By combining state-of-the-art transformer architectures with explainability techniques, this work aims to improve both the effectiveness and transparency of multilingual sentiment analysis and hate speech detection systems.
Similar Papers
Advancing Hate Speech Detection with Transformers: Insights from the MetaHate
Machine Learning (CS)
Finds mean online words faster than before.
Multilingual Hate Speech Detection in Social Media Using Translation-Based Approaches with Large Language Models
Computation and Language
Stops online hate speech in Urdu, English, and Spanish.
MemeIntel: Explainable Detection of Propagandistic and Hateful Memes
Computation and Language
Helps computers spot fake news in pictures.