Bilingual Social Media Text Hate Speech Detection For Afaan Oromo and Amharic Languages Using Deep Learning

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

ASTU

Abstract

Hate speech on social media has become a challenging problem in the past few years. Despite the importance of social media for communication and information sharing purposes, it is observable that people are using it as a playground for hate speech against a targeted group based on their protected identity such as religion, ethnic group, and gender. Recently, hate speech have become a common problem also in Ethiopia. The problem can lead to violent actions in society when it is shared by a large community. Detecting hate speech texts on social media is a tedious and complex task due to the unstructured format of social media content. In the past recent years hate speech texts on social media have become more challenging and it attracted several researchers to work on hate speech detection. Due to the success of deep learning algorithms in natural language processing tasks, some researchers proposed deep learning models for hate speech detection. Even if the problem of hate speech is not language-specific, most of the studies are explored only for high resource languages like English except some studies recently proposed also for low resource languages. Also, most works on this problem are focused on a single language, and as a solution, this study proposed bilingual hate speech detection for Afaan Oromo and Amharic text on social media using deep learning. A bilingual dataset prepared from newly collected Afaan Oromo text from the Facebook platform and the existing binary Amharic dataset is adopted to develop models. The newly collected data is obtained by scraping selected Facebook pages using Fcepager graph API and annotated by 5 annotators. The prepared dataset contains binary classes “Hate” and “Free”. Bidirectional RNNs and attention mechanisms are implemented using Word2vec as feature representation. The word2vec model is trained based on skip-gram model due to its suitability for representing non frequent keywords of limited size dataset. Additionally LSTM and GRU networks are also implemented for model comparison. The models are trained using 5-fold and also 10-fold cross validation. The results shows that models achieved good performance when using 5-fold cross validation on our dataset. Then, several experiments are employed to select the best-performing model and finally, the BiLSTM model outperformed all other models with an accuracy of 94.3% and f1_score of 94.2%.

Description

Citation

Collections

Endorsement

Review

Supplemented By

Referenced By