Bilingual Social Media Text Hate Speech Detection For Afaan Oromo and Amharic Languages Using Deep Learning
Loading...
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
ASTU
Abstract
Hate speech on social media has become a challenging problem in the past few years. Despite the
importance of social media for communication and information sharing purposes, it is observable that
people are using it as a playground for hate speech against a targeted group based on their protected
identity such as religion, ethnic group, and gender. Recently, hate speech have become a common
problem also in Ethiopia. The problem can lead to violent actions in society when it is shared by a
large community. Detecting hate speech texts on social media is a tedious and complex task due to the
unstructured format of social media content. In the past recent years hate speech texts on social media
have become more challenging and it attracted several researchers to work on hate speech detection.
Due to the success of deep learning algorithms in natural language processing tasks, some researchers
proposed deep learning models for hate speech detection. Even if the problem of hate speech is not
language-specific, most of the studies are explored only for high resource languages like English
except some studies recently proposed also for low resource languages. Also, most works on this
problem are focused on a single language, and as a solution, this study proposed bilingual hate speech
detection for Afaan Oromo and Amharic text on social media using deep learning. A bilingual dataset
prepared from newly collected Afaan Oromo text from the Facebook platform and the existing binary
Amharic dataset is adopted to develop models. The newly collected data is obtained by scraping
selected Facebook pages using Fcepager graph API and annotated by 5 annotators. The prepared
dataset contains binary classes “Hate” and “Free”. Bidirectional RNNs and attention mechanisms
are implemented using Word2vec as feature representation. The word2vec model is trained based on
skip-gram model due to its suitability for representing non frequent keywords of limited size dataset.
Additionally LSTM and GRU networks are also implemented for model comparison. The models are
trained using 5-fold and also 10-fold cross validation. The results shows that models achieved good
performance when using 5-fold cross validation on our dataset. Then, several experiments are
employed to select the best-performing model and finally, the BiLSTM model outperformed all other
models with an accuracy of 94.3% and f1_score of 94.2%.
