N-Gram Word Prediction For Afan Oromo Words

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

This study presents word prediction for Afan Oromo words. Word prediction is an application of Natural language processing that is used to auto complete the words that the user types the first letter or letters of a word and the system provides one or more higher probable words. The main objective of this study is to develop word prediction that predict and auto complete the words in the corpus. The developed corpus collected from governmental medias, news, cultural documents, history of societies for the sake of this study only. In order to develop the model, we have used unsupervised machine learning due to lack of large training corpus for this language. The idea behind the approach is to overcome the problem of a bottleneck, while unsupervised approach can be suitable when there is scarcity of training data. This makes our approach suitable for prediction when there is lack of resources. The algorithm that used in this study was N-grams algorithms (Unigram, Bigram and Trigram) for auto completing a word by predicting a correct word in a sentence which saves time and keystrokes of typing and also reduces misspelling. We used small data corpus of Afan Oromo language of different word types to predict correct word with the accuracy as much as possible. The result and finding are promising. We hope that our work will impact current state of understanding for automated Afan Oromo typing. This work describes how we improve word entry information, through word prediction, as an assistive technology for people with motion impairment using the regular keyboard, to eliminate the overhead needed for the learning process. We also present evaluation metrics to compare different models being used in our study. The result argued that WP yields an accuracy of 90% in Unsupervised Machine Learning Approach. The achieved result was encouraging, despite it is less resource requirement. Yet; further experiments using different N-gram approaches that extend this work are needed for a better performance.

Description

Keywords

Citation

Collections

Endorsement

Review

Supplemented By

Referenced By