Research Article | Open Access | Download PDF
Volume 74 | Issue 7 | Year 2026 | Article Id. IJCTT-V74I7P105 | DOI : https://doi.org/10.14445/22312803/IJCTT-V74I7P105A Distilbert Case-Based Deep Learning Model for Part-of-Speech Tagging for Under-Resourced Kenyan Language: A Case of Dholuo
Maureen Otieno, Lilian Wanzare, Calvins Otieno
| Received | Revised | Accepted | Published |
|---|---|---|---|
| 28 May 2026 | 30 Jun 2026 | 17 Jul 2026 | 31 Jul 2026 |
Citation :
Maureen Otieno, Lilian Wanzare, Calvins Otieno, "A Distilbert Case-Based Deep Learning Model for Part-of-Speech Tagging for Under-Resourced Kenyan Language: A Case of Dholuo," International Journal of Computer Trends and Technology (IJCTT), vol. 74, no. 7, pp. 52-60, 2026. Crossref, https://doi.org/10.14445/22312803/IJCTT-V74I7P105
Abstract
One of the basic Natural Language Processing (NLP) tasks is Part-of-Speech (POS) tagging, which helps in various applications like sentiment analysis and information retrieval. However, creating accurate POS taggers for low-resource African languages continues to be difficult due to the scarcity of linguistic resources that are annotated. Its contribution is a deep learning method for POS tagging of Dholuo, a less-resourced Western Nilotic language, spoken by about four million people in Kenya and Tanzania. The suggested system uses DistilBERT, a small transformer model, in addition to FastText and Word2Vec vector representations of words that are used to capture the context and meaning of a word. The KenCorpus Dholuo POS dataset was carefully preprocessed, normalized, and standardized with the Universal POS tags and balanced using a hybrid resampling strategy to bring about class representation. The proposed method combines contextual transformer representations with complementary word representations and a training strategy that is optimized for the linguistic features of Dholuo, while previous studies primarily used multilingual transformer models or traditional sequence-labeling methods. The framework developed is a computationally efficient one that is well-suited for low-resource language processing. The results of the experiments reveal that the proposed DistilBERT-based model outperforms the baseline Conditional Random Field (CRF) and Bidirectional Long Short-Term Memory (BiLSTM) models with an accuracy 79.05%, precision 80.63%, recall 79.05% and F1-score 79.24%. To our best knowledge, these results are the best reported for Dholuo POS tagging, and for under-resourced languages in Africa in general, highlighting the suitability of lightweight transformer architectures.
Keywords
Deep Learning, Dholuo , DistilBERT, Natural language processing, POS Tagging.
References
[1] Mary Dibits, and Pius A. Owolawi, and Sunday O. Ojo, “An
Hybrid Part of Speech Tagger for Setswana Language using Voting Method,” International
Conference on Intelligent and Innovative Computing Applications, pp.
245-253, 2022.
[CrossRef] [Google Scholar]
[2] Cheikh M. Bamba Dione et al., “MasakhaPOS: Part-of-Speech
Tagging for Typologically Diverse African Languages,” arXiv Preprint,
pp. 1-18, 2023.
[CrossRef] [Google Scholar] [Publisher
Link]
[3] Ikechukwu I. Ayogu et al., “A Comparative Study of Hidden
Markov Model and Conditional Random Fields on a Yorùbá part-of-Speech Tagging
Task,” 2017 International Conference on Computing Networking and Informatics
(ICCNI), Lagos, Nigeria, pp. 1-6, 2017.
[CrossRef] [Google Scholar] [Publisher
Link]
[4] Tusarkanta Dalai, Tapas Kumar Mishra, and Pankaj K. Sa,
“Part-of-Speech Tagging of Odia Language using Statistical and Deep Learning
based Approaches,” ACM Transactions on Asian and Low-Resource Language
Information Processing, vol. 22, no. 6, pp. 1-24, 2023.
[CrossRef] [Google Scholar] [Publisher
Link]
[5] Barack Wanjawa et al., “KenCorpus: A Kenyan Language Corpus
of Swahili, Dholuo and Luhya for Natural Language Processing Tasks,” Journal
for Language Technology and Computational Linguistics, vol. 36, no. 2, pp.
1-27, 2023.
[CrossRef] [Google Scholar] [Publisher
Link]
[6] Iheanetu Olamma, Michael Kingsley, and Sunday Ojo, “Hidden
Markov-based Part-of-Speech Tagger for Igbo Resource-Scarce African Language,” Proceedings
of the First International Workshop on NLP Solutions for Under Resourced
Languages,
Association for Computational Linguistics, Trento, Italy, pp. 118-123, 2019.
[Google Scholar] [Publisher
Link]
[7] Alebachew Chiche, Hiwot Kadi, and Tibebu Bekele, “A Hidden
Markov Model-based Part of Speech Tagger for Shekki'noono Language,” International
Journal of Computing, vol. 20, no. 4, pp. 587-595, 2021.
[CrossRef] [Google Scholar] [Publisher Link]
[8] Benson N. Kituku, Musumba George, and Peter Wagacha, “Kamba
Part of Speech Tagger using Memory-based Approach,” International Journal on
Natural Language Computing, vol. 4, no. 2, pp. 43-53, 2015.
[CrossRef] [Google Scholar]
[9] Rkia Bani et al., “Toward Accurate Amazigh Part-of-Speech
Tagging,” IAES International Journal of Artificial Intelligence, vol.
13, no. 1, pp. 572-580, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[10] Mohamed Amine Cheragui, Abdelhalim Hafedh Dahou, and Amin
Abdedaiem, “Exploring BERT Models for Part-of-Speech Tagging in the Algerian
Dialect: A Comprehensive Study,” Proceedings of the 6th
International Conference on Natural Language and Speech Processing (ICNLSP), Association for Computational Linguistics, pp. 140-150, 2023
[Google Scholar] [Publisher
Link]
[11] Guy De Pauw, Naomi Maajabu, and Peter Waiganjo Wagacha, “A
Knowledge-light Approach to Luo Machine Translation and Part-of-Speech
Tagging,” Proceedings of the Second Workshop on African Language Technology
(AfLaT), Valletta, Malta, pp. 15-20, 2010.
[Google Scholar]