International Journal of Computer
Trends and Technology

Research Article | Open Access | Download PDF
Volume 74 | Issue 7 | Year 2026 | Article Id. IJCTT-V74I7P105 | DOI : https://doi.org/10.14445/22312803/IJCTT-V74I7P105

A Distilbert Case-Based Deep Learning Model for Part-of-Speech Tagging for Under-Resourced Kenyan Language: A Case of Dholuo


Maureen Otieno, Lilian Wanzare, Calvins Otieno

Received Revised Accepted Published
28 May 2026 30 Jun 2026 17 Jul 2026 31 Jul 2026

Citation :

Maureen Otieno, Lilian Wanzare, Calvins Otieno, "A Distilbert Case-Based Deep Learning Model for Part-of-Speech Tagging for Under-Resourced Kenyan Language: A Case of Dholuo," International Journal of Computer Trends and Technology (IJCTT), vol. 74, no. 7, pp. 52-60, 2026. Crossref, https://doi.org/10.14445/22312803/IJCTT-V74I7P105

Abstract

One of the basic Natural Language Processing (NLP) tasks is Part-of-Speech (POS) tagging, which helps in various applications like sentiment analysis and information retrieval. However, creating accurate POS taggers for low-resource African languages continues to be difficult due to the scarcity of linguistic resources that are annotated. Its contribution is a deep learning method for POS tagging of Dholuo, a less-resourced Western Nilotic language, spoken by about four million people in Kenya and Tanzania. The suggested system uses DistilBERT, a small transformer model, in addition to FastText and Word2Vec vector representations of words that are used to capture the context and meaning of a word. The KenCorpus Dholuo POS dataset was carefully preprocessed, normalized, and standardized with the Universal POS tags and balanced using a hybrid resampling strategy to bring about class representation. The proposed method combines contextual transformer representations with complementary word representations and a training strategy that is optimized for the linguistic features of Dholuo, while previous studies primarily used multilingual transformer models or traditional sequence-labeling methods. The framework developed is a computationally efficient one that is well-suited for low-resource language processing. The results of the experiments reveal that the proposed DistilBERT-based model outperforms the baseline Conditional Random Field (CRF) and Bidirectional Long Short-Term Memory (BiLSTM) models with an accuracy 79.05%, precision 80.63%, recall 79.05% and F1-score 79.24%. To our best knowledge, these results are the best reported for Dholuo POS tagging, and for under-resourced languages in Africa in general, highlighting the suitability of lightweight transformer architectures.

Keywords

Deep Learning, Dholuo , DistilBERT, Natural language processing, POS Tagging.

References

[1] Mary Dibits, and Pius A. Owolawi, and Sunday O. Ojo, “An Hybrid Part of Speech Tagger for Setswana Language using Voting Method,” International Conference on Intelligent and Innovative Computing Applications, pp. 245-253, 2022.
  [
CrossRef] [Google Scholar]

[2] Cheikh M. Bamba Dione et al., “MasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African Languages,” arXiv Preprint, pp. 1-18, 2023.
  [
CrossRef] [Google Scholar] [Publisher Link]

[3] Ikechukwu I. Ayogu et al., “A Comparative Study of Hidden Markov Model and Conditional Random Fields on a Yorùbá part-of-Speech Tagging Task,” 2017 International Conference on Computing Networking and Informatics (ICCNI), Lagos, Nigeria, pp. 1-6, 2017.
  [
CrossRef] [Google Scholar] [Publisher Link]

[4] Tusarkanta Dalai, Tapas Kumar Mishra, and Pankaj K. Sa, “Part-of-Speech Tagging of Odia Language using Statistical and Deep Learning based Approaches,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 22, no. 6, pp. 1-24, 2023.
  [
CrossRef] [Google Scholar] [Publisher Link]

[5] Barack Wanjawa et al., “KenCorpus: A Kenyan Language Corpus of Swahili, Dholuo and Luhya for Natural Language Processing Tasks,” Journal for Language Technology and Computational Linguistics, vol. 36, no. 2, pp. 1-27, 2023.
   [
CrossRef] [Google Scholar] [Publisher Link]

[6] Iheanetu Olamma, Michael Kingsley, and Sunday Ojo, “Hidden Markov-based Part-of-Speech Tagger for Igbo Resource-Scarce African Language,” Proceedings of the First International Workshop on NLP Solutions for Under Resourced Languages,
Association for Computational Linguistics, Trento, Italy, pp. 118-123, 2019.
  [
Google Scholar] [Publisher Link]

[7] Alebachew Chiche, Hiwot Kadi, and Tibebu Bekele, “A Hidden Markov Model-based Part of Speech Tagger for Shekki'noono Language,” International Journal of Computing, vol. 20, no. 4, pp. 587-595, 2021.
  [
CrossRef] [Google Scholar] [Publisher Link]

[8] Benson N. Kituku, Musumba George, and Peter Wagacha, “Kamba Part of Speech Tagger using Memory-based Approach,” International Journal on Natural Language Computing, vol. 4, no. 2, pp. 43-53, 2015.
  [
CrossRef] [Google Scholar]

[9] Rkia Bani et al., “Toward Accurate Amazigh Part-of-Speech Tagging,” IAES International Journal of Artificial Intelligence, vol. 13, no. 1, pp. 572-580, 2024.
  [
CrossRef] [Google Scholar] [Publisher Link]

[10] Mohamed Amine Cheragui, Abdelhalim Hafedh Dahou, and Amin Abdedaiem, “Exploring BERT Models for Part-of-Speech Tagging in the Algerian Dialect: A Comprehensive Study,” Proceedings of the 6th International Conference on Natural Language and Speech Processing (ICNLSP), Association for Computational Linguistics, pp. 140-150, 2023
  [
Google Scholar] [Publisher Link]

[11] Guy De Pauw, Naomi Maajabu, and Peter Waiganjo Wagacha, “A Knowledge-light Approach to Luo Machine Translation and Part-of-Speech Tagging,” Proceedings of the Second Workshop on African Language Technology (AfLaT), Valletta, Malta, pp. 15-20, 2010.
  [
Google Scholar]