Development of a Bi Lingual Speech Translator for Igbo to English Using Hubert and Differential Evolution
DOI:
https://doi.org/10.70882/noun-ijcea.2026.1147Keywords:
Speech to speech translation, Igbo language, HuBERT Model, Discrete units, Differential Evolution, Low resource languageAbstract
The language barrier between Igbo and English speakers in Nigeria limits effective communication in education, governance, and social interactions. Most existing speech translation tools rely on text‑based or cascaded architectures that introduce latency, and fail to capture the nuances of spoken Igbo. This paper presents a direct speech‑to‑speech translation (S2ST) system that converts Igbo speech into English speech without intermediate text generation. The system adopts a speech‑to‑unit (S2UT) framework: source Igbo audio is transformed into 80‑dimensional log‑mel features, while target English audio is represented as discrete acoustic units obtained by quantizing HuBERT features. A Transformer‑based encoder–decoder, implemented in fairseq, learns to map the source features to target unit sequences. A HiFi‑GAN vocoder then synthesizes the final English waveform. Three optimization strategies were compared: no optimizer, the Adam Optimizer, and a Differential Evolution (DE) fine‑tuning applied to a subset of the decoder’s output projection weights. Experiments show that the Adam‑trained model achieves a BLEU4 score of 23.97, a Unit Error Rate (UER) of 66.18%, and a token‑level accuracy of 33.69%. The model without any optimizer fails to converge (BLEU4 = 0.15). though the DE fine‑tuning reduced the loss by a little margin but still yields identical performance, confirming that gradient‑based training already reaches a strong local optimum.
References
Abbott, J., & Martinus, L. (2018). Towards neural machine translation for African languages. Proceedings of the 2nd Workshop on African Natural Language Processing. https://doi.org/10.48550/arXiv.1811.05467
Abiola, O. B., Adetunmbi, A. O., & Oguntimilehin, A. (2015). Using hybrid approach for English-to-Yoruba text to text machine translation system. International Journal of Computer Science and Mobile Computing, 4(8), 308–313.
Dossou, B. F. P., & Emezue, C. C. (2021). OkwuGbé: End-to-end speech recognition for Fon and Igbo. arXiv. https://doi.org/10.48550/arXiv.2103.07762
Hsu, W.-N., Bolte, B., Tsai, Y.-H. H., Lakhotia, K., Salakhutdinov, R., & Mohamed, A. (2021). HuBERT: Self-supervised speech representation learning by masked prediction of hidden units. arXiv. https://doi.org/10.48550/arXiv.2106.07447
Jia, Y., Ramanovich, M. T., Remez, T., & Pomerantz, R. (2022). Translatotron 2: High-quality direct speech-to-speech translation with voice preservation. Proceedings of the 39th International Conference on Machine Learning, 10120–10134. https://doi.org/10.48550/arXiv.2107.08661
Jia, Y., Weiss, R. J., Biadsy, F., Macherey, W., Johnson, M., Chen, Z., & Wu, Y. (2019). Direct speech-to-speech translation with a sequence-to-sequence model. Proceedings of Interspeech 2019, 1123–1127. https://doi.org/10.21437/Interspeech.2019-1951
Kong, J., Kim, J., & Bae, J. (2020). HiFi-GAN: Generative adversarial networks for efficient and high-fidelity speech synthesis. Advances in Neural Information Processing Systems, 33, 17022–17033. https://doi.org/10.48550/arXiv.2010.05646
Lee, A., Chen, P.-J., Wang, C., Gu, J., Popuri, S., Ma, X., Polyak, A., Adi, Y., He, Q., Tang, Y., Pino, J., & Hsu, W.-N. (2022a). Direct speech-to-speech translation with discrete units. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3327–3339. https://doi.org/10.18653/v1/2022.acl-long.235
Lee, A., Gong, H., Duquenne, P.-A., Schwenk, H., Chen, P.-J., Wang, C., Popuri, S., Adi, Y., Pino, J., Gu, J., & Hsu, W.-N. (2022b). Textless speech-to-speech translation on real data. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 860–872. https://doi.org/10.18653/v1/2022.naacl-main.63
Li, X., Jia, Y., & Chiu, C.-C. (2022). Textless direct speech-to-speech translation with discrete speech representation. arXiv preprint arXiv:2211.00115. https://doi.org/10.48550/arXiv.2211.00115
Lopez, A. (2008). Statistical machine translation. ACM Computing Surveys, 40(3), Article 8. https://doi.org/10.1145/1380584.1380586
Maltais, M., et al. (2026). NaijaS2ST: A multi-accent benchmark for speech-to-speech translation in low-resource Nigerian languages. arXiv. https://doi.org/10.48550/arXiv.2604.16287
Nwonu, E. G., Nkamaigbo, L. C., & Ezenwafor-Afuecheta, C. I. (2023). Multilingualism and insecurity in Nigeria. International Journal of Advanced Academic Studies, 10(2), 45–58.
Omachonu, G. S. (2012, October 20). Igala language studies and development: Progress, issues and challenges [Paper presentation]. 12th Igala Education Summit, Kogi State University, Anyigba, Nigeria.
Popuri, S., Chen, P.-J., Wang, C., Pino, J., Adi, Y., Gu, J., Hsu, W.-N., & Lee, A. (2022). Enhanced direct speech-to-speech translation using self-supervised pre-training and data augmentation. Proceedings of Interspeech 2022, 5195–5199. https://doi.org/10.21437/Interspeech.2022-11032
Rashidi, S., & Sameti, H. (2025). Improving direct Persian–English speech-to-speech translation with discrete units and synthetic parallel data. arXiv. https://doi.org/10.48550/arXiv.2511.12690
Storn, R., & Price, K. (1997). Differential evolution – A simple and efficient heuristic for global optimization over continuous spaces. Journal of Global Optimization, 11(4), 341–359. https://doi.org/10.1023/A:1008202821328
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008. https://doi.org/10.48550/arXiv.1706.03762
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Frederick Chukwuma Alu, Jumoke Falilat Ajao, Sulaiman Olaniyi Abdulsallam, Maryam Adeola Abdullahi (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

