Development of a Bi Lingual Speech Translator for Igbo to English Using Hubert and Differential Evolution

Authors

  • Frederick Chukwuma Alu Kwara State University, Malete Author
  • Jumoke Falilat Ajao Kwara State University, Malete Author
  • Sulaiman Olaniyi Abdulsallam Kwara State University, Malete Author
  • Maryam Adeola Abdullahi Kwara State University, Malete Author

DOI:

https://doi.org/10.70882/noun-ijcea.2026.1147

Keywords:

Speech to speech translation, Igbo language, HuBERT Model, Discrete units, Differential Evolution, Low resource language

Abstract

The language barrier between Igbo and English speakers in Nigeria limits effective communication in education, governance, and social interactions. Most existing speech translation tools rely on text‑based or cascaded architectures that introduce latency, and fail to capture the nuances of spoken Igbo. This paper presents a direct speech‑to‑speech translation (S2ST) system that converts Igbo speech into English speech without intermediate text generation. The system adopts a speech‑to‑unit (S2UT) framework: source Igbo audio is transformed into 80‑dimensional log‑mel features, while target English audio is represented as discrete acoustic units obtained by quantizing HuBERT features. A Transformer‑based encoder–decoder, implemented in fairseq, learns to map the source features to target unit sequences. A HiFi‑GAN vocoder then synthesizes the final English waveform. Three optimization strategies were compared: no optimizer, the Adam Optimizer, and a Differential Evolution (DE) fine‑tuning applied to a subset of the decoder’s output projection weights. Experiments show that the Adam‑trained model achieves a BLEU4 score of 23.97, a Unit Error Rate (UER) of 66.18%, and a token‑level accuracy of 33.69%. The model without any optimizer fails to converge (BLEU4 = 0.15). though the DE fine‑tuning reduced the loss by a little margin but still yields identical performance, confirming that gradient‑based training already reaches a strong local optimum.

References

Abbott, J., & Martinus, L. (2018). Towards neural machine translation for African languages. Proceedings of the 2nd Workshop on African Natural Language Processing. https://doi.org/10.48550/arXiv.1811.05467

Abiola, O. B., Adetunmbi, A. O., & Oguntimilehin, A. (2015). Using hybrid approach for English-to-Yoruba text to text machine translation system. International Journal of Computer Science and Mobile Computing, 4(8), 308–313.

Dossou, B. F. P., & Emezue, C. C. (2021). OkwuGbé: End-to-end speech recognition for Fon and Igbo. arXiv. https://doi.org/10.48550/arXiv.2103.07762

Hsu, W.-N., Bolte, B., Tsai, Y.-H. H., Lakhotia, K., Salakhutdinov, R., & Mohamed, A. (2021). HuBERT: Self-supervised speech representation learning by masked prediction of hidden units. arXiv. https://doi.org/10.48550/arXiv.2106.07447

Jia, Y., Ramanovich, M. T., Remez, T., & Pomerantz, R. (2022). Translatotron 2: High-quality direct speech-to-speech translation with voice preservation. Proceedings of the 39th International Conference on Machine Learning, 10120–10134. https://doi.org/10.48550/arXiv.2107.08661

Jia, Y., Weiss, R. J., Biadsy, F., Macherey, W., Johnson, M., Chen, Z., & Wu, Y. (2019). Direct speech-to-speech translation with a sequence-to-sequence model. Proceedings of Interspeech 2019, 1123–1127. https://doi.org/10.21437/Interspeech.2019-1951

Kong, J., Kim, J., & Bae, J. (2020). HiFi-GAN: Generative adversarial networks for efficient and high-fidelity speech synthesis. Advances in Neural Information Processing Systems, 33, 17022–17033. https://doi.org/10.48550/arXiv.2010.05646

Lee, A., Chen, P.-J., Wang, C., Gu, J., Popuri, S., Ma, X., Polyak, A., Adi, Y., He, Q., Tang, Y., Pino, J., & Hsu, W.-N. (2022a). Direct speech-to-speech translation with discrete units. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3327–3339. https://doi.org/10.18653/v1/2022.acl-long.235

Lee, A., Gong, H., Duquenne, P.-A., Schwenk, H., Chen, P.-J., Wang, C., Popuri, S., Adi, Y., Pino, J., Gu, J., & Hsu, W.-N. (2022b). Textless speech-to-speech translation on real data. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 860–872. https://doi.org/10.18653/v1/2022.naacl-main.63

Li, X., Jia, Y., & Chiu, C.-C. (2022). Textless direct speech-to-speech translation with discrete speech representation. arXiv preprint arXiv:2211.00115. https://doi.org/10.48550/arXiv.2211.00115

Lopez, A. (2008). Statistical machine translation. ACM Computing Surveys, 40(3), Article 8. https://doi.org/10.1145/1380584.1380586

Maltais, M., et al. (2026). NaijaS2ST: A multi-accent benchmark for speech-to-speech translation in low-resource Nigerian languages. arXiv. https://doi.org/10.48550/arXiv.2604.16287

Nwonu, E. G., Nkamaigbo, L. C., & Ezenwafor-Afuecheta, C. I. (2023). Multilingualism and insecurity in Nigeria. International Journal of Advanced Academic Studies, 10(2), 45–58.

Omachonu, G. S. (2012, October 20). Igala language studies and development: Progress, issues and challenges [Paper presentation]. 12th Igala Education Summit, Kogi State University, Anyigba, Nigeria.

Popuri, S., Chen, P.-J., Wang, C., Pino, J., Adi, Y., Gu, J., Hsu, W.-N., & Lee, A. (2022). Enhanced direct speech-to-speech translation using self-supervised pre-training and data augmentation. Proceedings of Interspeech 2022, 5195–5199. https://doi.org/10.21437/Interspeech.2022-11032

Rashidi, S., & Sameti, H. (2025). Improving direct Persian–English speech-to-speech translation with discrete units and synthetic parallel data. arXiv. https://doi.org/10.48550/arXiv.2511.12690

Storn, R., & Price, K. (1997). Differential evolution – A simple and efficient heuristic for global optimization over continuous spaces. Journal of Global Optimization, 11(4), 341–359. https://doi.org/10.1023/A:1008202821328

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008. https://doi.org/10.48550/arXiv.1706.03762

Downloads

Published

2026-09-14

Issue

Section

Articles