Vol. 36 (2026): Volumen 36
Artículos de investigación

Development of a hñähñu–spanish translator using a statistical algorithm for neural networks and vectorized words

José Manuel Cruz Olguín Universidad Autónoma del Estado de Hidalgo
Salvador Santos Romero Universidad Autónoma del Estado de Hidalgo
Virgilio López-Morales Universidad Autónoma del Estado de Hidalgo
Manuel Alejandro Ojeda Misses Tecnólogico Nacional de México

Published 2026-09-30

How to Cite

Cruz Olguín, J. M., Santos Romero, S., López-Morales, V., & Ojeda Misses, M. A. (2026). Development of a hñähñu–spanish translator using a statistical algorithm for neural networks and vectorized words. Acta Universitaria, 36, 1-20. https://doi.org/10.15174/au.2026.4778

Abstract

This work presents a statistical-probabilistic (EyP) learning algorithm for training artificial neural networks, applied to Hñähñu–Spanish lexical translation. The proposed method uses statistical normalization of the prediction error and the normal cumulative distribution function to generate adaptive parameters for updating the synaptic weights and bias of the neural network. For validation, a neural network was trained using a bilingual corpus composed of words represented by fixed-length numerical vectors. The performance of the proposed algorithm was evaluated using the mean squared error and compared with the Backpropagation and Levenberg–Marquardt algorithms. The results show that the EyP algorithm is capable of learning lexical associations between the two languages while maintaining stable behavior during the training process. Furthermore, the findings suggest that the proposed methodology represents a viable alternative for scenarios with limited data availability and demonstrate its potential for the development of computational tools aimed at the preservation and digital revitalization of indigenous languages.

References

  1. Abdar, M., Pourpanah, F., Hussain, S., Rezazadegan, D., Liu, L., Ghavamzadeh, M., Fieguth, P., Cao, X., Khosravi, A., Acharya, U. R., Makarenkov, V., & Nahavandi, S. (2021). A review of uncertainty quantification in deep learning: techniques, applications and challenges. Information Fusion, 76, 243–297. https://doi.org/10.1016/j.inffus.2021.05.008
  2. Abdulkadirov, R., Lyakhov, P., & Nagornov, N. (2023). Survey of optimization algorithms in modern neural networks. Mathematics, 11(11), 2466. https://doi.org/10.3390/math11112466
  3. Afifah, K., Yulita, I. N., & Sarathan, I. (2021). Sentiment analysis on telemedicine app reviews using XGBoost classifier. International Conference on Artificial Intelligence and Big Data Analytics (ICAIBDA), 22–27. https://doi.org/10.1109/ICAIBDA53487.2021.9689735
  4. Aguilar, C. A., & García, H. A. (2023). Tecnologías del lenguaje aplicadas al procesamiento de lenguas indígenas en México: una visión general. Lingüística y Literatura, 44(84), 79–102. https://doi.org/10.17533/udea.lyl.n84a04
  5. Anderson, D. R., Sweeney, D. J., Williams, T. A., Camn, J. D., & Martin, K. (2008). Estadística para administración y economía (10ª ed.). Cengage Learning.
  6. Baruch, I. S., Arellano-Quintana, V., & Pérez-Reynaud, E. (2017). Complex-valued neural network topology and learning applied for identification and control of nonlinear systems. Neurocomputing, 233, 104–115. https://doi.org/10.1016/j.neucom.2016.09.109
  7. Baruch, I. S., & Gaspar, M. C. R. (2009). A Levenberg-Marquardt learning applied for recurrent neural identification and control of a wastewater treatment bioprocess. International Journal of Intelligent Systems, 24, 1094–1114. https://doi.org/10.1002/int.20377
  8. Çelik, Ö., & Koç, B. C. (2021). TF IDF, Word2vec ve Fasttext Vektör Model Yöntemleri ile Türkçe Haber Metinlerinin Sınıflandırılması. Dokuz Eylül Üniversitesi Mühendislik Fakültesi Fen ve Mühendislik Dergisi, 23(67), 121–127. https://doi.org/10.21205/deufmd.2021236710
  9. Cho, K., van Merriënboer, B., Bahdanau, D., & Bengio, Y. (2014a). On the properties of neural machine translation: encoder–decoder approaches. Proceedings of the Eighth Workshop on Syntax, Semantics and Structure in Statistical Translation, 103–111. https://doi.org/10.3115/v1/W14-4012
  10. Cho, K., van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., & Bengio, Y. (2014b). Learning phrase representations using RNN encoder–decoder for statistical machine translation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), 1724–1734. https://doi.org/10.3115/v1/D14-1179
  11. De la Vega, M. (2017). Aprendiendo otomí (Hñähñu). Comisión Nacional para el Desarrollo de los Pueblos Indígenas (CDI). https://www.gob.mx/cms/uploads/attachment/file/302157/aprendiendo-otomi-temoaya-estado-de-mexico-web.pdf
  12. Dyson, L. E., Grant, S., & Hendriks, M. (2016). Indigenous people and mobile technologies. Routledge.
  13. Dyson, L. E., Hendriks, M., & Grant, S. (2007). Information technology and indigenous people. En L. E. Dyson, M. Hendriks, & S. Grant (eds.), Information Technology and Indigenous People (pp. 1–12). IGI Global. https://doi.org/10.4018/978-1-59904-298-5
  14. Escorza-Sánchez, Y. M., Martínez-Martín, G., Saldaña-Tapia, Y., & Maldonado-Catalán, O. (2018). Aplicación móvil para reforzar el aprendizaje de la lengua Hñähñu. Revista de Tecnología y Educación, 2(6), 23–31. https://www.ecorfan.org/republicofperu/research_journals/Revista_de_Tecnologia_y_Educacion/vol2num6/Revista_de_Tecnolog%C3%ADa_y_Educaci%C3%B3n_V2_N6_4.pdf
  15. Fortuin, V. (2022). Priors in Bayesian deep learning: a review. International Statistical Review, 90(3), 563-591. https://doi.org/10.1111/insr.12502
  16. Gallardo, P. (2012). Ritual, palabra y cosmos otomí: yo soy costumbre, yo soy de antigua. UNAM. https://historicas.unam.mx/publicaciones/publicadigital/libros/ritualpalabra/cosmos.html
  17. Gawlikowski, J., Tassi, C. R. N., Ali, M., Lee, J., Humt, M., Feng, J., Kruspe, A., Triebel, R., Jung, P., Roscher, R., Shahzad, M., Yang, W., Bamler, R., & Zhu, X. X. (2023). A survey of uncertainty in deep neural networks. Artificial Intelligence Review, 56, 1513–1589. https://doi.org/10.1007/s10462-023-10562-9
  18. HaCohen-Kerner, Y., Miller, D., & Yigal, Y. (2020). The influence of preprocessing on text classification using a bag-of-words representation. PLOS ONE, 15(5), e0232525. https://doi.org/10.1371/journal.pone.0232525
  19. Hernández, Y., Guzmán, M. G., & Simón, E. (2020). Un libro electrónico en lengua hñähñu para promover el rescate y la ingesta de alimentos saludables. Revista Lengua y Cultura, 1(2), 77–84. https://doi.org/10.29057/lc.v1i2.5432
  20. Hernández-Cruz, L., Victoria-Torquemada, M., & Sinclair-Crawford, D. (2010). Diccionario del Hñähñu (otomí) del Valle del Mezquital, Estado de Hidalgo. http://docencia.uaeh.edu.mx/estudios-pertinencia/docs/hidalgo-municipios/Valle-Del-Mezquital-Diccionario-Hnahnu.pdf
  21. Hinton, L., & Hale, K. (eds.). (2001). The green book of language revitalization in practice. Academic Press. https://www.journals.uchicago.edu/doi/10.1086/424555
  22. Instituto Nacional de Estadística y Geografía (INEGI). (2015). Encuesta Intercensal 2015. Lenguas indígenas nacionales y su distribución geográfica. Instituto Nacional de Estadística y Geografía. https://www.inegi.org.mx/programas/intercen-sal/2015/
  23. Instituto Nacional de Estadística y Geografía (INEGI). (2024). Estadísticas a propósito del día Internacional de los Pueblos Indígenas [Comunicado de prensa núm. 472/24]. https://www.inegi.org.mx/contenidos/saladeprensa/aproposito/2024/EAP_PueblosInd24.pdf
  24. Kabir, H. M. D., Kebria, P. M., Khosravi, A., & Nahavandi, S. (2019). Red neuronal para el cálculo de densidad de probabilidad en series temporales. En 2019 IEEE International Conference on Cloud Computing Technology and Science (CloudCom), Sydney, Australia (pp. 341–346). https://doi.org/10.1109/CloudCom.2019.00059
  25. Khoussi, S., Heckert, A., Battou, A., & Bensalem, S. (2021). A neural networks-based methodology for fitting data to probability distributions. En 2021 IEEE/ACS 18th International Conference on Computer Systems and Applications (AICCSA), Tangier, Morocco (pp. 1–7). IEEE. https://doi.org/10.1109/AICCSA53542.2021.9686821
  26. Kurniawan, F. W., & Maharani, W. (2020). Indonesian Twitter sentiment analysis using Word2Vec. En 2020 International Conference on Data Science and Its Applications (ICoDSA), Bandung, Indonesia (pp. 1-6). https://doi.org/10.1109/icodsa50139.2020.9212906
  27. Lastra, Y. (2010). Los otomíes, su lengua y su historia. IIA-UNAM. https://editorialiia.unam.mx/omp/index.php/publicaciones/catalog/book/otomies_lengua_historia
  28. Li, C., Liu, Y., & Ma, Z. M. (2024). Neural networks taking probability distributions as input: a framework for analyzing exchangeable networks. Neurocomputing, 584, 127572. https://doi.org/10.1016/j.neucom.2024.127572
  29. Linka, K., Schäfer, A., Meng, X., Zou, Z., Karniadakis, G. E., & Kuhl, E. (2022). Bayesian physics-informed neural networks for real world nonlinear dynamical systems. Computer Methods in Applied Mechanics and Engineering, 402, 115346. https://doi.org/10.1016/j.cma.2022.115346
  30. Mohebali, B., Tahmassebi, A., Meyer-Baese, A., & Gandomi, A. H. (2020). Probabilistic neural networks: a brief overview of theory, implementation, and application. En P. Samui, D. T. Bui, S. Chakraborty, & R. C. Deo (eds.), Handbook of probabilistic models (pp. 347–367). Butterworth-Heinemann. https://doi.org/10.1016/B978-0-12-816514-0.00014-X
  31. Nalisnick, E., Smyth, P., & Tran, D. (2023). A brief tour of deep learning from a statistical perspective. Annual Review of Statistics and its Application, 10, 219–246. https://doi.org/10.1146/annurev-statistics-032921-013738
  32. Nurdin, A., Seno, B. A., Bustamin, A., & Abidin, Z. (2020). Perbandingan kinerja word embedding Word2Vec, GloVe, dan FastText pada klasifikasi teks. Jurnal Tekno Kompak, 14(2), 74-79. https://doi.org/10.33365/jtk.v14i2.732
  33. Oakland, J., & Oakland, R. (2024). Statistical process control and data analytics (8a ed.). Routledge. https://doi.org/10.4324/9781003439080
  34. Ojeda, M. A., Solomon, I., & Soria, A. (2019). A real-time identification for hand-based movements using Recurrent Complex-Valued Neural Networks. 2019 IEEE 4th Colombian Conference on Automatic Control (CCAC), Medellin, Colombia (pp. 1–6). https://doi.org/10.1109/CCAC.2019.8920864
  35. Ojeda-Misses, M. A., Martines-Arano, H., López-Morales, V., Franco-Árcega, A., & Márquez-Grajales, A. (2024). Self-tuned closed-loop controller based on statistical data using a servomechanism. 2024 XXVI Robotics Mexican Congress (COMRob), Torreón, Coahuila, México (pp. 27–32). https://doi.org/10.1109/COMRob64055.2024.10777440
  36. Singh, S., Kumar, K., & Kumar, B. (2022). Sentiment analysis of Twitter data using TF IDF and machine learning techniques. En 2022 International Conference on Machine Learning, Big Data, Cloud and Parallel Computing (COM IT CON), Faridabad, India (pp. 252-255). https://doi.org/10.1109/com-it-con54601.2022.9850477
  37. Vargas, I. (2017). Experiencias de un proyecto de revitalización lingüística del hñähñu (otomí) del Valle del Mezquital, Hidalgo: de actores, discursos y prácticas. Zeitschrift für romanische Philologie, 133(4), 1064–1090. https://doi.org/10.1515/zrp-2017-0055