Multiscript Language Technologies and Digital Inclusion: A Sociotechnical Study of Kazakh Script Conversion
DOI:
https://doi.org/10.63646/PMTI7305Keywords:
Kazakh; multiscript conversion; digital inclusion; sociotechnical systems; language technology; loanword prompts; transformerAbstract
Kazakh is a major Turkic language written and read through Arabic-based, Cyrillic-based, and Latin-based scripts across different communities, regions, archives, and digital platforms. This multiscript condition is often treated as a narrow technical problem of transliteration accuracy, yet it also shapes who can search for public information, access education, preserve family records, participate in e-government, and maintain linguistic identity in data-driven societies. This paper develops a sociotechnical study of Kazakh script conversion by connecting neural conversion methods, loanword-aware prompting, corpus governance, and digital inclusion. Building on a secondary analysis of recent Kazakh multiscript conversion benchmarks, the study reinterprets character error rate (CER) and word error rate (WER) as proxies for accessibility friction, institutional reliability, and cultural continuity. The analysis shows that prompt-constrained Transformer conversion substantially reduces word-level friction across six conversion directions but also reveals that model accuracy alone is insufficient for inclusive deployment. Script conversion systems affect people through interface design, standardization choices, metadata practices, education policies, data rights, and community trust. The paper contributes a multilayer framework that links script ecology, linguistic resources, conversion services, access settings, governance arrangements, and inclusion outcomes. It further proposes design principles for transparent, auditable, and community-sensitive Kazakh language technologies. The findings suggest that multiscript conversion should be governed not merely as an automation service but as digital public infrastructure for linguistic equity.
How to Cite
References
Star, S. L., & Ruhleder, K. (1996). Steps toward an ecology of infrastructure: Design and access for large information spaces. Information Systems Research, 7(1), 111-134. https://doi.org/10.1287/isre.7.1.111
Kornai, A. (2013). Digital language death. PLOS ONE, 8(10), e77056. https://doi.org/10.1371/journal.pone.0077056
Bahdanau, D., Cho, K., & Bengio, Y. (2015). Neural machine translation by jointly learning to align and translate. International Conference on Learning Representations. https://doi.org/10.48550/arXiv.1409.0473
Dong, J., Jiang, T., Cheng, L., Anwar, A., & Yang, Y. (2018). A compromise Arabic-Kazakh coded character processing method based on the OpenType font format. Computer Standards & Interfaces, 55, 1-7. https://doi.org/10.1016/j.csi.2017.02.005
Karimi, S., Scholer, F., & Turpin, A. (2011). Machine transliteration survey. ACM Computing Surveys, 43(3), Article 17. https://doi.org/10.1145/1922649.1922654
Al-Onaizan, Y., & Knight, K. (2002). Machine transliteration of names in Arabic texts. Proceedings of the ACL-02 Workshop on Computational Approaches to Semitic Languages. https://doi.org/10.3115/1118637.1118642
Akhmed-Zaki, D., Mansurova, M., Madiyeva, G., Kadyrbek, N., & Kyrgyzbayeva, M. (2021). Development of the information system for the Kazakh language preprocessing. Cogent Engineering, 8(1), 1896418. https://doi.org/10.1080/23311916.2021.1896418
Bapna, A., & Firat, O. (2019). Simple, scalable adaptation for neural machine translation. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, 1538-1548. https://doi.org/10.18653/v1/D19-1165
Kumaran, A., & Kellner, T. (2007). A generic framework for machine transliteration. Proceedings of the 30th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 721-722. https://doi.org/10.1145/1277741.1277876
Knight, K., & Graehl, J. (1998). Machine transliteration. Computational Linguistics, 24(4), 599-612. https://doi.org/10.5555/972764.972767
Sennrich, R., Haddow, B., & Birch, A. (2016). Neural machine translation of rare words with subword units. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, 1715-1725. https://doi.org/10.18653/v1/P16-1162
Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., & Bengio, Y. (2014). Learning phrase representations using RNN encoder-decoder for statistical machine translation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, 1724-1734. https://doi.org/10.3115/v1/D14-1179
Hargittai, E. (2002). Second-level digital divide: Differences in people's online skills. First Monday, 7(4). https://doi.org/10.5210/fm.v7i4.942
Selwyn, N. (2004). Reconsidering political and popular understandings of the digital divide. New Media & Society, 6(3), 341-362. https://doi.org/10.1177/1461444804042519
Helsper, E. J. (2012). A corresponding fields model for the links between social and digital exclusion. Communication Theory, 22(4), 403-426. https://doi.org/10.1111/j.1468-2885.2012.01416.x
van Deursen, A. J. A. M., & van Dijk, J. A. G. M. (2014). The digital divide shifts to differences in usage. New Media & Society, 16(3), 507-526. https://doi.org/10.1177/1461444813487959
Scheerder, A., van Deursen, A., & van Dijk, J. (2017). Determinants of internet skills, uses and outcomes: A systematic review of the second- and third-level digital divide. Telematics and Informatics, 34(8), 1607-1624. https://doi.org/10.1016/j.tele.2017.07.007
Orlikowski, W. J. (1992). The duality of technology: Rethinking the concept of technology in organizations. Organization Science, 3(3), 398-427. https://doi.org/10.1287/orsc.3.3.398
Joshi, P., Santy, S., Budhiraja, A., Bali, K., & Choudhury, M. (2020). The state and fate of linguistic diversity and inclusion in the NLP world. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 6282-6293. https://doi.org/10.18653/v1/2020.acl-main.560
Ranathunga, S., Lee, E. S. A., Skenduli, M. P., Shekhar, R., Alam, M., & Kaur, R. (2023). Neural machine translation for low-resource languages: A survey. ACM Computing Surveys, 55(11), Article 229. https://doi.org/10.1145/3567592
Haddow, B., Bawden, R., Miceli Barone, A. V., Helcl, J., & Birch, A. (2022). Survey of low-resource machine translation. Computational Linguistics, 48(3), 673-732. https://doi.org/10.1162/coli_a_00446
Dror, R., Baumer, G., Shlomov, S., & Reichart, R. (2018). The hitchhiker's guide to testing statistical significance in natural language processing. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, 1383-1392. https://doi.org/10.18653/v1/P18-1128
Galla, C. K. (2016). Indigenous language revitalization, promotion, and education: Function of digital technology. Computer Assisted Language Learning, 29(7), 1137-1151. https://doi.org/10.1080/09588221.2016.1166137
Toiganbayeva, N., Kasem, M., Abdimanap, G., Bostanbekov, K., Abdallah, A., Alimova, A., & Nurseitov, D. (2022). KOHTD: Kazakh offline handwritten text dataset. Signal Processing: Image Communication, 108, 116827. https://doi.org/10.1016/j.image.2022.116827
Yeshpanov, R., Polonskaya, A., & Varol, H. A. (2024). KazParC: Kazakh parallel corpus for machine translation. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, 9633-9644. https://doi.org/10.18653/v1/2024.lrec-main.842
Veitsman, Y., & Hartmann, M. (2025). Recent advancements and challenges of Turkic Central Asian language processing. Proceedings of the First Workshop on Language Models for Low-Resource Languages, 309-324. https://doi.org/10.18653/v1/2025.loreslm-1.25
Senel, L. K., Ebing, B., Baghirova, K., Schuetze, H., & Glavas, G. (2024). Kardes-NLU: Transfer to low-resource languages with the help of a high-resource cousin. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics, 1776-1790. https://doi.org/10.18653/v1/2024.eacl-long.100
Orlikowski, W. J. (2000). Using technology and constituting structures: A practice lens for studying technology in organizations. Organization Science, 11(4), 404-428. https://doi.org/10.1287/orsc.11.4.404.14600
Geels, F. W. (2004). From sectoral systems of innovation to socio-technical systems: Insights about dynamics and change from sociology and institutional theory. Research Policy, 33(6-7), 897-920. https://doi.org/10.1016/j.respol.2004.01.015
Markard, J., Raven, R., & Truffer, B. (2012). Sustainability transitions: An emerging field of research and its prospects. Research Policy, 41(6), 955-967. https://doi.org/10.1016/j.respol.2012.02.013
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency, 220-229. https://doi.org/10.1145/3287560.3287596
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daume III, H., & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86-92. https://doi.org/10.1145/3458723
Papineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002). BLEU: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311-318. https://doi.org/10.3115/1073083.1073135
Snover, M., Dorr, B., Schwartz, R., Micciulla, L., & Makhoul, J. (2006). A study of translation edit rate with targeted human annotation. Proceedings of AMTA 2006, 223-231. https://doi.org/10.3115/1626355.1626389
Post, M. (2018). A call for clarity in reporting BLEU scores. Proceedings of the Third Conference on Machine Translation, 186-191. https://doi.org/10.18653/v1/W18-6319
Marie, B., Fujita, A., & Rubino, R. (2021). Scientific credibility of machine translation research: A meta-evaluation of 769 papers. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics, 7297-7306. https://doi.org/10.18653/v1/2021.acl-long.566
Freitag, M., Foster, G., Grangier, D., Ratnakar, V., Tan, Q., & Macherey, W. (2021). Experts, errors, and context: A large-scale study of human evaluation for machine translation. Transactions of the Association for Computational Linguistics, 9, 1460-1474. https://doi.org/10.1162/tacl_a_00437
Mathur, N., Baldwin, T., & Cohn, T. (2020). Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 4984-4997. https://doi.org/10.18653/v1/2020.acl-main.448
Rei, R., Stewart, C., Farinha, A. C., & Lavie, A. (2020). COMET: A neural framework for MT evaluation. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, 2685-2702. https://doi.org/10.18653/v1/2020.emnlp-main.213
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2020). BERTScore: Evaluating text generation with BERT. International Conference on Learning Representations. https://doi.org/10.48550/arXiv.1904.09675
Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D., & Barnes, P. (2020). Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 33-44. https://doi.org/10.1145/3351095.3372873
Holstein, K., Wortman Vaughan, J., Daume III, H., Dudik, M., & Wallach, H. (2019). Improving fairness in machine learning systems: What do industry practitioners need? Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 1-16. https://doi.org/10.1145/3290605.3300830
Selbst, A. D., Boyd, D., Friedler, S. A., Venkatasubramanian, S., & Vertesi, J. (2019). Fairness and abstraction in sociotechnical systems. Proceedings of the Conference on Fairness, Accountability, and Transparency, 59-68. https://doi.org/10.1145/3287560.3287598
Goyal, N., Gao, C., Chaudhary, V., Chen, P. J., Wenzek, G., Ju, D., Krishnan, S., Ranzato, M., Guzman, F., & Fan, A. (2022). The FLORES-101 evaluation benchmark for low-resource and multilingual machine translation. Transactions of the Association for Computational Linguistics, 10, 522-538. https://doi.org/10.1162/tacl_a_00474
Johnson, M., Schuster, M., Le, Q. V., Krikun, M., Wu, Y., Chen, Z., Thorat, N., Viegas, F., Wattenberg, M., Corrado, G., Hughes, M., & Dean, J. (2017). Google's multilingual neural machine translation system: Enabling zero-shot translation. Transactions of the Association for Computational Linguistics, 5, 339-351. https://doi.org/10.1162/tacl_a_00065
Costa-jussa, M. R., Cross, J., Celebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., Sun, A., Wang, S., Wenzek, G., Youngblood, A., Akula, B., Barrault, L., Mejia-Gonzalez, G., Hansanti, P., Hoffman, J., Jarrett, S., et al. (2022). No language left behind: Scaling human-centered machine translation. arXiv. https://doi.org/10.48550/arXiv.2207.04672
Guzman, F., Chen, P. J., Ott, M., Pino, J., Lample, G., Koehn, P., Chaudhary, V., & Ranzato, M. (2019). The FLORES evaluation datasets for low-resource machine translation: Nepali-English and Sinhala-English. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 6098-6111. https://doi.org/10.18653/v1/D19-1632
Kreutzer, J., Caswell, I., Wang, L., Wahab, A., van Esch, D., Ulzii-Orshikh, N., Tapo, A., Subramani, N., Sokolov, A., Sikasote, C., Setyawan, M., Sarin, S., Samb, S., Sagot, B., Rivera, C., Rios, A., Papadimitriou, I., Osei, S., Ortiz Suarez, P., et al. (2022). Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics, 10, 50-72. https://doi.org/10.1162/tacl_a_00447
Nekoto, W., Marivate, V., Matsila, T., Fasubaa, T., Kolawole, T., Fagbohungbe, T., Akinola, S. O., Muhammad, S. H., Kabongo, S., Osei, S., Freshia, S., Niyongabo, R. A., Macharm, R., Ogayo, P., Ahia, O., et al. (2020). Participatory research for low-resourced machine translation: A case study in African languages. Findings of the Association for Computational Linguistics: EMNLP 2020, 2144-2160. https://doi.org/10.18653/v1/2020.findings-emnlp.195
Mager, M., Gutierrez-Vasques, X., Sierra, G., & Meza, I. (2018). Challenges of language technologies for the indigenous languages of the Americas. Proceedings of the 27th International Conference on Computational Linguistics, 55-69. https://doi.org/10.18653/v1/C18-1006
Tiedemann, J., & Thottingal, S. (2020). OPUS-MT: Building open translation services for the world. Proceedings of the 22nd Annual Conference of the European Association for Machine Translation, 479-480. https://doi.org/10.18653/v1/2020.eamt-1.61
Mager, M., Oncevay, A., Ebrahimi, A., Ortega, J., Rios, A., Fan, A., Gutierrez-Vasques, X., Chiruzzo, L., Gimenez-Lugo, G., Ramos, R., et al. (2021). Findings of the AmericasNLP 2021 shared task on open machine translation for indigenous languages of the Americas. Proceedings of the First Workshop on Natural Language Processing for Indigenous Languages of the Americas, 202-217. https://doi.org/10.18653/v1/2021.americasnlp-1.23
Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610-623. https://doi.org/10.1145/3442188.3445922
Bird, S. (2020). Decolonising speech and language technology. Proceedings of the 28th International Conference on Computational Linguistics, 3504-3519. https://doi.org/10.18653/v1/2020.coling-main.313
Blodgett, S. L., Barocas, S., Daume III, H., & Wallach, H. (2020). Language (technology) is power: A critical survey of bias in NLP. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5454-5476. https://doi.org/10.18653/v1/2020.acl-main.485
Wu, S., & Dredze, M. (2020). Are all languages created equal in multilingual BERT? Proceedings of the 5th Workshop on Representation Learning for NLP, 120-130. https://doi.org/10.18653/v1/2020.repl4nlp-1.16
Dabre, R., Chu, C., & Kunchukuttan, A. (2020). A survey of multilingual neural machine translation. ACM Computing Surveys, 53(5), Article 99. https://doi.org/10.1145/3406095
Aharoni, R., Johnson, M., & Firat, O. (2019). Massively multilingual neural machine translation. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, 3874-3884. https://doi.org/10.18653/v1/N19-1388
Lample, G., Ott, M., Conneau, A., Denoyer, L., & Ranzato, M. (2018). Phrase-based and neural unsupervised machine translation. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 5039-5049. https://doi.org/10.18653/v1/D18-1549
Firat, O., Sankaran, B., Al-Onaizan, Y., Vural, F. T. Y., & Cho, K. (2016). Zero-resource translation with multi-lingual neural machine translation. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 268-277. https://doi.org/10.18653/v1/D16-1026
Zoph, B., Yuret, D., May, J., & Knight, K. (2016). Transfer learning for low-resource neural machine translation. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 1568-1575. https://doi.org/10.18653/v1/D16-1163
Koehn, P., & Knowles, R. (2017). Six challenges for neural machine translation. Proceedings of the First Workshop on Neural Machine Translation, 28-39. https://doi.org/10.18653/v1/W17-3204
Koehn, P., Hoang, H., Birch, A., Callison-Burch, C., Federico, M., Bertoldi, N., Cowan, B., Shen, W., Moran, C., Zens, R., Dyer, C., Bojar, O., Constantin, A., & Herbst, E. (2007). Moses: Open source toolkit for statistical machine translation. Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics Companion Volume, 177-180. https://doi.org/10.3115/1557769.1557821
Junczys-Dowmunt, M., Grundkiewicz, R., Dwojak, T., Hoang, H., Heafield, K., Neckermann, T., Seide, F., Germann, U., Fikri Aji, A., Bogoychev, N., Martins, A. F. T., & Birch, A. (2018). Marian: Fast neural machine translation in C++. Proceedings of ACL 2018, System Demonstrations, 116-121. https://doi.org/10.18653/v1/P18-4020
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., & Auli, M. (2019). fairseq: A fast, extensible toolkit for sequence modeling. Proceedings of NAACL-HLT 2019: Demonstrations, 48-53. https://doi.org/10.18653/v1/N19-4009
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., et al. (2020). Transformers: State-of-the-art natural language processing. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 38-45. https://doi.org/10.18653/v1/2020.emnlp-demos.6
Pires, T., Schlinger, E., & Garrette, D. (2019). How multilingual is multilingual BERT? Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 4996-5001. https://doi.org/10.18653/v1/P19-1493
Ruder, S., Vulic, I., & Sogaard, A. (2019). A survey of cross-lingual word embedding models. Journal of Artificial Intelligence Research, 65, 569-631. https://doi.org/10.1613/jair.1.11640
Lu, Y. (2019). Artificial intelligence: A survey on evolution, models, applications and future trends. Journal of Management Analytics, 6(1), 1-29. https://doi.org/10.1080/23270012.2019.1570365
Chen, Y., Lu, Y., Bulysheva, L., & Kataev, M. Y. (2024). Applications of blockchain in Industry 4.0: A review. Information Systems Frontiers, 26(5), 1715-1729. https://doi.org/10.1007/s10796-022-10248-7
Zhang, C., & Lu, Y. (2021). Study on artificial intelligence: The state of the art and future prospects. Journal of Industrial Information Integration, 23, 100224. https://doi.org/10.1016/j.jii.2021.100224
Kou, G., & Lu, Y. (2025). FinTech: A literature review of emerging financial technologies and applications. Financial Innovation, 11, Article 1. https://doi.org/10.1186/s40854-024-00668-6
Xu, L. D., Lu, Y., & Li, L. (2021). Embedding blockchain technology into IoT for security: A survey. IEEE Internet of Things Journal, 8(13), 10452-10473. https://doi.org/10.1109/JIOT.2021.3060508