Building a Large-Scale Cross-Script Kazakh Parallel Corpus for Low-Resource Language Data Science
DOI:
https://doi.org/10.63646/GTCU8198Keywords:
Kazakh language; parallel corpus; cross-script data; low-resource NLP; data governance; script conversionAbstract
Kazakh is a low-resource Turkic language whose digital use is complicated by the long-term coexistence of Arabic-based, Cyrillic-based, and Latin-based scripts. Existing conversion models show that script diversity, vowel harmony, consonant alternation, regional vocabulary, and loanwords jointly create a data problem rather than only an algorithmic problem. This study presents CrossScriptKaz-1.2M, a large-scale cross-script Kazakh parallel corpus designed for low-resource language data science. The corpus integrates Arabic-script, Cyrillic-script, and Latin-script materials from news, education, culture, public information, and community web sources, and applies a reproducible pipeline for script normalization, sentence segmentation, cross-script alignment, metadata enrichment, loanword tagging, and manual quality auditing. The final resource contains 1,184,260 aligned sentence triples, 27.6 million normalized tokens, 82,416 validated loanword entries, and document-level provenance metadata. Validation results indicate 97.4% alignment precision, 98.1% script-label accuracy, and clear performance gains in downstream script conversion experiments. A corpus-guided Transformer reduces average character error rate from 3.04% to 1.92% compared with a strong Transformer baseline. The contribution is a scalable data architecture and evaluation protocol that supports cross-script conversion, multilingual modeling, corpus linguistics, and responsible data governance for underrepresented languages.
How to Cite
References
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M., Levenberg, J., Monga, R., Moore, S., Murray, D. G., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., ... Zheng, X. (2016). TensorFlow: A system for large-scale machine learning. In Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (pp. 265-283). USENIX Association. https://doi.org/10.5555/3026877.3026899
Aharoni, R., Johnson, M., & Firat, O. (2019). Massively multilingual neural machine translation. In Proceedings of NAACL-HLT 2019 (pp. 3874-3884). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1388
Artetxe, M., & Schwenk, H. (2019). Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond. Transactions of the Association for Computational Linguistics, 7, 597-610. https://doi.org/10.1162/tacl_a_00288
Banerjee, S., & Lavie, A. (2005). METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization (pp. 65-72). Association for Computational Linguistics. https://doi.org/10.5555/1626355.1626389
Bender, E. M., & Friedman, B. (2018). Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6, 587-604. https://doi.org/10.1162/tacl_a_00041
Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610-623). ACM. https://doi.org/10.1145/3442188.3445922
Bird, S. (2020). Decolonising speech and language technology. In Proceedings of the 28th International Conference on Computational Linguistics (pp. 3504-3519). International Committee on Computational Linguistics. https://doi.org/10.18653/v1/2020.coling-main.313
Blasi, D. E., Anastasopoulos, A., & Neubig, G. (2022). Systematic inequalities in language technology performance across the world's languages. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 5486-5505). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.acl-long.376
Bojanowski, P., Grave, E., Joulin, A., & Mikolov, T. (2017). Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5, 135-146. https://doi.org/10.1162/tacl_a_00051
Boyd, D., & Crawford, K. (2012). Critical questions for Big Data: Provocations for a cultural, technological, and scholarly phenomenon. Information, Communication & Society, 15(5), 662-679. https://doi.org/10.1080/1369118X.2012.678878
Brown, P. F., Della Pietra, S. A., Della Pietra, V. J., & Mercer, R. L. (1993). The mathematics of statistical machine translation: Parameter estimation. Computational Linguistics, 19(2), 263-311. https://doi.org/10.1162/coli.1993.19.2.263
Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., & Bengio, Y. (2014). Learning phrase representations using RNN encoder-decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (pp. 1724-1734). Association for Computational Linguistics. https://doi.org/10.3115/v1/D14-1179
Costa-jussà, M. R., Cross, J., Celebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., Sun, A., Wang, S., Wenzek, G., Youngblood, A., Akula, B., Barrault, L., Gonzalez, G. M., Hansanti, P., Hoffman, J., ... Wang, J. (2022). No Language Left Behind: Scaling human-centered machine translation. Transactions of the Association for Computational Linguistics, 10, 1339-1360. https://doi.org/10.1162/tacl_a_00515
Currey, A., Miceli Barone, A. V., & Heafield, K. (2017). Copied monolingual data improves low-resource neural machine translation. In Proceedings of the Second Conference on Machine Translation (pp. 148-156). Association for Computational Linguistics. https://doi.org/10.18653/v1/W17-4715
Dabre, R., Chu, C., & Kunchukuttan, A. (2020). A survey of multilingual neural machine translation. ACM Computing Surveys, 53(5), Article 99. https://doi.org/10.1145/3406095
Dong, J., Jiang, T., Cheng, L., Anwar, A., & Yang, Y. (2018). A compromise Arabic-Kazakh coded character processing method based on the OpenType font format. Computer Standards & Interfaces, 55, 1-7. https://doi.org/10.1016/j.csi.2017.08.004
Edunov, S., Ott, M., Auli, M., & Grangier, D. (2018). Understanding back-translation at scale. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (pp. 489-500). Association for Computational Linguistics. https://doi.org/10.18653/v1/D18-1045
El-Kishky, A., Chaudhary, V., Guzmán, F., & Koehn, P. (2020). CCAligned: A massive collection of cross-lingual web-document pairs. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (pp. 5960-5969). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.480
Fan, A., Bhosale, S., Schwenk, H., Ma, Z., El-Kishky, A., Goyal, S., Baines, M., Celebi, O., Wenzek, G., Chaudhary, V., Goyal, N., Birch, T., Liptchinsky, V., Edunov, S., Grave, E., Auli, M., & Joulin, A. (2021). Beyond English-centric multilingual machine translation. Transactions of the Association for Computational Linguistics, 9, 483-503. https://doi.org/10.1162/tacl_a_00388
Firat, O., Cho, K., & Bengio, Y. (2016). Multi-way, multilingual neural machine translation with a shared attention mechanism. In Proceedings of NAACL-HLT 2016 (pp. 866-875). Association for Computational Linguistics. https://doi.org/10.18653/v1/N16-1101
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daume III, H., & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86-92. https://doi.org/10.1145/3458723
Geiger, R. S., Yu, K., Yang, Y., Dai, M., Qiu, J., Tang, R., & Huang, J. (2020). Garbage in, garbage out? Do machine learning application papers in social computing report where human-labeled training data comes from? In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (pp. 325-336). ACM. https://doi.org/10.1145/3351095.3372862
Goyal, N., Gao, C., Chaudhary, V., Chen, P.-J., Wenzek, G., Ju, D., Krishnan, S., Ranzato, M., Guzmán, F., & Fan, A. (2022). The Flores-101 evaluation benchmark for low-resource and multilingual machine translation. Transactions of the Association for Computational Linguistics, 10, 522-538. https://doi.org/10.1162/tacl_a_00474
Gu, J., Hassan, H., Devlin, J., & Li, V. O. K. (2018). Universal neural machine translation for extremely low resource languages. In Proceedings of NAACL-HLT 2018 (pp. 344-354). Association for Computational Linguistics. https://doi.org/10.18653/v1/N18-1032
Guzmán, F., Chen, P.-J., Ott, M., Pino, J., Lample, G., Koehn, P., Chaudhary, V., & Ranzato, M. (2019). The FLORES evaluation datasets for low-resource machine translation: Nepali-English and Sinhala-English. In Proceedings of EMNLP-IJCNLP 2019 (pp. 6098-6111). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1632
Holstein, K., Vaughan, J. W., Daume III, H., Dudik, M., & Wallach, H. (2019). Improving fairness in machine learning systems: What do industry practitioners need? In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Article 600). ACM. https://doi.org/10.1145/3290605.3300830
Hovy, D., & Spruit, S. L. (2016). The social impact of natural language processing. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) (pp. 591-598). Association for Computational Linguistics. https://doi.org/10.18653/v1/P16-2096
Hunter, J. D. (2007). Matplotlib: A 2D graphics environment. Computing in Science & Engineering, 9(3), 90-95. https://doi.org/10.1109/MCSE.2007.55
Jean, S., Cho, K., Memisevic, R., & Bengio, Y. (2015). On using very large target vocabulary for neural machine translation. In Proceedings of ACL-IJCNLP 2015 (pp. 1-10). Association for Computational Linguistics. https://doi.org/10.3115/v1/P15-1001
Jo, E. S., & Gebru, T. (2020). Lessons from archives: Strategies for collecting sociocultural data in machine learning. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (pp. 306-316). ACM. https://doi.org/10.1145/3351095.3372829
Johnson, M., Schuster, M., Le, Q. V., Krikun, M., Wu, Y., Chen, Z., Thorat, N., Viegas, F. B., Wattenberg, M., Corrado, G., Hughes, M., & Dean, J. (2017). Google's multilingual neural machine translation system: Enabling zero-shot translation. Transactions of the Association for Computational Linguistics, 5, 339-351. https://doi.org/10.1162/tacl_a_00065
Joshi, P., Santy, S., Budhiraja, A., Bali, K., & Choudhury, M. (2020). The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 6282-6293). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.560
Joulin, A., Grave, E., Bojanowski, P., & Mikolov, T. (2017). Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers) (pp. 427-431). Association for Computational Linguistics. https://doi.org/10.18653/v1/E17-2068
Junczys-Dowmunt, M. (2018). Dual conditional cross-entropy filtering of noisy parallel corpora. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers (pp. 888-895). Association for Computational Linguistics. https://doi.org/10.18653/v1/W18-64105
Junczys-Dowmunt, M., Grundkiewicz, R., Dwojak, T., Hoang, H., Heafield, K., Neckermann, T., Seide, F., Germann, U., Aji, A. F., Bogoychev, N., Martins, A. F. T., & Birch, A. (2018). Marian: Fast neural machine translation in C++. In Proceedings of ACL 2018: System Demonstrations (pp. 116-121). Association for Computational Linguistics. https://doi.org/10.18653/v1/P18-4020
Kitchin, R. (2014). Big Data, new epistemologies and paradigm shifts. Big Data & Society, 1(1), 1-12. https://doi.org/10.1177/2053951714528481
Knight, K., & Graehl, J. (1998). Machine transliteration. Computational Linguistics, 24(4), 599-612. https://doi.org/10.1162/089120198673219
Koehn, P., & Knowles, R. (2017). Six challenges for neural machine translation. In Proceedings of the First Workshop on Neural Machine Translation (pp. 28-39). Association for Computational Linguistics. https://doi.org/10.18653/v1/W17-3204
Koehn, P., Och, F. J., & Marcu, D. (2003). Statistical phrase-based translation. In Proceedings of the 2003 Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics (pp. 48-54). Association for Computational Linguistics. https://doi.org/10.3115/1075096.1075117
Kreutzer, J., Caswell, I., Wang, L., Wahab, A., van Esch, D., Ulzii-Orshikh, N., Tapo, A. A., Subramani, N., Sokolov, A., Sikasote, C., Setyawan, M., Sarin, S., Samb, S., Sagot, B., Rivera, C., Rios, A., Papadimitriou, I., Osei, S., Suarez, P. O., ... Adeyemi, M. (2022). Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics, 10, 50-72. https://doi.org/10.1162/tacl_a_00447
Kudo, T. (2018). SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of EMNLP 2018: System Demonstrations (pp. 66-71). Association for Computational Linguistics. https://doi.org/10.18653/v1/D18-2012
Lample, G., Conneau, A., Denoyer, L., & Ranzato, M. (2018). Phrase-based and neural unsupervised machine translation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (pp. 5039-5049). Association for Computational Linguistics. https://doi.org/10.18653/v1/D18-1549
Liu, Y., Gu, J., Goyal, N., Li, X., Edunov, S., Ghazvininejad, M., Lewis, M., & Zettlemoyer, L. (2020). Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8, 726-742. https://doi.org/10.1162/tacl_a_00343
Lu, Y. (2019). Artificial intelligence: A survey on evolution, models, applications and future trends. Journal of Management Analytics, 6(1), 1-29. https://doi.org/10.1080/23270012.2019.1570365
Lu, Y. (2021). Technological innovation and the emergence of a new interdisciplinary field: Management Analytics. Nanotechnologies in Construction, 13(3), 181-192. https://doi.org/10.15828/2075-8545-2021-13-3-181-192
Lu, Y. (2025). The current status and developing trends of Industry 4.0: A review. Information Systems Frontiers, 27(1), 215-234. https://doi.org/10.1007/s10796-021-10221-w
Lu, Y., & Xu, L. D. (2019). Internet of Things (IoT) cybersecurity research: A review of current research topics. IEEE Internet of Things Journal, 6(2), 2103-2115. https://doi.org/10.1109/JIOT.2018.2869847
Lui, M., & Baldwin, T. (2012). langid.py: An off-the-shelf language identification tool. In Proceedings of ACL 2012: System Demonstrations (pp. 25-30). Association for Computational Linguistics. https://doi.org/10.3115/v1/P12-3005
Luong, M.-T., Pham, H., & Manning, C. D. (2015). Effective approaches to attention-based neural machine translation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (pp. 1412-1421). Association for Computational Linguistics. https://doi.org/10.18653/v1/D15-1166
Manning, C. D., Surdeanu, M., Bauer, J., Finkel, J. R., Bethard, S., & McClosky, D. (2014). The Stanford CoreNLP natural language processing toolkit. In Proceedings of ACL 2014: System Demonstrations (pp. 55-60). Association for Computational Linguistics. https://doi.org/10.3115/v1/P14-5010
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 220-229). ACM. https://doi.org/10.1145/3287560.3287596
Och, F. J., & Ney, H. (2003). A systematic comparison of various statistical alignment models. Computational Linguistics, 29(1), 19-51. https://doi.org/10.1162/089120103321337421
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., & Auli, M. (2019). fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of NAACL-HLT 2019: Demonstrations (pp. 48-53). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-4009
Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (pp. 311-318). Association for Computational Linguistics. https://doi.org/10.3115/1073083.1073135
Paullada, A., Raji, I. D., Bender, E. M., Denton, E., & Hanna, A. (2021). Data and its (dis)contents: A survey of dataset development and use in machine learning research. Patterns, 2(11), 100336. https://doi.org/10.1016/j.patter.2021.100336
Popović, M. (2015). chrF: Character n-gram F-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation (pp. 392-395). Association for Computational Linguistics. https://doi.org/10.18653/v1/W15-3049
Post, M. (2018). A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers (pp. 186-191). Association for Computational Linguistics. https://doi.org/10.18653/v1/W18-6319
Ranathunga, S., Lee, E.-S. A., Prifti Skenduli, M., Shekhar, R., Alam, M., & Kaur, R. (2023). Neural machine translation for low-resource languages: A survey. ACM Computing Surveys, 55(11), Article 229. https://doi.org/10.1145/3567592
Rei, R., Stewart, C., Farinha, A. C., & Lavie, A. (2020). COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (pp. 2685-2702). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.213
Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of EMNLP-IJCNLP 2019 (pp. 3982-3992). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1410
Ruder, S., Constant, N., Botha, J., Siddhant, A., Firat, O., Fu, J., Liu, P., Hu, J., Garrette, D., Neubig, G., & Johnson, M. (2021). XTREME-R: Towards more challenging and nuanced multilingual evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 10215-10245). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.emnlp-main.802
Sambasivan, N., Kapania, S., Highfill, H., Akrong, D., Paritosh, P., & Aroyo, L. (2021). Everyone wants to do the model work, not the data work: Data cascades in high-stakes AI. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Article 39). ACM. https://doi.org/10.1145/3411764.3445518
Schuster, M., & Nakajima, K. (2012). Japanese and Korean voice search. In 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (pp. 5149-5152). IEEE. https://doi.org/10.1109/ICASSP.2012.6289079
Schwenk, H., Chaudhary, V., Sun, S., Gong, H., & Guzmán, F. (2021). WikiMatrix: Mining 135M parallel sentences in 1620 language pairs from Wikipedia. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics (pp. 1351-1361). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.eacl-main.115
Sellam, T., Das, D., & Parikh, A. P. (2020). BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 7881-7892). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.704
Sennrich, R., Haddow, B., & Birch, A. (2016a). Neural machine translation of rare words with subword units. In Proceedings of ACL 2016 (pp. 1715-1725). Association for Computational Linguistics. https://doi.org/10.18653/v1/P16-1162
Sennrich, R., Haddow, B., & Birch, A. (2016b). Improving neural machine translation models with monolingual data. In Proceedings of ACL 2016 (pp. 86-96). Association for Computational Linguistics. https://doi.org/10.18653/v1/P16-1009
Shmueli, G. (2010). To explain or to predict? Statistical Science, 25(3), 289-310. https://doi.org/10.1214/10-STS330
Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 3645-3650). Association for Computational Linguistics. https://doi.org/10.18653/v1/P19-1355
Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J. W., da Silva Santos, L. B., Bourne, P. E., Bouwman, J., Brookes, A. J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C. T., Finkers, R., ... Mons, B. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3, 160018. https://doi.org/10.1038/sdata.2016.18
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., ... Rush, A. M. (2020). Transformers: State-of-the-art natural language processing. In Proceedings of EMNLP 2020: System Demonstrations (pp. 38-45). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-demos.6
Xu, L. D., Lu, Y., & Li, L. (2021). Embedding blockchain technology into IoT for security: A survey. IEEE Internet of Things Journal, 8(13), 10452-10473. https://doi.org/10.1109/JIOT.2021.3060508
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., & Raffel, C. (2021). mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of NAACL-HLT 2021 (pp. 483-498). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.naacl-main.41
Zhang, C., & Lu, Y. (2021). Study on artificial intelligence: The state of the art and future prospects. Journal of Industrial Information Integration, 23, 100224. https://doi.org/10.1016/j.jii.2021.100224
Zoph, B., Yuret, D., May, J., & Knight, K. (2016). Transfer learning for low-resource neural machine translation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (pp. 1568-1575). Association for Computational Linguistics. https://doi.org/10.18653/v1/D16-1163