Representing emotive discourse in Ukrainian-English literary translation: A multi-method performance evaluation of Large Language Models, Neural Machine Translation and Computer-Assisted Translation tools

Authors

DOI:

https://doi.org/10.29038/kar

Keywords:

quality evaluation, neural machine translation, literary translation, BLEU, CAT tools, Large Language Model, emotive discourse

Abstract

The study examines the capacity of modern translation technologies to render Ukrainian literary texts into English. Lesya Ukrainka’s Letter to Serhii Merzhynskyi was chosen as the original text for translation analysis. It is a piece of emotive discourse marked by vivid imagery, nuanced stylistic features, expressive syntactic patterns and archaic vocabulary. Six translation services were tested. They included general-purpose Neural Machine Translation services, Computer-Assisted Translation tools and Large Language Models. Their output was evaluated using a three-step methodological framework. First, automatic evaluation was conducted using a Bilingual Evaluation Understudy (BLEU) metric to provide initial quantitative comparability across the systems’ output. Second, a qualitative analysis was undertaken through the concept of literariness, focusing on literature-specific features, aesthetic and stylistic peculiarities that distinguish literary texts from non-literary ones. In the final stage, human evaluation was employed, with five human annotators – native speakers with advanced linguistic proficiency, professional translators and scholars – ranking sentences to assess MT performance. The results of human evaluation and qualitative analysis revealed that the top-performing translation technologies were LLMs ChatGPT-5 and DeepSeek, which not only met a baseline level of translation adequacy but also consistently surpassed human translation in contextual and emotional sensitivity and overall naturalness and fluency. By contrast, automatic evaluation using the BLEU metric assigned the highest score to Google Translate output, highlighting the metric's limitations for literary text. Despite the notable efficiency of modern translation technologies, certain errors persist to varying degrees across all tested tools. These errors are connected with rendering imagery, handling syntactic constructions with long-range dependencies, translating pronouns, handling register mismatches, disrupting tone and other similar issues.

Acknowledgements

The authors express sincere gratitude to Tony Palmer, Russell Mackenzie, Nataliia Dobzhanska-Knight, and Nataliia Voloshynovych for their selfless assistance in conducting the human annotation work.

Disclosure Statement

As the authors of this paper and members of the Editorial Board of the EEJPL, Olena Karpina and Serhii Zasiekin declare that they have recused themselves from all editorial discussions and decisions concerning this manuscript.

Downloads

Download data is not yet available.

Author Biography

  • Olena Karpina *, Lesya Ukrainka Volyn National University, Ukraine

    Corresponding author. Email: [email protected]

References

Alghamdi, E. A., Zakraoui, J., & Abanmy, F. A. (2024). Domain adaptation for Arabic machine translation: Financial texts as a case study. Applied Sciences, 14(16), 7088. https://doi.org/10.3390/app14167088

Castilho, S., & Knowles, R. (2025). A survey of context in neural machine translation and its evaluation. Natural Language Processing, 31(4), 986–1016. Cambridge University Press. https://doi.org/10.1017/nlp.2024.7

Costa, A., Ling, W., Luís, T., Correia, R., & Coheur, L. (2017). A linguistically motivated taxonomy for machine translation error analysis. Machine Translation, 31(3–4), 229–244. https://doi.org/10.1007/s10590-017-9205-4

Federico, M., Cattelan, A., & Trombetti, M. (2014). The MateCat tool. In Proceedings of the COLING 2014: System Demonstrations (pp. 129–132). Association for Computational Linguistics. https://www.aclanthology.org/C14-2028.pdf

Fonteyne, M., Tezcan, A., & Macken, L. (2020). Literary machine translation under the magnifying glass: Assessing the quality of an NMT-translated detective novel on document level. In Proceedings of the Twelfth Language Resources and Evaluation Conference (LREC 2020) (pp. 3790–3798). European Language Resources Association. https://aclanthology.org/2020.lrec-1.468/

Freitag, M., Foster, G., Grangier, D., Ratnakar, V., Tan, Q., & Macherey, W. (2021). Experts, errors, and context: A large-scale study of human evaluation for machine translation. Transactions of the Association for Computational Linguistics, 9, 1460–1474. https://doi.org/10.1162/tacl_a_00437

Guerberof-Arenas, A., & Toral, A. (2022). Creativity in translation: Machine translation as a constraint for literary texts. Translation Spaces, 11(2), 184–212. https://doi.org/10.1075/ts.21025.gue.

Jacobs, A. M., & Kinder, A. (2022). Computational analyses of the topics, sentiments, literariness, creativity and beauty of texts in a large corpus of English literature. arXiv Preprint arXiv:2201.04356. https://arxiv.org/abs/2201.04356

Karpina, O. (2020). Komparatyvnyi analiz literaturnoho i mashynnoho perekladiv (na materiali frahmentiv romanu S. Plath “The Bell Jar”) [Comparative analysis of literary and machine translations (a case study of the excerpts The Bell Jar by S. Plath)]. Current Issues of Foreign Philology, 3, 94–101. https://doi.org/10.32782/2410-0927-2020-12-16

Karpina, O. (2023). Evaluating the quality of machine translation output with HTER in domain-specific textual environment. Linguistic Studies, 46, 85–99. https://doi.org/10.31558/1815-3070.2023.46.8

Karpinska, M., & Iyyer, M. (2023). Large language models effectively leverage document-level context for literary translation, but critical errors persist. arXiv Preprint arXiv:2304.03245. https://arxiv.org/abs/2304.03245

Kiritchenko, S., & Mohammad, S. M. (2017). Best–worst scaling more reliable than rating scales: A case study on sentiment intensity annotation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) (pp. 465–470). Association for Computational Linguistics. https://doi.org/10.18653/v1/P17-2074

Kocmi, T., & Federmann, C. (2023). GEMBA-MQM: Detecting translation quality error spans with GPT-4. In Proceedings of the Eighth Conference on Machine Translation (pp. 768–775). Association for Computational Linguistics. https://arxiv.org/abs/2310.13988

Koliada, I. (2021, February 25–26). “Ukrainska Biatrice” i “Donka Prometeia”: Zhertovne kokhannia u zhytti Hanny Barvinok i Lesi Ukrainky [“Ukrainian Beatrice” and “Daughter of Prometheus”: Sacrificial love in the life of Hanna Barvinok and Lesya Ukrainka]. In Ideolohynia natsionalnoi arystokratii (na poshanu 150-richchia vid dnia narodzhennia Lesi Ukrainky) [The Ideologist of the National Aristocracy (in honor of the 150th anniversary of Lesya Ukrainka’s birth)], Proceedings of the International Conference (pp. 326–335). Lviv Danylo Halytskyi National Medical University.

Lan, M., & Zhao, L. (2021). Contrasting and analyzing machine and human translation: A case study on Red Sorghum. In Proceedings of the 2021 6th International Conference on Modern Management and Education Technology (MMET 2021) (pp. 684–690). Atlantis Press. https://doi.org/10.2991/assehr.k.211011.123

Lihus, O., & Grinchenko, B. (2021). Ukrainian Romanticism in the context of European culture of the 19th – early 20th centuries: Interdisciplinary historiographical analysis. Mundo Eslavo, 20, 147–157. https://revistaseug.ugr.es/index.php/meslav/article/view/21627/22597/79552

Macken, L. (2024). Evaluating ChatGPT’s Ability to Automatically Post-Edit Literary Texts. In B. Vanroy, M.-A. Lefer, L. Macken, & P. Ruffo (Eds.), Proceedings of the 1st Workshop on Creative-text Translation and Technology (pp. 65–81). European Association for Machine Translation. https://aclanthology.org/2024.ctt-1.7/

Mienye, I. D., Swart, T. G., & Obaido, G. (2024). Recurrent neural networks: A comprehensive review of architectures, variants, and applications. Information, 15(9), 517. https://doi.org/10.3390/info15090517

Noll, R., Berger, A., Kieu, D., et al. (2025). Assessing GPT and DeepL for terminology translation in the medical domain: A comparative study on the human phenotype ontology. BMC Medical Informatics and Decision Making, 25, 237. https://doi.org/10.1186/s12911-025-03075-8

Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). BLEU: A method for automatic evaluation of machine translation. In P. Isabelle, E. Charniak, & D. Lin (Eds.), Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (pp. 311–318). Association for Computational Linguistics. https://doi.org/10.3115/1073083.1073135

Rybicki, J. (2025). Can machine translation of literary texts fool stylometry? Digital Scholarship in the Humanities, 40(1), 268–276. https://doi.org/10.1093/llc/fqaf010

Sun, S., Liu, K., & Moratto, R. (2025). Navigating the paradigm shift-translation studies in the age of AI. In S. Sun, K. Liu, & R. Moratto (Eds.), Translation Studies in the Age of Artificial Intelligence (pp. 1–17). Routledge.

Thai, K., Karpinska, M., Krishna, K., Ray, W., Inghilleri, M., Wieting, J., & Iyyer, M. (2022). Exploring document-level literary machine translation with parallel paragraphs from world literature. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (pp. 9882–9902). Association for Computational Linguistics. https://aclanthology.org/2022.emnlp-main.672

Toral, A., & Way, A. (2014). Is machine translation ready for literature? In Translating and the Computer 36 (pp. 174–176). Dublin City University. https://aclanthology.org/2014.tc-1.23.pdf

Toral, A., & Way, A. (2018). What level of quality can neural machine translation attain on literary text? arXiv Preprint arXiv:1801.04962. https://doi.org/10.48550/arXiv.1801.04962

Toral, A., van Cranenburgh, A. V., & Nutters, T. (2024). Literary-adapted machine translation in a well-resourced language pair: Explorations with More Data and Wider Contexts. In A. Rothwell, A. Way, & R. Youdale (Eds.), Computer-assisted literary translation (pp. 27–52). Routledge. https://doi.org/10.4324/9781003357391-3

Van Cranenburgh, A., & Bod, R. (2017). A Data-Oriented Model of Literary Language. In M. Lapata, P. Blunsom, & A. Koller (Eds.), Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (Long Papers) (pp. 1228–1238). Association for Computational Linguistics. https://aclanthology.org/E17-1115

Vennita, R. & Hasnah, Y. (2024). A probe into the comparison of human translation and deep translation in translating English text into Indonesian. Journal of English Development, 4(2), 303-330. https://doi.org/10.25217/jed.v3i01.4503

Wu, M., Xu, J., Yuan, Y., Haffari, G., Wang, L., Luo, W., & Zhang, K. (2024). (Perhaps) beyond human translation: Harnessing multi-agent collaboration for translating ultra-long literary texts. arXiv Preprint arXiv:2405.11804. https://arxiv.org/abs/2405.11804

Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., ... & Dean, J. (2016). Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv Preprint arXiv:1609.08144. https://doi.org/10.48550/arXiv.1609.08144

Xu, Y., Zhang, B., et al. (2024). DeepSeek LLM: Scaling open-source language models with longtermism. arXiv Preprint arXiv:2401.02954. https://arxiv.org/abs/2401.02954

Yulianto, A., & Supriatnaningsih, R. (2021). Google Translate vs. DeepL: A quantitative evaluation of close-language pair translation (French to English). AJELP: Asian Journal of English Language and Pedagogy, 9(2), 109–127. https://doi.org/10.37134/ajelp.vol9.2.9.2021

Zasiekin, S. (2019). Investigating cognitive and psycholinguistic features of translation universals. Psycholinguistics, 26(2), 114–134. https://doi.org/10.31470/2309-1797-2019-26-2-114-134

Zasiekin, S. & Kalishchuk , D. (2025). Can machines communicate psychotrauma? Affective and cognitive shifts in AI-translated Russia-Ukraine war narratives. Psycholinguistics, 38(2), 58-76. https://doi.org/10.31470/2309-1797-2025-38-2-58-76

Zhang, R., Zhao, W., & Eger, S. (2025). How good are LLMs for literary translation, really? Literary translation evaluation with humans and LLMs. arXiv Preprint arXiv:2410.18697. https://doi.org/10.48550/arXiv.2410.18697

Zhang, R., Zhao, W., Macken, L., & Eger, S. (2025, May 8). TransProQA: An LLM-based literary translation evaluation metric with professional question answering. arXiv Preprint arXiv:2505.05423. https://doi.org/10.48550/arXiv.2505.05423

Sources

Collins Dictionary. (n.d.). Withered. In Collins English Dictionary. https://www.collinsdictionary.com/dictionary/english/withered

DeepL. (2024, October 9). DeepL is 2024’s most-used machine translation provider worldwide among language service companies [Press release]. PR Newswire. https://www.prnewswire.com/

DeepSeek. (2025, June 27). https://chat.deepseek.com/

Komarnyckyj, S. (2022). Lesya Ukrainka’s “Your Letters Always Smell of Withered Roses” (1947): A new translation. Volupté: Interdisciplinary Journal of Decadence Studies, 5(1), 92–94. https://doi.org/10.25602/GOLD.v.v5i1.1624.g1738

Matecat. (n.d.). Free online CAT tool with integrated MT. https://www.matecat.com

OpenAI. (2025, August 7). Introducing GPT-5. https://openai.com/index/introducing-gpt-5

Oxford Learner’s Dictionaries. (n.d.). Withered. In Oxford Learner’s Dictionaries. https://www.oxfordlearnersdictionaries.com/definition/english/withered?q=withered

Smartcat. (n.d.). AI-powered CAT tool. Free online translator. https://www.smartcat.com/cat-tool/

Interactive BLEU score evaluator (n. d.). Tilde. https://translate.tilde.ai/bleu#/

Українка Л. (2021) Повне академічне зібрання творів: у 14 томах. Том 12. Листи. 1897–1901 / ред. О. Полюхович; упоряд. В. Прокіп (Савчук); комент. В. Прокіп (Савчук), В. Агеєва. Луцьк: Волинський національний університет імені Лесі Українки,. 608 с., С. 321–322.

Ukrainka, L. (2021). Povne akademichne zibrannia tvoriv: u 14 tomakh. Tom 12. Lysty, 1897–1901 [Complete academic collection of works: in 14 volumes. Vol. 12. Letters, 1897–1901] / O. Poliukhovych (Ed.), compiled by V. Prokip (Savchuk) with comments from V. Prokip (Savchuk), V. Aheieva. (pp. 321–322). Lesya Ukrainka Volyn National University.

Downloads

Published

2025-12-29

Issue

Section

Vol. 12 No. 2 (2025)

How to Cite

Karpina, O., & Zasiekin, S. (2025). Representing emotive discourse in Ukrainian-English literary translation: A multi-method performance evaluation of Large Language Models, Neural Machine Translation and Computer-Assisted Translation tools. East European Journal of Psycholinguistics , 12(2), 178-203. https://doi.org/10.29038/kar

Similar Articles

1-10 of 339

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)

1 2 3 4 5 > >>