<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE root>
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" article-type="research-article" dtd-version="1.2" xml:lang="en"><front><journal-meta><journal-id journal-id-type="publisher-id">RUDN Journal of Language Studies, Semiotics and Semantics</journal-id><journal-title-group><journal-title xml:lang="en">RUDN Journal of Language Studies, Semiotics and Semantics</journal-title><trans-title-group xml:lang="ru"><trans-title>Вестник Российского университета дружбы народов. Серия: Теория языка. Семиотика. Семантика</trans-title></trans-title-group></journal-title-group><issn publication-format="print">2313-2299</issn><issn publication-format="electronic">2411-1236</issn><publisher><publisher-name xml:lang="en">Peoples’ Friendship University of Russia named after Patrice Lumumba (RUDN University)</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="publisher-id">52660</article-id><article-id pub-id-type="doi">10.22363/2313-2299-2026-17-2-557-568</article-id><article-id pub-id-type="edn">LSTNOS</article-id><article-categories><subj-group subj-group-type="toc-heading" xml:lang="en"><subject>ARTIFICIAL INTELLIGENCE IN APPLIED LINGUISTICS: CURRENT CHALLENGES AND PROSPECTS</subject></subj-group><subj-group subj-group-type="toc-heading" xml:lang="ru"><subject>АКТУАЛЬНЫЕ ВЫЗОВЫ ИСКУССТВЕННОГО ИНТЕЛЛЕКТА В ПРИКЛАДНОЙ ЛИНГВИСТИКЕ</subject></subj-group><subj-group subj-group-type="article-type"><subject>Research Article</subject></subj-group></article-categories><title-group><article-title xml:lang="en">Token as an Algorithmic ‘Shadow’ of a Morpheme</article-title><trans-title-group xml:lang="ru"><trans-title>Токен как алгоритмическая «тень» морфемы</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-2488-4616</contrib-id><contrib-id contrib-id-type="scopus">59236754500</contrib-id><contrib-id contrib-id-type="spin">5871-9336</contrib-id><name-alternatives><name xml:lang="en"><surname>Zhikulina</surname><given-names>Kristina P.</given-names></name><name xml:lang="ru"><surname>Жикулина</surname><given-names>Кристина Петровна</given-names></name></name-alternatives><bio xml:lang="en"><p>PhD in Philology, Аssistant of the Department of General and Russian Linguistics, Faculty of Philology</p></bio><bio xml:lang="ru"><p>кандидат филологических наук, ассистент кафедры общего и русского языкознания филологического факультета</p></bio><email>zhikulina-kp@rudn.ru</email><xref ref-type="aff" rid="aff1"/></contrib></contrib-group><aff-alternatives id="aff1"><aff><institution xml:lang="en">RUDN University</institution></aff><aff><institution xml:lang="ru">Российский университет дружбы народов</institution></aff></aff-alternatives><pub-date date-type="pub" iso-8601-date="2026-09-30" publication-format="electronic"><day>30</day><month>09</month><year>2026</year></pub-date><volume>17</volume><issue>2</issue><issue-title xml:lang="en">ARTIFICIAL INTELLIGENCE IN APPLIED LINGUISTICS: CURRENT CHALLENGES AND PROSPECTS</issue-title><issue-title xml:lang="ru">АКТУАЛЬНЫЕ ВЫЗОВЫ ИСКУССТВЕННОГО ИНТЕЛЛЕКТА В ПРИКЛАДНОЙ ЛИНГВИСТИКЕ</issue-title><fpage>557</fpage><lpage>568</lpage><history><date date-type="received" iso-8601-date="2026-10-06"><day>06</day><month>10</month><year>2026</year></date></history><permissions><copyright-statement xml:lang="en">Copyright ©; 2026, Zhikulina K.P.</copyright-statement><copyright-statement xml:lang="ru">Copyright ©; 2026, Жикулина К.П.</copyright-statement><copyright-year>2026</copyright-year><copyright-holder xml:lang="en">Zhikulina K.P.</copyright-holder><copyright-holder xml:lang="ru">Жикулина К.П.</copyright-holder><ali:free_to_read xmlns:ali="http://www.niso.org/schemas/ali/1.0/"/><license><ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/">https://creativecommons.org/licenses/by-nc/4.0</ali:license_ref></license></permissions><self-uri xlink:href="https://journals.rudn.ru/semiotics-semantics/article/view/52660">https://journals.rudn.ru/semiotics-semantics/article/view/52660</self-uri><abstract xml:lang="en"><p>The relevance of the study of a critical stage of text processing in the era of the dominance of large language models is due to the systemic divergence between statistically identified token boundaries and linguistically justified morpheme boundaries in tokenizers, which is particularly evident when working with rich morphology languages. The scientific novelty of the study lies in considering the token not as technical part, but as a linguistic artifact - an algorithmic ‘shadow’ of a morpheme that has lost its semantic motivation but preserved its external resemblance. The aim of the research is to identify linguistic losses that occur during statistical segmentation of words and to assess the degree of divergence between token and morpheme boundaries. Research methods: comparative linguistic analysis, corpus analysis, critical discourse analysis and typological modeling. During the research, a typology of morphological distortions in BPE tokenization has been developed and empirically substantiated, describing three types of distortions and divergence between token and morpheme boundaries: fragmentation, contamination, and aberration. Using the lexemes ‘podorozhnik’ and ‘vysokotekhnologichny’ as material, the three-stage degradation of morphological transparency in the transition from monolingual to bilingual and multilingual BPE models has been experimentally proven. It has been established that multilingual tokenizers completely lose morphological properties in relation to a particular language multiplying statistical artifacts that are not identified by a classical set of morphological units and native speakers of this language.</p></abstract><trans-abstract xml:lang="ru"><p>Актуальность исследования критического этапа обработки текста в эпоху доминирования больших языковых моделей обусловлена системным расхождением между статистически выделенными границами токенов и лингвистически обоснованными границами морфем в токенизаторах, что проявляется в работе с языками с богатой морфологией. Научная новизна исследования заключается в рассмотрении токена не как технической единицы, а как лингвистического артефакта - алгоритмической «тени» морфемы , утратившей семантическую мотивированность, но сохранившую внешнее сходство. Цель исследования - выявить лингвистические потери, возникающие при статистической сегментации слов и оценить степень расхождения между токенными и морфемными границами. Методы исследования: сопоставительный лингвистический анализ, корпусный анализ, критический дискурс-анализ и типологическое моделирование. В ходе исследований разработана и эмпирически обоснована типология морфологических искажений при BPE-токенизации, описывающая три типа искажений и расхождений между морфемными и токенными границами: фрагментация, контаминация и аберрация. На материале лексем подорожник и высокотехнологичный экспериментально доказана трехступенчатая деградация морфологической прозрачности при переходе от одноязычных к двуязычным и мультиязычным BPE-моделzv. Установлено, что мультиязычные токенизаторы полностью утрачивают морфологические свойства в отношении отдельного языка, множа статистические артефакты, которые не идентифицируются классическим набором единиц морфологии и носителями данного языка.</p></trans-abstract><kwd-group xml:lang="en"><kwd>artificial intelligence</kwd><kwd>computational linguistics</kwd><kwd>machine learning</kwd><kwd>BPE-tokenization</kwd><kwd>NLP-systems</kwd><kwd>LLM</kwd><kwd>subword segmentation</kwd><kwd>morphology</kwd><kwd>morphemization</kwd></kwd-group><kwd-group xml:lang="ru"><kwd>искусственный интеллект</kwd><kwd>компьютерная лингвистика</kwd><kwd>машинное обучение</kwd><kwd>BPE-токенизация</kwd><kwd>НЛП-системы</kwd><kwd>большая языковая модель</kwd><kwd>сублексемная сегментация</kwd><kwd>морфология</kwd><kwd>морфемизация</kwd><kwd>подслово</kwd></kwd-group><funding-group><award-group><funding-source><institution-wrap><institution xml:lang="ru">Публикация подготовлена в рамках грантовой поддержки научных  проектов РУДН 050227-0-000</institution></institution-wrap><institution-wrap><institution xml:lang="en">The publication has been supported by the RUDN University Scientific Projects Grant System, project No 050227-0-000</institution></institution-wrap></funding-source></award-group></funding-group></article-meta><fn-group/></front><body></body><back><ref-list><ref id="B1"><label>1.</label><citation-alternatives><mixed-citation xml:lang="en">Sennrich, R., Haddov, B., &amp; Brich, A. (2016). Neural Machine Translation of Rare Words with Subword Units. In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (pp. 1715–1725). N-Y.: Cornell University publ. https://doi.org/10.48550/erXiv.1508.07909</mixed-citation><mixed-citation xml:lang="ru">Sennrich R., Haddov B., Brich A. Neural Machine Translation of Rare Words with Subword Units // Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics. N.-Y. : Cornell University, 2016 P. 1715-1725. https://doi.org/10.48550/erXiv.1508.07909</mixed-citation></citation-alternatives></ref><ref id="B2"><label>2.</label><citation-alternatives><mixed-citation xml:lang="en">Maksimenko, O.I. (2024). Formalized and Digital Linguistics. Moscow: State University of Education publ. (In Russ.). EDN: NPXJZB</mixed-citation><mixed-citation xml:lang="ru">Максименко О.И. Формализованная и цифровая лингвистика. М. : Государственный университет просвещения, 2024. EDN: NPXJZB</mixed-citation></citation-alternatives></ref><ref id="B3"><label>3.</label><citation-alternatives><mixed-citation xml:lang="en">Nikitina, E.N., Stankevich, M.A., &amp; Larionov, D.S. (2025). Psych Verbs Used with Negation: A Way of Interpretation of Automatic Text Analysis Data. Media Linguistics, 12(2), 174–198. (In Russ.). https://doi.org/10.21638/spbu22.2025.201 EDN: TDQIWC</mixed-citation><mixed-citation xml:lang="ru">Никитина Е.Н., Станкевич М.А., Ларионов Д.С. Эмотивные глаголы с отрицанием: опыт интерпретации данных автоматического анализа текстов // Медиалингвистика. 2025. Т. 12. № 2. С. 174-198. https://doi.org/10.21638/spbu22.2025.201 EDN: TDQIWC</mixed-citation></citation-alternatives></ref><ref id="B4"><label>4.</label><citation-alternatives><mixed-citation xml:lang="en">Gudkov, V.V., Mitrenina, O.V., Sokolov, E.G., &amp; Koval, A.A. (2023). Language-Based Transfer Learning Approaches for Part-of-Speech Tagging on Saint Petersburg Corpus of Hagiographic Texts (Skat). Vestnik of Saint Petersburg University. Language and Literature, 20(2), 268—282. (In Russ.). https://doi.org/10.21638/spbu09.2023.205 EDN: VQHOZM</mixed-citation><mixed-citation xml:lang="ru">Гудков В.В., Митренина О.В., Соколов Е.Г., Коваль А.А. Языковой перенос нейросетевого обучения для частеречной разметки Санкт-Петербургского корпуса агиографических текстов (СКАТ) // Вестник Санкт-Петербургского университета. Язык и литература. 2023. Т. 20. № 2. C. 268-282. https://doi.org/10.21638/spbu09.2023.205 EDN: VQHOZM</mixed-citation></citation-alternatives></ref><ref id="B5"><label>5.</label><citation-alternatives><mixed-citation xml:lang="en">Polshchykova, O.N. (2022). Formation of Computational Linguistics Terminology. Issues in Journalism, Education, Linguistics, 41(3), 590–607. (In Russ.). https://doi.org/10.52575/2712-7451-2022-41-3-590-607 EDN: HTWXDI</mixed-citation><mixed-citation xml:lang="ru">Польщикова О.Н. Становление и формирование терминологии компьютерной лингвистики // Вопросы журналистики, педагогики, языкознания. 2022. Т. 41. № 3. C. 590-607. https://doi.org/10.52575/2712-7451-2022-41-3-590-607 EDN: HTWXDI</mixed-citation></citation-alternatives></ref><ref id="B6"><label>6.</label><citation-alternatives><mixed-citation xml:lang="en">Gorozhanov, A.I., &amp; Stepanova, D.V. (2022). Work of Fiction Interpretation:  Corpus Approach. Philology. Theory &amp; Practice, 15(1), 203–208. (In Russ.). https://doi.org/10.30853/phil20220020 EDN: TCZLAF</mixed-citation><mixed-citation xml:lang="ru">Горожанов А.И., Степанова Д.В. Интерпретация художественного произведения: корпусный подход // Филологические науки. Вопросы теории и практики. 2022. № 15 (1). C. 203-208. https://doi.org/10.30853/phil20220020 EDN: TCZLAF</mixed-citation></citation-alternatives></ref><ref id="B7"><label>7.</label><citation-alternatives><mixed-citation xml:lang="en">Kudo, T. (2018). Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates // arXiv.org. URL: https://arxiv.org/abs/1804.10959 (accessed: 18.12.2025).</mixed-citation><mixed-citation xml:lang="ru">Kudo T. Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates // arXiv.org. 2018. Режим доступа: https://arxiv.org/abs/1804.10959 (дата обращения: 18.12.2025).</mixed-citation></citation-alternatives></ref><ref id="B8"><label>8.</label><citation-alternatives><mixed-citation xml:lang="en">Kudo, T., &amp; Richardson, J. (2018). SentencePiece: A Simple and Language Independent Subword Tokenizer and Detokenizer for Neural Text Processing. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (pp. 66–71). Brussels. https://doi.org/10.48550/arXiv.1808.06226</mixed-citation><mixed-citation xml:lang="ru">Kudo T., Richardson J. SentencePiece: A Simple and Language Independent Subword Tokenizer and Detokenizer for Neural Text Processing // Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Brussels. P. 66-71. 2018 https://doi.org/10.48550/arXiv.1808.06226</mixed-citation></citation-alternatives></ref><ref id="B9"><label>9.</label><citation-alternatives><mixed-citation xml:lang="en">Lomakina, O.V. (2025). Metalanguage of Modern Linguaxiology: Paremiological Level. RUDN Journal of Language Studies. Semiotics and Semantics, 16(3), 622–637. (In Russ.). https://doi.org/10.22363/2313-2299-2025-16-3-622-637 EDN: DQGURR</mixed-citation><mixed-citation xml:lang="ru">Ломакина О.В. Метаязык современной лингвокультурологии: паремиологический уровень // Вестник Российского университета дружбы народов. Серия: Теория языка. Семиотика. Семантика. 2025. Т. 16. № 3. С. 622-637. https://doi.org/10.22363/2313-2299-2025-16-3-622-637 EDN: DQGURR</mixed-citation></citation-alternatives></ref><ref id="B10"><label>10.</label><citation-alternatives><mixed-citation xml:lang="en">Zonova, D.Yu. (2023). The Principle of Operation and Problems of “Generative Pre-trained Transformer Artificial Intelligence”. Bulletin of Science and Education, 8 (139), 28–31.  (In Russ.). EDN: AFNBRE</mixed-citation><mixed-citation xml:lang="ru">Зонова Д.Ю. Принцип работы и проблемы «generative pre-trained transformer artificial intelligence» // Вестник науки и образования. 2023. № 8 (139). С. 28-31. EDN: AFNBRE</mixed-citation></citation-alternatives></ref></ref-list></back></article>
