<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE root>
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" article-type="research-article" dtd-version="1.2" xml:lang="en"><front><journal-meta><journal-id journal-id-type="publisher-id">RUDN Journal of Language Studies, Semiotics and Semantics</journal-id><journal-title-group><journal-title xml:lang="en">RUDN Journal of Language Studies, Semiotics and Semantics</journal-title><trans-title-group xml:lang="ru"><trans-title>Вестник Российского университета дружбы народов. Серия: Теория языка. Семиотика. Семантика</trans-title></trans-title-group></journal-title-group><issn publication-format="print">2313-2299</issn><issn publication-format="electronic">2411-1236</issn><publisher><publisher-name xml:lang="en">Peoples’ Friendship University of Russia named after Patrice Lumumba (RUDN University)</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="publisher-id">52659</article-id><article-id pub-id-type="doi">10.22363/2313-2299-2026-17-2-534-456</article-id><article-id pub-id-type="edn">LQDNOM</article-id><article-categories><subj-group subj-group-type="toc-heading" xml:lang="en"><subject>ARTIFICIAL INTELLIGENCE IN APPLIED LINGUISTICS: CURRENT CHALLENGES AND PROSPECTS</subject></subj-group><subj-group subj-group-type="toc-heading" xml:lang="ru"><subject>АКТУАЛЬНЫЕ ВЫЗОВЫ ИСКУССТВЕННОГО ИНТЕЛЛЕКТА В ПРИКЛАДНОЙ ЛИНГВИСТИКЕ</subject></subj-group><subj-group subj-group-type="article-type"><subject>Research Article</subject></subj-group></article-categories><title-group><article-title xml:lang="en">Modeling Demographically Conditioned Language in Large Language Models: Towards a Synthetic Respondent for Linguistic Research</article-title><trans-title-group xml:lang="ru"><trans-title>Моделирование демографически обусловленной речи в больших языковых моделях: построение синтетического респондента для лингвистических исследований</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-1748-7666</contrib-id><contrib-id contrib-id-type="spin">6174-8813</contrib-id><name-alternatives><name xml:lang="en"><surname>Vanichkina</surname><given-names>Alexandra S.</given-names></name><name xml:lang="ru"><surname>Ваничкина</surname><given-names>Александра Савельевна</given-names></name></name-alternatives><bio xml:lang="en"><p>PhD in director, Institute of Information Sciences</p></bio><bio xml:lang="ru"><p>кандидат филологических наук, директор, Институт информационных наук</p></bio><email>alexvanichkina@gmail.com</email><xref ref-type="aff" rid="aff1"/></contrib><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0009-0006-9151-0269</contrib-id><name-alternatives><name xml:lang="en"><surname>Izyumskaya-Kapitonova</surname><given-names>Veronika V.</given-names></name><name xml:lang="ru"><surname>Изюмская-Капитонова</surname><given-names>Вероника Викторовна</given-names></name></name-alternatives><bio xml:lang="en"><p>Junior Researcher, Corpus Linguistics Laboratory</p></bio><bio xml:lang="ru"><p>младший научный сотрудник, лаборатория корпусной лингвистики</p></bio><email>v.v.linguist.ik@gmail.com</email><xref ref-type="aff" rid="aff1"/></contrib></contrib-group><aff-alternatives id="aff1"><aff><institution xml:lang="en">Moscow State Linguistic University</institution></aff><aff><institution xml:lang="ru">Московский государственный лингвистический университет</institution></aff></aff-alternatives><pub-date date-type="pub" iso-8601-date="2026-09-30" publication-format="electronic"><day>30</day><month>09</month><year>2026</year></pub-date><volume>17</volume><issue>2</issue><issue-title xml:lang="en">ARTIFICIAL INTELLIGENCE IN APPLIED LINGUISTICS: CURRENT CHALLENGES AND PROSPECTS</issue-title><issue-title xml:lang="ru">АКТУАЛЬНЫЕ ВЫЗОВЫ ИСКУССТВЕННОГО ИНТЕЛЛЕКТА В ПРИКЛАДНОЙ ЛИНГВИСТИКЕ</issue-title><fpage>534</fpage><lpage>456</lpage><history><date date-type="received" iso-8601-date="2026-10-06"><day>06</day><month>10</month><year>2026</year></date></history><permissions><copyright-statement xml:lang="en">Copyright ©; 2026, Vanichkina A.S., Izyumskaya-Kapitonova V.V.</copyright-statement><copyright-statement xml:lang="ru">Copyright ©; 2026, Ваничкина А.С., Изюмская-Капитонова В.В.</copyright-statement><copyright-year>2026</copyright-year><copyright-holder xml:lang="en">Vanichkina A.S., Izyumskaya-Kapitonova V.V.</copyright-holder><copyright-holder xml:lang="ru">Ваничкина А.С., Изюмская-Капитонова В.В.</copyright-holder><ali:free_to_read xmlns:ali="http://www.niso.org/schemas/ali/1.0/"/><license><ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/">https://creativecommons.org/licenses/by-nc/4.0</ali:license_ref></license></permissions><self-uri xlink:href="https://journals.rudn.ru/semiotics-semantics/article/view/52659">https://journals.rudn.ru/semiotics-semantics/article/view/52659</self-uri><abstract xml:lang="en"><p>Using large language models (LLMs) to construct synthetic respondents is increasingly common in marketing and the social sciences, yet their applicability to modeling demographically conditioned variation in natural language remains underexplored in linguistics. This article aims to assess the extent to which synthetic respondents, instantiated via sociodemographic persona prompting, can reproduce gender-marked features of Russian diary writing. The data comprise diary entries from the Russian National Corpus (RNC) after 1950, split into female and male subcorpora and used as reference data, as well as 108 synthetic diary texts generated by three large language models (GPT-5.2, DeepSeek, GigaChat) across 32 persona-prompt configurations. The methods include corpus-based analysis of gender variation in language using a set of lexico-grammatical and pragmatic features, automated morphological and lexico-pragmatic annotation of synthetic texts with the Stanza library, computation of a parametric and a lexical index of gender proximity, and their aggregation into an integral measure. The results show that specific combinations of models and persona-prompt formulas (above all GPT-5.2 and DeepSeek with direct role adoption and structured or name-based demographic priming) generate texts whose statistical profiles approximate the female diary subcorpus of the RNC, whereas other configurations systematically shift towards the male profile. The study concludes that synthetic respondents can serve as a research instrument for modeling gender-marked language use, but their validity is tightly dependent on the choice of model and persona prompt and requires prior calibration against reference corpora, thereby specifying the scope and conditions for the responsible use of synthetic respondents in linguistics.</p></abstract><trans-abstract xml:lang="ru"><p>Использование больших языковых моделей (LLM) для построения синтетических респондентов активно осваивается в маркетинге и социальных науках, однако их применимость к моделированию демографически обусловленной вариативности естественной речи остается в лингвистике недостаточно исследованной. Целью статьи является оценка того, в какой мере синтетические респонденты, задаваемые с помощью личностного социодемографического промптинга, способны воспроизводить гендерно маркированные параметры русскоязычного дневникового письма. Материалом служат дневниковые записи Национального корпуса русского языка (НКРЯ) после 1950 г., разделенные на женский и мужской подкорпусы и использованные как референсные данные, а также 108 синтетических дневниковых текстов, сгенерированных тремя большими языковыми моделями (GPT-5.2, DeepSeek, GigaChat) по 32 конфигурациям личностного промптинга. Методы включают корпусный анализ гендерной вариативности речи на основе набора лексико-грамматических и прагматических показателей, автоматизированную морфологическую и лексико-прагматическую разметку синтетических текстов с помощью библиотеки Stanza, расчет параметрического и лексемного индексов гендерной близости и их агрегирование в интегральный показатель. Результаты показывают, что отдельные комбинации моделей и формул личностного промптинга (прежде всего GPT-5.2 и DeepSeek при непосредственном принятии роли и структурированной или именной стратегии прайминга) генерируют тексты, статистический профиль которых сближается с женским дневниковым подкорпусом НКРЯ, тогда как другие конфигурации систематически смещаются к мужскому профилю. Заключение. Показано, что синтетические респонденты могут выступать исследовательским инструментом для моделирования гендерно маркированного речевого поведения, но их валидность строго зависит от выбора модели и личностного промпта и требует предварительной калибровки относительно референсных корпусов; тем самым статья уточняет границы и условия ответственного использования синтетических респондентов в лингвистике.</p></trans-abstract><kwd-group xml:lang="en"><kwd>artificial intelligence</kwd><kwd>semiotics</kwd><kwd>semantics</kwd><kwd>simulation of meaning</kwd><kwd>perplexity</kwd></kwd-group><kwd-group xml:lang="ru"><kwd>cинтетические респонденты</kwd><kwd>большие языковые модели</kwd><kwd>личностный социодемографический промптинг</kwd><kwd>гендерная вариативность речи</kwd><kwd>корпусная лингвистика</kwd><kwd>генерация синтетических текстов</kwd></kwd-group><funding-group/></article-meta><fn-group/></front><body></body><back><ref-list><ref id="B1"><label>1.</label><citation-alternatives><mixed-citation xml:lang="en">Shrestha, P., Krpan, D., Koaik, F., Schnider, R., Sayess, D., &amp; Binbaz, M.S. (2024). Beyond WEIRD: Can Synthetic Survey Participants Substitute for Humans in Global Policy Research? Behavioral Science &amp; Policy, 10(2), 26–5. https://doi.org/10.1177/23794607241311793 EDN: RJHEHB</mixed-citation><mixed-citation xml:lang="ru">Shrestha P., Krpan D., Koaik F., Schnider R., Sayess D., Binbaz M.S. Beyond WEIRD: Can synthetic survey participants substitute for humans in global policy research? // Behavioral Science &amp; Policy. 2024. Vol. 10. Iss. 2. P. 26-45. https://doi.org/10.1177/23794607241311793 EDN: RJHEHB</mixed-citation></citation-alternatives></ref><ref id="B2"><label>2.</label><citation-alternatives><mixed-citation xml:lang="en">Argyle, L.P., Busby, E.C., Fulda, N., Gubler, J.R., Rytting, C., &amp; Wingate, D. (2023). Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis, 31(3), 337–351. https://doi.org/10.1017/pan.2023.2 EDN: TOAPUN</mixed-citation><mixed-citation xml:lang="ru">Argyle L.P., Busby E.C., Fulda N., Gubler J.R., Rytting C., Wingate D. Out of One, Many: Using Language Models to Simulate Human Samples // Political Analysis. 2023. Vol. 31. Iss. 3. P. 337-351. https://doi.org/10.1017/pan.2023.2 EDN: TOAPUN</mixed-citation></citation-alternatives></ref><ref id="B3"><label>3.</label><citation-alternatives><mixed-citation xml:lang="en">Puzanova, Zh.V., Kozhoridze, G.G., &amp; Kozhoridze, D.G. (2025). AI and Sociology: an Analysis of the Technological Capabilities of Virtual Respondents. Sotsiologiya: Metodologiya, Metody, Matematicheskoe Modelirovanie (Sotsiologiya: 4M), 60, 216–246. (In Russ.). https://doi.org/10.19181/4m.2025.34.1.6 EDN: PRPHTP</mixed-citation><mixed-citation xml:lang="ru">Пузанова Ж.В., Кожоридзе Г.Г., Кожоридзе Д.Г. ИИ и социология: анализ технологических возможностей виртуальных респондентов // Социология: методология, методы, математическое моделирование (Социология: 4М). 2025. № 60. С. 216-246. https://doi.org/10.19181/4m.2025.34.1.6 EDN: PRPHTP</mixed-citation></citation-alternatives></ref><ref id="B4"><label>4.</label><citation-alternatives><mixed-citation xml:lang="en">Aher, G., Arriaga, R.I., &amp; Kalai, A.T. (2023). Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies. In: Proceedings of the 40th International Conference on Machine Learning (Vol. 202, pp. 358–390). PMLR.</mixed-citation><mixed-citation xml:lang="ru">Aher G., Arriaga R.I., Kalai A.T. Using large language models to simulate multiple humans and replicate human subject studies // Proceedings of the 40th International Conference on Machine Learning. PMLR. 2023. Vol. 202. P. 358-390.</mixed-citation></citation-alternatives></ref><ref id="B5"><label>5.</label><citation-alternatives><mixed-citation xml:lang="en">Dubnyakova, O.A., &amp; Kashina, T.A. (2017). Communicative and Pragmatic Particularities of Personal Diary. MCU Journal of Philology. Theory of Linguistics. Linguistic Education, 3(23), 43–49. (In Russ.). EDN: YJGTID</mixed-citation><mixed-citation xml:lang="ru">Дубнякова О.А., Кашина Т.А. Коммуникативно-прагматические особенности личного дневника // Вестник МГПУ. Серия «Филология. Теория языка. Языковое образование». 2017. № 3 (23). С. 43-49. EDN: YJGTID</mixed-citation></citation-alternatives></ref><ref id="B6"><label>6.</label><citation-alternatives><mixed-citation xml:lang="en">Tannen, D. (1990). You Just Don't Understand: Women and Men in Conversation. New York: Ballantine Books. (In Russ.).</mixed-citation><mixed-citation xml:lang="ru">Таннен Д. Ты меня не понимаешь! : женщины и мужчины в разговоре / пер. с англ. Т. Печурко. М. : Вече ; Персей ; ACT, 1996.</mixed-citation></citation-alternatives></ref><ref id="B7"><label>7.</label><citation-alternatives><mixed-citation xml:lang="en">Kirilina, A.V. (1999). Gender: Linguistic Aspects. Moscow: Institute of Sociology, Russian Academy of Sciences. (In Russ.). EDN: OXTFTD</mixed-citation><mixed-citation xml:lang="ru">Кирилина А.В. Гендер: лингвистические аспекты : монография. М. : Ин-т социологии РАН, 1999. EDN: OXTFTD</mixed-citation></citation-alternatives></ref><ref id="B8"><label>8.</label><citation-alternatives><mixed-citation xml:lang="en">Potapov, V.V. (2002). A Multilevel Strategy in Linguistic Genderology. Topics in the Study of Language, 1, 103–126. (In Russ.). EDN: WVGGEK</mixed-citation><mixed-citation xml:lang="ru">Потапов В.В. Многоуровневая стратегия в лингвистической гендерологии // Вопросы языкознания. 2002. № 1. С. 103-126. EDN: WVGGEK</mixed-citation></citation-alternatives></ref><ref id="B9"><label>9.</label><citation-alternatives><mixed-citation xml:lang="en">Gender and language: An Anthology. (2005). Moscow: Yazyki slavyanskoy kultury. (In Russ.).</mixed-citation><mixed-citation xml:lang="ru">Гендер и язык : антология. М. : Языки славянской культуры, 2005.</mixed-citation></citation-alternatives></ref><ref id="B10"><label>10.</label><citation-alternatives><mixed-citation xml:lang="en">Popova, T.I. (2022). Social Roles of the Speaker in Everyday Communication in Russian: A Gender Perspective [PhD in Philology]. Saint Petersburg: Saint Petersburg State University. (In Russ.).</mixed-citation><mixed-citation xml:lang="ru">Попова Т.И. Социальные роли говорящего в повседневной коммуникации на русском языке: гендерный аспект : дис.. канд. филол. наук. СПб., 2022.</mixed-citation></citation-alternatives></ref><ref id="B11"><label>11.</label><citation-alternatives><mixed-citation xml:lang="en">Lutz, M., Sen, I., Ahnert, G., Rogers, E., &amp; Strohmaier, M. (2025). The Prompt Makes the Person(a): a Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models. arXiv. https://arxiv.org/abs/2507.16076</mixed-citation><mixed-citation xml:lang="ru">Lutz M., Sen I., Ahnert G., Rogers E., Strohmaier M. The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models. URL: arXiv.org. 2025. https://arxiv.org/abs/2507.16076</mixed-citation></citation-alternatives></ref><ref id="B12"><label>12.</label><citation-alternatives><mixed-citation xml:lang="en">Lans, J.C.W. (2025). The “Magic Word” for LLMS: A Study on the Effect of Politeness on LLM Performance [PhD Thesis]. Leiden: Leiden University, Leiden Institute of Advanced Computer Science, Leiden University.</mixed-citation><mixed-citation xml:lang="ru">Lans J.C.W. The “Magic Word” for LLMs: A Study on the Effect of Politeness on LLM Performance [PhD Thesis]. Leiden : Leiden Institute of Advanced Computer Science (LIACS), Leiden University, 2025.</mixed-citation></citation-alternatives></ref><ref id="B13"><label>13.</label><citation-alternatives><mixed-citation xml:lang="en">Dobariya, O., &amp; Kumar, A. (2025). Investigating How Prompt Politeness Affects LLM Accuracy (Short Paper). arXiv https://arxiv.org/abs/2510.04950</mixed-citation><mixed-citation xml:lang="ru">Dobariya O., Kumar A. Investigating how prompt politeness affects LLM accuracy (short paper). URL: arXiv. 2025. https://arxiv.org/abs/2510.04950</mixed-citation></citation-alternatives></ref><ref id="B14"><label>14.</label><citation-alternatives><mixed-citation xml:lang="en">Qi, P., Zhang, Y., Zhang, Y., Bolton, J., &amp; Manning, C.D. (2020). Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. arXiv. https://arxiv.org/abs/2003.07082</mixed-citation><mixed-citation xml:lang="ru">Шмелeва Т.В. Речевой жанр (возможности описания и использования в преподавании языка) // Russistik. Русистика: научный журнал актуальных проблем преподавания русского языка. 1990. № 2. С. 20-32.</mixed-citation></citation-alternatives></ref><ref id="B15"><label>15.</label><citation-alternatives><mixed-citation xml:lang="en">Shmeleva, T.V. (1990). Speech Genre: Possibilities of Description and Use in Language Teaching. Russistik. Russian Language Studies: A Journal of Current Problems of Teaching Russian, 2, 20–32. (In Russ.).</mixed-citation><mixed-citation xml:lang="ru">Моисеева Т.Ф., Огорелков И.В. Современная автороведческая экспертиза. М. : Российский государственный университет правосудия (РГУП), 2022.</mixed-citation></citation-alternatives></ref><ref id="B16"><label>16.</label><mixed-citation>Moiseeva, T.F., &amp; Ogorelkov, I.V. (2022). Modern Authorship Examination. Moscow: Russian State University of Justice (RGUP). (In Russ.).</mixed-citation></ref></ref-list></back></article>
