Transformation-Based Generation of English Language Exercises: a Deterministic and CEFR-Guided Approach
- Authors: Popova M.V.1
-
Affiliations:
- Moscow State Linguistic University
- Issue: Vol 17, No 2 (2026): ARTIFICIAL INTELLIGENCE IN APPLIED LINGUISTICS: CURRENT CHALLENGES AND PROSPECTS
- Pages: 569-584
- Section: ARTIFICIAL INTELLIGENCE IN APPLIED LINGUISTICS: CURRENT CHALLENGES AND PROSPECTS
- URL: https://journals.rudn.ru/semiotics-semantics/article/view/52661
- DOI: https://doi.org/10.22363/2313-2299-2026-17-2-569-584
- EDN: https://elibrary.ru/LTBNXB
- ID: 52661
Cite item
Full Text
Abstract
The developed model provides a deterministic, transformation-based approach to generating English language exercises from user-provided texts. The system is designed as a rule-based AI tool that restructures existing textual material into pedagogically organized tasks while preserving structural and contextual coherence. Unlike large probabilistic language models that autonomously generate discourse, the proposed approach relies on explicit linguistic rules and CEFR-calibrated constraints (A1-C2), ensuring reproducibility and transparency of output. The study conceptualizes AI-assisted exercise generation as a structured process of text transformation rather than autonomous content production. The model maintains traceability between source text elements and generated tasks and avoids probabilistic text synthesis. CEFR integration functions as a constraint-based filtering mechanism that regulates lexical and grammatical selection according to proficiency descriptors. A qualitative pilot study conducted in a tertiary EFL context demonstrates that the system produces linguistically valid and pedagogically usable exercises derived from authentic texts. The findings indicate that rule-based AI systems can support instructional material preparation while preserving teacher control and structural stability. The study contributes to ongoing discussions in applied linguistics on transparent and theory-informed integration of artificial intelligence into foreign language education.
Full Text
Introduction
Artificial intelligence has emerged as a central methodological and epistemological issue in contemporary language education. While recent advances in large language models have enabled autonomous text generation at unprecedented scale, their integration into foreign language pedagogy raises fundamental questions regarding interpretability, pedagogical control, and construct validity [1; 2]. Student perceptions of AI in foreign language learning contexts frequently reflect concerns regarding transparency, reliability, and pedagogical trust [3]. Within applied linguistics, the discussion increasingly distinguishes between generative AI systems that synthesize new discourse and controlled computational tools designed to support instructional design [4; 5].
From a computer-assisted language learning (CALL) perspective, AI technologies must be evaluated not solely in terms of technical sophistication but in relation to their alignment with second language acquisition (SLA) theory and task design principles [6; 7]. Language pedagogy is grounded in structured input, meaningful task engagement, and construct-valid assessment practices [8; 9]. Consequently, fully autonomous probabilistic generation may conflict with pedagogical requirements of traceability, contextual integrity, and proficiency calibration.
In this context, the present study proposes a transformation-based, deterministic model for generating English language exercises from teacher-selected texts. Instead of synthesizing new discourse, the system restructures existing textual material into controlled pedagogical formats governed by explicit linguistic rules and CEFR-calibrated constraints (A1–C2). The model conceptualizes AI-assisted exercise generation as a structured process of secondary textual transformation embedded within established proficiency frameworks.
The study pursues three objectives: (1) to describe the system architecture, (2) to position it within contemporary AI-supported language education research, and (3) to evaluate its pedagogical viability through a qualitative pilot implementation in a tertiary EFL context.
By foregrounding interpretability, reproducibility, and proficiency-sensitive filtering, the study contributes to ongoing discussions in applied linguistics concerning theory-informed and pedagogically responsible AI integration.
Theoretical Framework
AI in Foreign Language Education
The rapid development of artificial intelligence technologies has significantly reshaped the landscape of foreign language education. Recent advances in large language models (LLMs) have enabled autonomous text generation, adaptive feedback, and real-time conversational interaction, expanding the range of AI-supported instructional tools [1; 2]. However, the integration of such systems into pedagogical contexts raises theoretical and methodological concerns regarding transparency, controllability, and construct alignment.
Contemporary research distinguishes between generative, adaptive, and assistive AI systems in education [4; 5]. Generative systems prioritize autonomous content production; adaptive systems focus on personalization through data-driven modeling; assistive systems support specific instructional functions without replacing pedagogical decision-making. While generative AI offers flexibility and scalability, its probabilistic architecture may introduce opacity, unpredictability, and reduced structural correspondence of output — factors that complicate pedagogical validation [2].
Within applied linguistics, these concerns intersect with broader theoretical questions about the relationship between technological mediation and language acquisition processes. From a CALL perspective, technological tools must be evaluated not solely in terms of technical sophistication but in relation to established principles of second language acquisition and instructional design [6].
Alignment with SLA and Task-Based Language Teaching
Second language acquisition research emphasizes the importance of structured input, meaningful task engagement, and form-meaning mapping in the development of communicative competence [7]. Task-based language teaching conceptualizes pedagogical tasks as goal-oriented activities grounded in authentic discourse and designed to promote language processing within meaningful contexts.
Unconstrained AI text generation may risk producing linguistically correct yet pedagogically misaligned content, particularly when task formats lack explicit linkage to input discourse. In contrast, rule-governed transformation architectures that derive exercises directly from teacher-selected texts maintain alignment between input, processing demands, and output formats. Such alignment reflects principles of instructional coherence central to task-based pedagogy [7].
From a vocabulary acquisition perspective, systematic lexical selection and frequency control are essential components of effective instructional design [9; 10]. AI systems that incorporate structured lexical filtering mechanisms may therefore better support vocabulary development than systems relying on unconstrained generative output.
Construct Validity and Assessment-Oriented Task Design
Exercise generation in language education cannot be considered independently of assessment theory. According to Bachman and Palmer (1996) [8], the validity of language tasks depends on construct representation, authenticity, and the relationship between input and expected response behavior. Selected-response formats, cloze procedures, and controlled production tasks each target distinct aspects of language ability and must be grounded in clearly defined constructs. From this perspective, AI-mediated task generation should preserve construct validity and ensure that response formats are meaningfully derived from the linguistic features of the source text. The preservation of textual anchoring enhances authenticity and reduces construct-irrelevant variance, particularly in vocabulary-focused and comprehension-oriented tasks [11].
Thus, the theoretical requirement is not merely technological feasibility but principled alignment between linguistic analysis, task format, and proficiency descriptors.
CEFR, Readability, and Proficiency-Sensitive Filtering
The Common European Framework of Reference for Languages (CEFR) provides internationally recognized descriptors of communicative competence across proficiency levels1. In computational language learning systems, CEFR alignment is often operationalized through readability classification and lexical profiling [12].
Research in NLP-informed CALL demonstrates that proficiency-sensitive filtering mechanisms can be implemented through structured lexical resources and readability modeling [13]. Such systems prioritize controlled selection of linguistic units rather than probabilistic generation. Early ICALL architectures based on modular NLP component reuse further established the viability of linguistically grounded exercise construction [14].
These approaches highlight an important distinction: proficiency calibration may function either as adaptive personalization driven by statistical modeling or as constraint-based filtering embedded within deterministic rule systems. The latter approach emphasizes reproducibility, interpretability, and architectural explicability.
Toward Deterministic, Linguistically Grounded AI Architectures
Recent debates in AI-supported education increasingly stress the need for explainable, controllable, and pedagogically accountable systems [1; 2]. While generative AI demonstrates impressive fluency, its stochastic nature complicates reproducibility and instructional validation.
Deterministic transformation-based models represent an alternative paradigm in which AI functions as a structured processing mechanism rather than an autonomous discourse generator. By integrating linguistic preprocessing, CEFR-calibrated filtering, and template-based instantiation, such systems maintain input-task alignment. This architecture aligns with SLA theory, task-based pedagogy, and assessment validity principles, thereby positioning deterministic AI as a theoretically grounded complement to generative systems.
The present study builds upon this theoretical foundation by proposing a controlled text-transformation system that operationalizes proficiency-sensitive filtering within a rule-governed, reproducible pipeline.
Model Description
General Architecture and Design Principles
The proposed system is an AI-assisted tool for generating English language learning exercises from user-provided texts. The model is implemented as an open-access Python-based software repository developed by the author in cooperation with L. Rybalko and made available for research and educational purposes2.
Technically, the system is written in Python and relies on standard NLP procedures for structural text analysis. The processing pipeline includes automated tokenization, sentence segmentation, part-of-speech tagging, and integration of CEFR-labelled lexical resources. These structural annotations serve as the basis for rule-based pedagogical transformations. Internally, the system operates as a sequence of rule-governed transformation modules acting on explicitly defined linguistic units (tokens, sentences, grammatical categories).
Unlike large-scale end-to-end generative language models that rely on probabilistic text synthesis, the present system follows a rule-based, text-driven architecture. It does not autonomously generate new discourse. Instead, it performs structured transformation of existing textual material into pedagogically organized tasks.
The generation mechanism consists of interconnected components, demonstrated in Figure 1.
- Linguistic preprocessing module performs segmentation, tokenization, lemmatization, part-of-speech tagging, and statistical feature extraction. This stage identifies structurally transformable linguistic units and generates morphosyntactic annotations required for downstream processing.
- Vectorization and feature extraction module converts selected textual units into structured numerical representations and extracts distributional and structural features that support controlled target selection and distractor generation.
- CEFR — constrained filtering module applies proficiency-level constraints (A1–C2), restricting lexical and grammatical elements according to CEFR descriptors and calibrated lexical resources.
- Rule — based target unit selection module identifies candidate lexical or grammatical elements suitable for pedagogical transformation based on explicitly defined linguistic rules.
- Template — based exercise generation module embeds selected units into predefined task structures, ensuring structural consistency and contextual anchoring.
- Output module produces fully instantiated exercises ready for instructional use.
Figure 1. Pipeline architecture of the CEFR-guided rule-based exercise generation system
Source: Compiled by Marianna V. Popova.
A distinctive feature of the model is the integration of CEFR standards. The system incorporates proficiency-level parameters ranging from A1 to C2. During exercise generation, the instructor may specify the intended CEFR level, which influences:
- selection of lexical items appropriate to the target level,
- filtering of grammatical constructions,
- control of task complexity and structural density.
The CEFR parameterization does not function as probabilistic adaptation but as a constraint-based filtering mechanism embedded within the transformation rules. Proficiency-sensitive filtering mechanisms are closely related to readability classification research in NLP-informed language learning systems [12]. Early ICALL systems implementing rule-based sentence selection and CEFR-calibrated readability assessment have demonstrated the pedagogical viability of linguistically controlled filtering [13]. The present model extends this logic from sentence selection to structured exercise generation.
The CEFR-calibrated filtering module relies on externally available CEFR-labelled lexical datasets. Specifically, the model integrates lexical frequency and level information derived from publicly accessible CEFR-annotated word lists, including the “10,000 English Words CEFR Labelled” dataset3 and the UniversalCEFR English lexical dataset4. These resources provide level-based lexical categorization (A1–C2), which is incorporated into the rule-based filtering mechanism to constrain lexical selection according to proficiency parameters. The model incorporates CEFR-labelled lexical resources as structured reference datasets within its deterministic filtering layer.
All transformation operations are reproducible: identical input under identical CEFR parameters yields identical output. This architectural decision ensures output stability, procedural reproducibility, and pedagogical reliability.
The system is designed according to three core principles:
- Text-dependence — all exercises are derived exclusively from the submitted source text.
- Textual traceability — each generated task can be linked to identifiable lexical or syntactic units in the original discourse.
- Pedagogical controllability — instructors retain full authority to review, interpret, and integrate outputs into instructional contexts.
The model may be understood as a controlled mechanism of structured transformation regulated not only by linguistic form but also by proficiency-based constraints. The integration of CEFR descriptors situates the system within established frameworks of communicative competence, reinforcing its alignment with contemporary language pedagogy.
This configuration positions the system as an architecturally explicit AI-mediated processing tool that combines structural linguistic analysis with proficiency-calibrated pedagogical transformation. All processing is executed locally without reliance on external large language model APIs.
Exercise Types and Generation Logic
The current implementation supports four pedagogically established task formats widely used in foreign language instruction and assessment. Each task type is derived through rule-governed transformations applied to the source text, ensuring alignment with communicative, lexical, and form-focused learning objectives [11; 16].
Selected-Response Vocabulary Items (Multiple-Choice Format)
The system generates selected-response vocabulary items in a multiple-choice format, a task type extensively discussed in contemporary language assessment research [11]. Target lexical units are extracted from the source text and presented as correct responses. Distractors are generated through controlled lexical variation procedures designed to maintain semantic proximity and grammatical plausibility while minimizing construct-irrelevant variance. From an assessment perspective, multiple-choice vocabulary items allow controlled measurement of receptive lexical knowledge when distractor quality is systematically regulated [8]. By restricting response options to contextually compatible material, the system preserves discourse integrity and avoids introducing content external to the input domain.
Cloze and Controlled Gap-Fill Tasks
Selected lexical or grammatical elements are systematically omitted from sentences based on predefined linguistic criteria (e.g., part of speech, grammatical function, CEFR level). The preservation of the original syntactic frame supports contextual inference and reinforces form-meaning mapping, consistent with usage-based perspectives on second language acquisition [7]. Such tasks constitute semi-controlled production activities in which retrieval occurs under discourse-level constraints instead of decontextualized rule application.
Meaning-Focused Comprehension Tasks (Drag-and-Drop Format)
The system also generates structured comprehension tasks implemented in a drag-and-drop format. These tasks correspond to meaning-focused reading activities within task-based language teaching frameworks [7]. By requiring learners to interpret sentence-level and discourse-level meaning, these tasks engage interpretive and discourse competencies as described in communicative competence models [8]. The preservation of textual cohesion ensures alignment with authenticity-oriented instructional approaches [15].
Lexical Puzzle Tasks (Word Reconstruction)
Finally, the system produces lexical reconstruction and word manipulation tasks designed to reinforce orthographic awareness and lexical retrieval. From a vocabulary pedagogy perspective, such tasks contribute to consolidation within balanced lexical instruction models [9; 16]. These exercises integrate attention to form within meaningful textual input, aligning with form-focused instruction embedded in communicative contexts [7].
Across all formats, the generation mechanism preserves source-text anchoring and structural coherence. Tasks are derived exclusively from the original discourse, ensuring referential integrity and contextual consistency. From an assessment perspective, this design reduces construct-irrelevant variance and supports alignment between task demands and the input domain [8]. From an SLA perspective, it integrates meaning-focused and form-focused strands of instruction within a unified rule-governed transformation framework [7; 16].
Workflow
The operational workflow of the system consists of the following sequential stages:
- The user uploads an English text (.txt) and specifies the target CEFR level (A1–C2).
- The system performs automated structural analysis of the input text, including tokenization, sentence segmentation, and part-of-speech tagging. This stage produces linguistically annotated textual units that serve as the basis for subsequent operations.
- Following preprocessing, the system applies a CEFR-calibrated filtering layer. Lexical items are evaluated against integrated CEFR-labelled reference datasets, and selection constraints are imposed according to the specified proficiency level. This stage functions as a constraint-based selection mechanism and does not rely on predictive classification models.
- Linguistically relevant elements (e.g., lexical items, grammatical constructions) are selected according to predefined transformation rules and proficiency constraints.
- Selected textual units are embedded into predefined exercise templates (e.g., selected-response items, cloze tasks, reconstruction tasks), resulting in structured pedagogical outputs.
- Generated exercises are presented for review and use.
Pedagogical Scope and Constraints
The model is intentionally limited to rule-based generation. It does not incorporate large-scale generative language models, probabilistic text synthesis, or adaptive learner profiling mechanisms.
These constraints serve three purposes:
- ensuring output stability,
- maintaining transparency of transformation procedures,
- preserving teacher control over instructional content.
The absence of autonomous generation reduces risks associated with semantic drift or contextually detached output, which are frequently discussed in relation to large language models [1; 2].
At present, difficulty differentiation remains instructor-mediated and does not involve algorithmic adaptation. The system is designed as a material preparation instrument and is not intended to serve as a personalized tutoring environment.
Cross-Linguistic Adaptability: Extension to German
Although the present implementation is designed for English language instruction, the architecture of the system is not language-specific. The modular NLP pipeline and constraint-based generation logic allow adaptation to other morphologically richer languages, including German. Such adaptation would require replacing the English preprocessing components with a German language model supporting tokenization, lemmatization, and morphosyntactic tagging, as well as integrating CEFR-aligned German lexical resources.
In the case of German, additional rule-layer refinement would be necessary to account for grammatical gender, case marking, separable verb constructions, and agreement phenomena. The CEFR-calibrated filtering mechanism could be extended using German CEFR-annotated lexical datasets, e.g., CEFRLex German, ensuring proficiency-sensitive constraint-based selection without introducing probabilistic modeling.
Importantly, the deterministic transformation architecture would remain unchanged. The system’s rule-governed design supports cross-linguistic scalability by isolating language-dependent resources within configurable preprocessing and filtering modules. This extensibility highlights the broader applicability of transformation-based exercise generation beyond English language pedagogy. A further development of the system could involve the implementation of a bilingual English–German version, enabling parallel constraint-based exercise generation across both languages within a unified, multilingual architectural framework.
Pilot Study
Research Design & Materials
The pilot implementation was designed as a qualitative exploratory study aimed at assessing the pedagogical feasibility and structural validity of the model. The objective was not to measure learning outcomes quantitatively but to examine the system’s performance in authentic instructional conditions and assess the structural and pedagogical properties of the generated exercises.
The pilot involved N = 5 university-level EFL instructors with teaching experience ranging from 3 to 15 years. All participants regularly design supplementary materials based on authentic English texts. Participation was voluntary.
A total of N = 25 authentic English texts were used during the pilot phase. The texts ranged from 200 to 400 words and represented three genres:
- informational academic prose,
- journalistic articles,
- short narrative excerpts.
The selected materials covered CEFR levels B1 to C1. No artificial simplification or preprocessing was applied prior to submission.
Procedure and Evaluation Criteria
Each instructor submitted 5 texts to the system and specified a CEFR level parameter (A2–C1 depending on instructional context). The model generated:
- selected-response vocabulary items (multiple-choice format);
- cloze and controlled gap-fill tasks;
- meaning-focused comprehension tasks implemented in a drag-and-drop format;
- lexical reconstruction tasks.
Generated outputs were analyzed without prior manual modification in order to assess immediate usability. Instructor reflections were collected through structured written commentary focusing on three evaluation criteria mentioned in the model description section.
The analysis combined:
- qualitative thematic coding of instructor feedback,
- structural review of generated exercises,
- comparison of CEFR parameter settings with observed task complexity.
No learner data were collected during this phase. The pilot was restricted to instructor-mediated evaluation of discursive restructuring quality.
Findings
The results indicate that the system consistently preserved traceable linkage between source texts and generated exercises. Tasks remained lexically and syntactically anchored in the source texts. CEFR calibration effectively constrained lexical density and grammatical complexity. Instructors reported high usability in contexts requiring rapid preparation of text-based practice materials. However, variability of exercise formulation remained limited due to the deterministic template structure.
The findings confirm that the model functions as a controlled transformation mechanism and not as an autonomous content generator. The exploratory design of the pilot study limits generalizability, and further controlled learner-based validation is required to assess learning effectiveness.
Discussion
The findings of the pilot implementation allow the proposed system to be situated within a distinct transformation-based paradigm of AI-supported language education. As outlined in the theoretical framework, contemporary research differentiates between generative, adaptive, and assistive AI systems [4; 5]. While generative systems prioritize autonomous discourse production and adaptive systems emphasize personalization, the present model operates through structured transformation of teacher-selected input. This architectural choice directly addresses concerns regarding transparency and pedagogical control raised in recent discussions of large language models in education [1; 2]. By maintaining structural traceability between input text and generated task formats, the system preserves source-text anchoring and reproducibility — features that are central to CALL design principles [6]. In doing so, the study addresses concerns regarding opacity, reduced structural traceability, diminished pedagogical control, and weakened construct alignment raised in current AI-in-education debates. In this framework, AI is not treated as an autonomous pedagogical agent but as a controlled linguistic processing mechanism integrated into instructional design principles.
From a second language acquisition perspective, pedagogical tasks must promote structured input processing and meaningful form–meaning integration [7]. The rule-governed transformation architecture supports this requirement by deriving all exercises directly from the source text, ensuring that learners engage with authentic discourse rather than decontextualized or externally generated content. The preservation of textual anchoring reinforces coherence between input domain and task demand. This alignment reduces the risk of construct-irrelevant variance and supports principled task construction. In contrast to unconstrained generative output, exercise generation maintains consistency between linguistic features of the text and cognitive processing required by the learner. Moreover, vocabulary-focused formats within the system rely on controlled lexical selection informed by frequency and proficiency descriptors. This design corresponds to established principles of vocabulary instruction emphasizing systematic exposure and level-appropriate lexical control [9; 10]. Thus, the system’s pedagogical strength lies not in generative capacity but in its ability to operationalize SLA-informed task design through computational structuring.
Language tasks function as instruments for eliciting evidence of underlying language ability. According to Bachman and Palmer (1996) [8], task validity depends on coherent construct representation, authenticity, and the relationship between input and expected response behavior. The present model preserves construct alignment by ensuring that all selected-response items, cloze tasks, and comprehension activities are structurally derived from the input text. This reduces construct-irrelevant content and enhances authenticity by grounding tasks in meaningful discourse. Selected-response vocabulary items and controlled gap-fill tasks correspond to established assessment formats targeting lexical knowledge and grammatical processing [11]. AI-assisted assessment formats may also influence learner motivation, anxiety, and affective engagement, thereby shaping task effectiveness beyond purely cognitive dimensions [17]. Their computational instantiation within a deterministic framework ensures reproducibility—an important but often underemphasized dimension of assessment design in AI-mediated contexts. By embedding CEFR-based filtering within the generation pipeline, the system further aligns task difficulty with internationally recognized proficiency descriptors. This proficiency-sensitive constraint mechanism reflects research in readability classification and CEFR-based ICALL systems [12; 13].
The system extends the tradition of ICALL architectures that rely on modular NLP components and linguistically controlled selection [14]. Similar to corpus-driven extraction systems such as GDEX [18], the model prioritizes controlled extraction over unconstrained generation. However, it advances this logic by integrating CEFR-calibrated filtering with template-based exercise instantiation. So, the model exemplifies an alternative trajectory within AI-supported education: rather than replacing pedagogical structure with probabilistic fluency, it operationalizes instructional design principles through deterministic computational procedures. This approach may be particularly relevant in higher education contexts where transparency, accountability, and alignment with standardized proficiency frameworks are critical.
The findings suggest that AI integration in language education does not necessitate autonomous generative capacity. Instead, linguistically grounded transformation architectures can provide scalable support for instructional material preparation while preserving teacher agency and theoretical coherence.
From an applied linguistics perspective, the model shows that computational systems can embody SLA theory, and assessment validity principles, moving beyond the replication of surface-level fluency. In this sense, AI functions as a mediating infrastructure that operationalizes established pedagogical constructs. Such architectures contribute to ongoing debates concerning interpretability, deterministic reproducibility, and responsible AI implementation in education [1; 2]. By foregrounding deterministic design and CEFR-constrained filtering, the study proposes a model of AI integration that prioritizes pedagogical structure over generative autonomy.
Conclusions
This study has presented a deterministic transformation model for generating CEFR-aligned English language exercises from teacher-selected texts. Within applied speech studies, the model demonstrates how computational architectures can restructure discourse while preserving referential and communicative integrity. The system was conceptualized not as an autonomous generative agent but as a linguistically grounded processing architecture that restructures existing discourse into pedagogically controlled task formats.
The theoretical positioning of the model draws on established principles of CALL research, task-based language teaching, vocabulary pedagogy, and construct validity in language assessment. Within this framework, AI-assisted exercise generation is understood as a structured transformation process governed by explicit linguistic rules and proficiency-sensitive constraints rather than probabilistic discourse synthesis.
The qualitative pilot implementation in a tertiary EFL context provides initial empirical support for the pedagogical viability of the system. By maintaining coherence between source input and generated formats, the model supports construct-consistent task design and reduces risks associated with opaque generative architectures.
From the perspective of applied linguistics, the study contributes to ongoing discussions on responsible and theory-informed AI integration in language education. It demonstrates that computational systems can operationalize SLA-informed pedagogical principles while preserving interpretability and teacher agency. The study therefore advocates for interpretable AI architectures that embed linguistic theory within computational design instead of relying on probabilistic autonomy.
Future research should extend the present findings through controlled learner-based evaluation, quantitative analysis of task effectiveness, and comparison with generative AI systems across proficiency levels. Further development may also explore expanded task formats and integration with adaptive feedback mechanisms while preserving architectural interpretability.
1 Council of Europe. (2001). Common European Framework of Reference for Languages: Learning, Teaching, Assessment. Cambridge: Cambridge University Press.
2 Rybalko, L., & Popova, M.V. (2025). English Exercise Generator [Computer software]. GitHub. URL: https://github.com/ludryb/generate_exercises/tree/main (accessed: 10.11.2025).
3 Nezahatkk. (n.d.). 10,000 English words CEFR labelled [Dataset]. Kaggle. URL:https://www.kaggle.com/datasets/nezahatkk/10-000-english-words-cerf-labelled (accessed: 10.11.2025).
4 UniversalCEFR. (n.d.). CEFR SP-EN dataset [Dataset]. Hugging Face. URL: https://huggingface.co/datasets/UniversalCEFR/cefr_sp_en (accessed: 10.11.2025).
About the authors
Marianna V. Popova
Moscow State Linguistic University
Author for correspondence.
Email: neunerin@gmail.com
ORCID iD: 0000-0003-4302-7054
SPIN-code: 9048-5511
ResearcherId: AEA-4819-2022
PhD in Philology, Associate Professor, Leading Researcher at the Experimental Phonetics and Forensic Linguistics Laboratory
38 build 1 Ostozhenka str., Moscow, Russian Federation, 119034References
- Kim, E.J. (2025). AI-Assisted English Learning: A Tool for All or Only а Select Few? Exploring Learner Difficulties and Group-Specific Effects in Korean EFL Contexts. Language Learning & Technology, 29(1), 1-22. https://doi.org/10.64152/10125/73633
- Kebble, P. (2023). A Chat with ChatGPT: The Potential Impact of Generative AI in Higher Education Learning, Teaching and Assessment, with Specific Reference to EAL/D Students. Journal of Academic Language and Learning, 17(1), T81-T91.
- Tikhonova, N.V., & Ilduganova, G.M. (2024). “I Am Afraid of How Fast Artificial Intelligence Is Developing”: Students’ Perceptions of AI in Foreign Language Learning. Higher Education in Russia, 33(4), 63-83. (In Russ.). https://doi.org/10.31992/0869-3617-2024-33-4-63-83 EDN: FNUAVR
- Pokrivčáková, S. (2019). Preparing Teachers for the Application of AI-Powered Technologies in Foreign Language Education. Journal of Language and Cultural Education, 7(3), 135-153. https://doi.org/10.2478/jolace-2019-0025
- Titova, S.V. (2024). AI-Based Technological Solutions in Foreign Language Education. Lomonosov Linguistics and Intercultural Communication Journal, 27(2), 18-37. (In Russ.). https://doi.org/10.55959/MSU-2074-1588-19-27-2-2. EDN: OWSQVG
- Chapelle, C.A. (2001). Computer Applications in Second Language Acquisition: Foundations for Teaching, Testing and Research. Cambridge: Cambridge University Press.
- Ellis, R. (2003). Task-Based Language Learning and Teaching. Oxford: Oxford University Press.
- Bachman, L.F., & Palmer, A.S. (1996). Language Testing in Practice: Designing and Developing Useful Language Tests. Oxford: Oxford University Press.
- Nation, I.S.P. (2001). Learning Vocabulary in Another Language. Cambridge: Cambridge University Press.
- Schmitt, N. (2008). Review Article: Instructed Second Language Vocabulary Learning. Language Teaching Research, 12(3), 329-363.
- Brown, J.D., & Abeywickrama, P. (2019). Language Assessment: Principles and Classroom Practices. Pearson.
- Vajjala, S., & Meurers, D. (2012). On Improving the Accuracy of Readability Classification Using Insights from Second Language Acquisition. In: Proceedings of the 7th Workshop on Innovative Use of NLP for Building Educational Applications (BEA7) (pp. 163-173).
- Pilán, I., Volodina, E., & Johansson, R. (2013). Automatic Selection of Suitable Sentences for Language Learning Exercises. In: 20 Years of EUROCALL: Learning from the Past, Looking to the Future. L. Bradley & S. Thouësny (eds.). Proceedings of the 2013 EUROCALL Conference, Évora, Portugal (pp. 218-225). Dublin-Voillans.
- Volodina, E., Borin, L., Loftsson, H., Arnbjörnsdóttir, B., & Leifsson, G.Ö. (2012). Waste Not, Want Not: Towards a System Architecture for ICALL Based on NLP Component Re-Use. In: Workshop on NLP in Computer-Assisted Language Learning. Linköping Electronic Conference Proceedings 80 (pp. 47-58). Linköping, Sweden.
- Gilmore, A. (2007). Authentic Materials and Authenticity in Foreign Language Learning. Language Teaching, 40(2), 97-118. https://doi.org/10.1017/S0261444807004144
- Nation, I.S.P. (2007). The Four Strands. Innovation in Language Learning and Teaching, 1(1), 2-13. https://doi.org/10.2167/illt039.0
- Biju, N., et al. (2024). Which One? AI-Assisted Language Assessment or Paper Format: An Exploration of the Impacts on Foreign Language Anxiety, Learning Attitudes, Motivation, and Writing Performance. Language Testing in Asia, 14(1), 45. https://doi.org/10.1186/s40468-024-00245-x
- Kilgarriff, A., Husák, M., McAdam, K., Rundell, M., & Rychlý, P. (2008). GDEX: Automatically Finding Good Dictionary Examples in a Corpus. In: Proceedings of Euralex 2008.
Supplementary files
Source: Compiled by Marianna V. Popova.








