Machine vs Human: Large Language Model Errors as a Means of Creating Stylistically Appropriate Text
- Authors: Zabolotskikh A.V.1, Sokolova E.E.2, Tavberidze D.V.1
-
Affiliations:
- RUDN University
- Moscow Institute of Physics and Technology
- Issue: Vol 17, No 2 (2026): ARTIFICIAL INTELLIGENCE IN APPLIED LINGUISTICS: CURRENT CHALLENGES AND PROSPECTS
- Pages: 585-604
- Section: ARTIFICIAL INTELLIGENCE IN APPLIED LINGUISTICS: CURRENT CHALLENGES AND PROSPECTS
- URL: https://journals.rudn.ru/semiotics-semantics/article/view/52662
- DOI: https://doi.org/10.22363/2313-2299-2026-17-2-585-604
- EDN: https://elibrary.ru/LTTOJA
- ID: 52662
Cite item
Full Text
Abstract
The relationship between humans and machines in education is analyzing the characteristic features of the LLM artificial language as a tool for teaching the creation of scientific written text through a critical analysis of errors made in texts created by man and machine. This departure from traditional ideas about learning to write in a foreign language reflects progressive processes in teaching, since identifying parallels between human learning and model training based on mistakes refers to a high-tech approach focused on the student rather than the teacher. LLM, as a tool capable of correcting and creating generated text, acts as an incentive for the development of critical thinking when analyzing inaccuracies or evaluating the text “appropriatness”. LLM’s ability to generate different results can be used as a error-based methodology for teaching academic writing. In this regard, academic writing training is a critical analysis and editing of text by visualizing errors, creating a sense of competition between a machine and a human and stimulating mental activity. This seems to be a clear example of how inconsistency and ambiguity can be used for development and creation.
Keywords
Full Text
Introduction
More and more articles devoted to the usage of Large Language Models (LLMs) for development critical thinking skills in education are overwhelming academia and exciting the minds of researchers. The advancement of LLMs such as Chat GPT 4, Gemini Ultra, Llama 3.1 (70B) [1. P. 1–75], etc. provide a unique opportunity to reconsider the methodological approach to teaching writing in higher education [2]. The user interaction with LLMs is usually performed as a series of natural language queries, which are answered by the model [3]. Internally, those systems use the transformer neural network architecture to predict the next words/tokens that are most likely to follow in such a dialogue format [4]. Thus, the use of these technologies in the methodology of teaching academic writing has begun the avalanche of enthusiastic research [5–10]. Furthermore, it is argued that the emergence of intelligent text generators may represent ‘a breakthrough in written communication format since the invention of the word processor’ (It is the biggest transformation of the writing process since the word processor) [11. P. 691]. Without detracting from LLMs merits as tools for creating and proofreading texts, we will try to consider the limitations which appeal to the integration of LLMs in assisting the learning process.
Terminological Inexactitude
The emergence of transformers which are actively used in natural language processing (NLP) and speech processing has introduced a certain ‘terminological inexactitude’ [12. P. 474] and confused the essence of concepts that refer to them. Primarily, LLMs were called an ‘artificial intelligence’ as part of an advertising campaign, which in fact was an artifice. This provoked the mass release of scientific articles where the authors draw direct analogies and compare the choice of language means of LLMs with the natural language used by humans, i.e. ‘humanise’ them [13]. From our perspective, this hobbles the efforts in understanding how humans should perceive the artificial language, what place should be offered to LLMs, and how we can define their scope and participation in higher education. The given paper considers LLMs as a tool in a pragmatically-orientated approach to teaching the scientific discourse generation, as the use of deep learning technology and multi-word sequencing of natural language has radically improved [14]. However, the tool which is widely used for writing abstracts for new preprints can produce inconsistent results: most abstracts contain false information, do not correspond to queries written in prompts, i.e. contain hallucinations or errors [15. P. 1–38; 16].
Use Errors in Your Favour
Error analysis has long been considered an effective technique in the field of English language teaching. Feedback on errors in recognizing and correcting mispronunciations for the learners was described as a powerful means to improve phonetics [17. P. 1–21). The support system may contribute to learning vocabulary by generating images corresponding to incorrect answers suggesting errors to the learner [18]. Error-driven teaching stems from the idea that the teacher’s impression of an output being an authentic representation of input is false [19]. In contrast, error analysis enables teachers to provide counterexamples encouraging reflection and understanding. Since errors are often seen as token-based being embedded in other errors [20] it seems possible to involve LLM in generating examples for error analysis.
Departure from Traditions
We need, therefore, to base our reasons on the previous teaching methods which has changed the educational landscape from a low-tech approach to learning to a high-tech one, as this departure from traditions reflects the dynamic processes in teaching, especially with the introduction of LLMs in it. Drawing parallels between teaching a human and training a model, we can observe similarity in both cases. LLMs’ reinforced learning with rewards and punishment can be attributed to a teacher-centred approach (a low-tech one), and error-driven learning applied for LLMs — to a student-centered approach (a high-tech one) [21]. The difference is also obvious. A human is capable of analysing rules, recognizing patterns, and identifying mistakes from just a few examples, whereas LLMs require an extensive corpus of texts usually created by huge teams of workers engaged in generating synthetic examples.
Contrary to popular belief, that the quality of learning outcomes based on generated texts depend on human-written prompts [22], we suggest that developing effective prompts can be a challenge, leading to potentially inaccurate or contextually inadequate responses. This emphasises the importance of understanding the mechanisms of developing prompts to achieve accurate results.
Problem Statement
Most of the works, which are devoted to the academic opportunities and challenges of using LLMs [23] describe examples of solving practical problems of teaching writing without theoretical framework which might provide deep insights into the potential risks of using LLMs and attributing intelligence to them.
Furthermore, understanding of this problem reveals a certain complexity due to its interdisciplinary context. At present, for language instructors who teach academic writing in higher education it is not enough to possess only linguistic knowledge of syntactic structures, generative grammar, and natural language semantics developed by the pioneers in these fields Noam Chomsky and Richard Montague (“English as a formal language”). Without understanding, at least in general terms, how large language models are built, it is impossible to interact effectively with LLMs in the learning process.
Considering the outlined challenges, the given article describes the approach to the integration of LLMs into English Academic Writing course for postgraduate students at MIPT and RUDN with the emphasis on teaching writing the basic elements of the Introduction section: research question, objectives, gap, hook, hypothesis, criteria for problem selection and a key message.
The article has the following structure: We begin with revealing a theoretical rationale for language description approaches to understand natural language (NL) and artificial language (AL) crossing points and distinctions to show that IT NL operates on different laws than Transformers. Then we try to determine the place and role of LLMs in education, as well as establish criteria for AL appropriateness in teaching scientific communication. Finally, we specify artificial text generation outcomes by several LLMs as a response to the same prompt to identify language adequacy as an example of an error-driven teaching method. The authors of the article believe that LLM can be utilised as an auxiliary tool to help teach pragmatics of scientific discourse, since many pragmatics learning functions are aimed at performing meaningful tasks that mimic students’ real needs.
Background of the Problem
Attempts to relate language to logical analysis had been made long before the emergence of LLMs. The origin can be attributed to Aristotle’s work ‘First Analytics’ in which much of his ideas can be interpreted as the analysis of language [24]. His thoughts gave rise to a large number of works that can be aligned as ‘logic and language’ (Husserl, 1900–1901, Wittgenstein, 1921, Crane, 1928, Ingarden, 1931, Peirce, F. De Saussure, etc.).
In the twentieth century, logic, driven by many divisions, regenerated and acquired a mathematical form, on which a new approach to understanding linguistics was formed. At that time, artificial languages of symbolic logic were perceived by their founders as ideal compared to natural languages. Theories using mathematical reasoning were built on artificial languages, as many scientists were mathematical logicians. In symbolism, the main approach to language analysis was the correctness of word usage in order to avoid ambiguity while interpreting scientific facts [25]. However, the further practical use of artificial language as the spoken language reflected the impossibility of such use, due to the lack of statics in language in the process of its development [26].
The Prague and Copenhagen Schools, based on the synchronous analysis in linguistics introduced by F. de Saussure, described the formal relations in language at a certain point in time [27]. Later, the ‘elimination of ambiguity and uncertainty’ between formalised and natural language was also discussed by Church [28. P. 106]. The American school of descriptive structural linguistics organized by L. Bloomfield (1887–1949) described the grammatical structure of a language as a formal structure, due to which similarities and differences between artificial language and natural language became obvious [29; 30].
A considerable amount of more recent research in the framework of logical analysis of language provided rational for the description of an artificial language by means of natural language and by a logical apparatus [31–33], etc. to name just a few.
In philosophical logic, in his ‘Logic and Conversation’, Grice divided the logicians into formalists and informalists, and described the existence of a divergence between formal symbols and their analogy in natural language. Formalists, believing that ideal language was the basis for the clear expression of scientific thought, described natural language by a system of formulae, and they considered the presence of elements that cannot be described, even by more complex formulae, in natural language as ‘undesirable excrescences’ that had no place in the reliable expression of scientific thought [34]. Their opponents rejected even the possibility of measuring the adequacy of language and its ability to produce clear definitions just to please scientific thought. This simplified and symbolized logic of the formalists should not displace the logic of natural languages, but should reinforce it.
It is in these broad strokes of generalisation, that the authors of this article see the difference between approaches to describing natural and artificial languages. To be more precise, natural language is described within the framework of linguistics, where the primary importance is assigned to the study of speech units and language units, whereas in artificial language, data structures, algorithms, and logic are considered. In other words, artificial language is limited to the creation of specific models based on a specific set of data within the framework of mathematical logic. As defined by the Encyclopaedia of Database Systems, “a language model assigns a probability to a piece of unseen text, based on some training data” [35].
Rehash of a Hundred-Year-Old Idea
In one of his implicatures, in ‘Logic and Conversation’, Grice gives an example of a dialogue between two colleagues about a mutual friend who works in a bank: “A asks B how C is getting on in his job, and B replies, Oh quite well, I think; he likes his colleagues, ~no! he hasn’t been to prison yet. At this point, A might well inquire what B was implying, what he was suggesting, or even what he meant by saying that C had not yet been to prison. The answer might be any one of such things as that C is the sort of person likely to yield to the temptation provided by his occupation, that C’s colleagues are really very unpleasant and treacherous people, and so forth. It might, of course, be quite unnecessary for A to make such an inquiry of B, the answer to it being, in the context, clear in advance” [36. P. 43]. In a question asked in natural language, we can ask for additional information to clarify the communicative intention of the speaker, or perceive it from the context.
A similar experiment was conducted in Apple company to test the accuracy of a number of large language models. A sentence was added to the mathematical task’s condition hypothetically having a similar meaning, but not affecting the logic of reasoning in order to get the correct answer. The results demonstrated a significant decrease in the performance of the tested models, as this addition violated the logic of phrase construction and did not correlate with the pattern embedded in the model itself. The authors’ conclusion reflected the idea that the models do not understand either linguistic or mathematical categories, but operate with data ‘without truly understanding their meaning’ [37. P. 10].
Participants and Procedure
Participants of the study were the 1st-year PhD students at MIPT with the level of English not lower than B2 according to the CEFR levels. No selection criteria were set to the gender, age, and field of study; however, the prior theoretical knowledge of the learners should include the understanding of the academic style. The general requirements for writing a text were given as a theoretical background to the students (see the Supplement). The emphasis was put, in particular, on the understanding of the appropriacy of the text, which is based on lexico-semantic, lexico-grammatical, and stylistical levels.
Research Methodology
Recognizing the existing contradictions and accepting the limitations of LLMs, we cannot deny that they were created to solve specific problems, in other words, they were based on a pragmatic approach. Moving to teaching pragmatics of scientific discourse, it is necessary to define what do we consider by appropriateness by which a text can be evaluated. In contemporary literature both human and machine evaluations and opinions may be perceived as equally trustworthy by learners [38. P. 1766–1776]. However, text appropriateness in terms of communication remains an open question [37].
To define the appropriateness of the text generated by LLMs, we need clear and conventional criteria which will correspond to proposed methodology. Taking into consideration that any text assessment is subjective, we see stylistic level as the most relevant for such a comparison. Stylistic peculiarities which include grammar and lexical aspects are well-established in scientific literature and may be considered as a standard for text appropriateness measurement [39; 40].
This research does not embrace the whole IMRAD structure and the paragraph length since a generated text is analyzed for stylistic appropriateness without considering the stage of planning. For LLM editor review assessment we will focus on the following criteria: logical coherence; the structure of the sentence; grammatical adequacy and lexis. This choice stems from the lack of substantial disagreement in science regarding academic text structure [41].
Coherence includes conventional logical indicators and connective elements. Clarity is the most significant prerequisite for good academic writing therefore the standard model of the English sentence (Subject-Predicate-Object) presupposes close connection of these structural elements without any insertions to avoid ambiguity. The 3d person singular and other impersonal components contribute to academic discourse objectiveness. On the lexical level formal writing is marked by colloquial expressions, slang, irrelevant abbreviations and the majority of phrasal verbs [41–43]. However, some researchers encourage avoiding professional jargon and pretentious vocabulary in scientific works to enhance understanding [44; 45], while others urge to the formal style usage [46; 47]. After selecting criteria for text appropriateness, we created the following algorithm for LLM efficiency testing which is presented in Figure 1.
For comparative analysis the students were suggested formulating the main content elements of the Introduction section, which included Purpose, Objectives, Research question, Key message, Gap, Hook, Hypothesis, Practical value of the research. Then, as a result of further discussion, a prompt was written stating guidelines for further text generation:
Figure 1. The proposed algorithm for LLMs efficiency testing
Source: compiled by Anna V. Zabolotskikh, Elena E. Sokolova, Daria V. Tavberidze
Please proofread and edit the following text for language issues and stylistic inconsistencies. Focus on language; don’t try to improve the argumentation or add and remove content unless doing so is necessary to resolve language and style issues. The text should be written in American English, in 7th edition APA Style. Present the edited text followed by a list of all the changes you made, with clear explanations of why you made each change.
An important condition of the query was the justification for the proposed changes. We chose three models based on HELM leaderboard, where the main win rates together with accuracy, and efficiency of models’ performance across different scenarios and metrics were presented. Two of them, GPT– 4o and Gemini Ultra are closed models and Llama 3.1 (70 B) is an open model.
Three LLMs were fed with PhD students’ “human-written” abstracts and a proposed prompt.
After receiving a response to the inquiry from LLMs, the PhD students analyzed the adjustments proposed by LLMs in accordance with the text appropriateness using guidelines.
Task 2. Evaluate the generated text according to the guidelines and characterize if the generated text corresponds to academic standards.
Task 3. Select or reject the proposed improvements
Task 4. Reorganize the text according to the identified errors.
Further we show the examples of errors identified by LLMs in the human-written abstracts written by PhD students.
Findings and Discussion
To generate material for error critical analysis we run human-written texts on LLM models. LLMs’ inconsistency was apparent to the learners and caused constructive criticism. Further we provide the examples of errors detected by different LLMs in human-written texts and students’ criticism after their comparative analysis.
Example I. Source human-written text: With the recent development of superconducting technology, researchers began using artificial superconducting quantum two-level systems (qubits) as an environment for quantum optics experiments. One of the experiments is quantum wave mixing. In a nutshell, when two coherent microwaves with frequencies close to the qubit resonance get mixed by it, the spectrum of the scattered field appears to be linear with amplitudes of the peaks corresponding to the photon number statistics in the incident coherent states of the signals. Provided that two arbitrary quantum states of light get mixed by a qubit, the resulting spectrum might demonstrate a similar correspondence to the photon number statistics of the scattered signals. Suppose we mix a superposition of a photon and a vacuum with a coherent microwave by a qubit. Will the amplitudes of the peaks in the resulting spectrum reveal the properties of the quantum superposition? Spectra in the experiment of mixing a superposition with a coherent microwave clearly correspond to the prepared superposition properties. However, we could not eliminate all spurious effects. Moreover, there is no accurate theoretical description for the scattering spectra in this particular experiment yet.
LLMs generated results:
Gemini Ultra: Changed “Provided that” to “When” for simplicity and clarity. Changed “get mixed” to “are mixed” for grammatical accuracy.
GPT 4-o: Changed “Provided that” to “If”: This substitution for a simpler conditional phrase improves readability and clarity.
Llama 3.1 (70B): Changed “Provided that” to “Given that” for a more direct and natural-sounding transition.
Changed “In a nutshell” → “In essence” (idiomatic expression) I replaced this with a more formal idiom, “in essence,” to maintain a consistent tone.
Changed “Two coherent microwaves with frequencies close to the qubit resonance get mixed by it” → “two coherent microwaves with frequencies close to the qubit’s resonance interact” (subject-verb agreement and possessive adjective) I changed the verb from “get mixed” to “interact” and added a possessive adjective (“qubit’s”) to make the sentence more specific.
Students’ constructive criticism: [As seen] in Figure 1, electronic tools assess the conjunction provided that in a different way, which makes the writing strategy not clear. The question for reflection, which arises: should academic writing be marked by formalized pretentious vocabulary or, vice versa, prosaic words to enhance better understanding? The genitive (possessive) is not generally used with non-human subjects, however in some cases the genitive is used with inanimate things for specificity.
Gemini Ultra: “Spectra in the experiment” to “Spectra from the experiment”: Changed the preposition to “from” to indicate the source of the spectra (see Figure 4).
GPT 4-o: “Spectra in the experiment of mixing a superposition” to “The spectra from the experiment mixing a superposition”: This rephrasing improves sentence structure and clarity (see Figure 2).
Figure 2. Chat GPT 4: corrections (answer to the prompt)
Source: compiled by Anna V. Zabolotskikh, Elena E. Sokolova, Daria V. Tavberidze
“researchers began using” → “researchers have begun using”: Changed “began” to “have begun” for grammatical correctness, indicating the ongoing relevance of the action.
“correspond to the prepared superposition properties” to “correspond to the properties of the prepared superposition”: This adjustment enhances clarity and formality.
Llama 3.1 (70B): “Spectra in the experiment” to “The experiment’s results show” to make the sentence more clear and concise, and to avoid using the phrase “Spectra in the experiment” which sounds somewhat awkward (see Figure 3).
Provided that two arbitrary quantum states of light get mixed by a qubit” → “Provided that two arbitrary quantum states of light are mixed by a qubit” (added a verb to make the sentence grammatically correct).
Students’ constructive criticism: [I found that] when detecting the same error in the original version, Gemini Ultra, GPT 4-o and Llama 3.1 (70B) demonstrated different levels of improvement, from a preposition change to rephrasing the entire structure with the subject replacement. Some tools suggested the Perfect grammar form instead of the Past Simple for currency. The grammatical structure to get + Participle II was replaced with a similar form to be + Participle II, the difference between these structures (the starting point of an action vs. a state) being not recognized. The noun + of + noun construction was suggested as a priority instead of “locked” noun +noun (GPT 4-o) for clarity.
Our comment: Comparative analysis provoked students’ reflection on the differences that make sense when selecting the right grammar form in a scientific text. With such an approach grammar acquisition starts from analyzing the context with potential errors after which the grammar rule is inferred from practical examples.
Figure 3. Llama 3.1 (70B): corrections (answer to the prompt)
Source: compiled by Anna V. Zabolotskikh, Elena E. Sokolova, Daria V. Tavberidze
Figure 4. Gemini Ultra: corrections (answer to the prompt)
Source: compiled by Anna V. Zabolotskikh, Elena E. Sokolova, Daria V. Tavberidze
Example II. Source human-written text: The Standard Model proved to be a reliable model but there are things it can’t explain, like dark matter. So, we surely know that there are more that we know, we need to extend our models or find new ones. My research question is “Is it possible to find new particles in future experiments?”. First of all, I’m interested in this question because revealing new frontiers of physics always brings new possibilities for mankind. A lot of new theoretical research seemed to be inapplicable but history proved otherwise. Scientists of the future will use our discoveries to invent technologies we would consider impossible today. Second, experiments for such research use the most advanced technologies of our time, most of them are invented specifically for this experiment. This means that new types of magnets or conductors or whatever designed for experiment could be used by engineers to make something useful for society. For example, the World Wide Web was invented by engineers in CERN.
To speak more specifically on my thesis, in the near future, 10–20 years, a lot of experiments designed specifically for finding new physics are going to start collecting data. Estimating the number of events in these experiments for a variety of models will let us understand if these models should be investigated. We surely should work on models we could prove. Also, predictions on models could save us time after these experiments present data, we would already know what we want to see there.
LLMs generated results:
GPT 4-o: “there are things it can’t explain” → “there are phenomena it cannot explain”: Changed “things” to “phenomena” to use more precise language. Replaced “can’t” with “cannot” for formality.
“revealing new frontiers of physics always brings new possibilities for mankind” → “revealing new frontiers of physics always brings new possibilities for humanity”: “Mankind” was changed to “humanity” for inclusiveness and modern language use.
“We surely should work on models we could prove.” → “We should focus on models that we can validate.”: Changed “surely” to “should” for a more assertive tone and rephrased for clarity.
“Scientists of the future” → “Future scientists”: Simplified for conciseness.
Gemini Ultra:
“but there are things it can’t explain” → “however, there are aspects it cannot explain” Explanation: “However” provides a smoother transition, and “aspects” is more formal than “things.” “Cannot” is preferred over “can’t” in academic writing.
“to invent technologies we would consider impossible today” → “to invent technologies that we would consider impossible today” Explanation: “That” is often preferred for restrictive clauses in academic writing.
Llama 3.1 (70B):
Changed “proved to be a reliable model” to “has proven to be a reliable framework.” This is a slight rewording to improve the clarity and concision of the sentence.
Students’ criticism: The suggested changes are focused on the substitution of words for more formal equivalents (GPT 4-o), the correction of linking words and the addition of restrictive relative pronouns (Gemini Ultra), the grammatical design of the sentence (Present Perfect to update the context instead of the factual Past Simple). Evaluative words are replaced by the modal verb should. A noun-phrase was suggested as preferable to the corresponding of-phrase (GPT 4-o). Long strings of nouns and adjectives are generally only used if they are names of pieces of equipment or methods.
Our comment: the students assess the register of the improved text on the lexical and grammatical levels.
Example III: Source human-written text: In this thesis, the nature of the underlying theory will be assumed to be string/M theory. This is because of the following reasons. At present, string theory is the only known consistent theory of quantum gravity, at least at the perturbative level. But for the purposes of this thesis, more importantly, it is the only known ultra-violet complete theory which can naturally give rise to effective four dimensional theories of particle physics with the most important features of the Standard Model, namely, non-abelian gauge fields and chiral fermions. In addition, string theory can address all open questions in particle physics and cosmology within its framework and hopefully, solve them. Therefore, it seems reasonable to assume the existence of such a theory providing an underlying microscopic theoretical structure to our universe.With the recent development of superconducting technology, researchers began using artificial superconducting quantum two-level systems (qubits) as an environment for quantum optics experiments. One of the experiments is quantum wave mixing. In a nutshell, when two coherent microwaves with frequencies close to the qubit resonance get mixed by it, the spectrum of the scattered field appears to be linear with amplitudes of the peaks corresponding to the photon number statistics in the incident coherent states of the signals. Provided that two arbitrary quantum states of light get mixed by a qubit, the resulting spectrum might demonstrate a similar correspondence to the photon number statistics of the scattered signals. Suppose we mix a superposition of a photon and a vacuum with a coherent microwave by a qubit. Will the amplitudes of the peaks in the resulting spectrum reveal the properties of the quantum superposition? Spectra in the experiment of mixing a superposition with a coherent microwave clearly correspond to the prepared superposition properties. However, we could not eliminate all spurious effects. Moreover, there is no accurate theoretical description for the scattering spectra in this particular experiment yet.
LLMs generated results:
GPT 4-o: “But for the purposes of this thesis, more importantly,” → “More importantly,”: Removed “But for the purposes of this thesis” as it was redundant and detracted from the sentence’s clarity.
Removed the comma before “and hopefully, solve them.”: Removed the comma for smoother flow and correct punctuation.
Gemini Ultra: “But for the purposes of this thesis, more importantly, it is the only known ultra-violet complete theory” → “More importantly, for the purposes of this thesis, it is the only known ultraviolet complete theory” Explanation: Removed “But” for a smoother transition and changed “ultra-violet” to “ultraviolet” for consistency with standard terminology.
“and hopefully, solve them” → “and potentially solve them”
Explanation: “Potentially” is more formal and objective than “hopefully,” which is more subjective.
Llama 3.1 (70B): “more importantly, it is the only known ultra-violet complete theory” → “Moreover, for the purposes of this thesis, string theory is the only known ultra-violet complete theory” Explanation: Added the word “Moreover” to introduce the additional reason, and clarified the purpose of the thesis using the phrase “for the purposes of this thesis”.
“which can naturally give rise to effective four dimensional theories” → “that can naturally give rise to effective four-dimensional theories” Explanation: Changed “which” to “that” to maintain consistency in using the correct pronoun for a clauses, and added a hyphen to “four-dimensional” to adhere to APA style.
“Suppose we mix a superposition of a photon and a vacuum with a coherent microwave by a qubit” → “Suppose we mix a superposition of a photon and a vacuum with a coherent microwave via a qubit” (preposition) I replaced “by” with “via” to use a more formal and precise preposition.
“Do the amplitudes of the peaks in the resulting spectrum reveal the properties of the quantum superposition” (original text added a question mark at the end of this sentence) Since this is the last question of the paragraph, it would be more conventional to end the paragraph with a period.
“there is no accurate theoretical description for the scattering spectra in this particular experiment yet” → “there is currently no accurate theoretical description for the scattering spectra in this particular experiment” (verb tense) I changed the verb tense from “is” to “is currently” to make the sentence more consistent with the surrounding text.
Students’ criticism: In terms of the text’s logical coherence, in one and the same structure some phrases can be considered redundant (GPT 4-o), there may be a replacement of only one conjunction (Gemini Ultra) or an introductory word addition (Llama 3.1 (70B)). Moreover, in some cases, hyphenated and solid spelling of words (Gemini Ultra, Llama 3.1 (70B)) may be noted. Punctuation changes (GPT 4-o) or more formal lexical substitution may take place in the same structures for greater clarity of presentation (Gemini Ultra). Prepositions cannot always be used interchangeably: ‘by’ often identifies the ‘how’, and can suggest cause and effect’; ‘via’ (and ‘through’ and ‘by means of’) may not imply causality and may indicate that what comes afterwards is one option or a step among many (Llama). The future tense is changed for the present simple for factual, not prediction usage to avoid ambiguity. Words, belonging to informal register (yet), should be avoided.
Our comment: students evaluate cohesive structure of generated text in terms of redundancy of linking elements and spelling.
Example IV. Source human-written text: Provided that sequence read counts at heterozygous genomic loci in the enriched sequencing experiments reflect the genomic copy numbers of alternative alleles, It is possible to estimate the Background imbalance of the allelic copy numbers from the enriched sequencing data alone. Research question: Is it possible to infer the genome-wide profiles of copy number imbalance in the alleles from enriched sequencing data? Key message: The relative background allelic dosage profiles can be reconstructed directly from heterozygous variant calls in enriched sequencing data and be used as a baseline in statistical evaluation of allele-specific transcription factor binding and chromatin accessibility.
LLMs generated results:
GPT 4-o:
“Background imbalance” → “background imbalance”: Changed “Background” to “background” to maintain consistency with standard capitalization for terms unless they begin a sentence.
“of allele-specific transcription factor binding and chromatin accessibility.” → “of allele-specific transcription factor binding and chromatin accessibility.”: Added “the” for grammatical correctness, creating “the statistical evaluation.”
Gemini Ultra:
Capitalized “Background” in the hypothesis and key message: In APA style, proper nouns and adjectives derived from proper nouns are capitalized.
Changed “allelic copy numbers” to “allelic dosage” in the key message: The term “allelic dosage” is more commonly used in this context to refer to the relative abundance of different alleles.
Added “the” before “statistical evaluation”: Change: “in statistical evaluation” to “in the statistical evaluation.” Explanation: Adding “the” clarifies that it refers to a specific evaluation, enhancing the precision of the statement.
Llama 3.1 (70B):
“It is possible to estimate the Background imbalance” → “it is possible to estimate the background imbalance” I removed the capitalization of “background” since it’s not a proper noun. I also changed “Background” to “background” for consistency.
“of the allelic copy numbers from the enriched sequencing data alone” → “of allelic copy numbers from the enriched sequencing data alone” I removed “the” before “allelic copy numbers” since it’s not necessary for clarity.
“in statistical evaluation of allele-specific transcription factor binding and chromatin accessibility” → “in statistical evaluations of allele-specific transcription factor binding and chromatin accessibility” I changed “evaluation” to “evaluations” since it’s a plural noun and “evaluation” would be incorrect in this context.
Students’ criticism: There is a lack of consistency regarding the standards for capitalizing nouns (GPT 4-o, Gemini Ultra). At the grammatical level, it can be noted that the article is added to abstract nouns for accuracy (GPT 4-o) or the singular form of the noun is replaced by the plural (Llama 3.1 (70B)). In addition, in some cases, more appropriate key terms (Gemini Ultra) may be suggested. It may be noted that there is a replacement of terminology, as well as a shortening of the sentence by replacing the construction with an impersonal It with a modal verb (Llama 3.1 (70B)). One and the same abstract noun was corrected differently and changed into the plural form “evaluations” (Llama 3.1 (70B) or specified with the definite article (Gemini Ultra). The plural form would focus attention on the result whereas the singular form emphasises the process.
Our comment: students compare the generated outcome with their original texts in terms of grammar: articles, plural/singular forms of nouns and lexis: the key terminology.
Discussion
In order to lay the foundations for a discussion around the argument that the teaching process may benefit from the adoption of LLMs, in this article we analyse error-driven learning methodology presented to the PhD students.
Error-driven learning proposes that students compared the human-written abstract with the generated one and, if there is a discrepancy, they analyzed each detected misconception using guidelines. The artificial error climate created a learning environment, in which the PhD students practiced critical and even skeptical analysis towards the generated results to evaluate information based on verifiability and falsifiability.
The discrepancies obtained from three LLMs can be used in teaching academic writing on the lexico-semantic, lexico-grammatical and stylistic levels. While analyzing generated outcomes the PhD students realized that an artificial language corrector is not a universal tool, as it reflects the source of information embedded in them. Running three LLMs on one and the same prompt can generate very different answers. Thus, using error-driven approach provides the learners not only with factual knowledge but stimulates them to reason with this knowledge.
Conceptually, error-driven learning as a reflection-based method in LLMs studies is aimed at significant improvements in various reasoning tasks [48] through addressing the shortcomings and leveraging the pre-designed prompts. This study suggests a modified teaching approach based on the same principle. However, in our case, artificial systematic errors produced by machines and introduced to human learners serve as an activation-based mechanism for observing and explaining a wide range of linguistic phenomena on lexico-semantic, lexico-grammatical and semantic levels. That said, while the focus of error-driven learning is computational modelling, its principles may be applicable across many other areas including teaching foreign languages.
Conclusions
In this article, we reflect on the relationship between humans and machines in education. LLMs’ ability to produce different generated outcomes may be applied as error-driven methodology for academic writing teaching. The main purpose of this investigation was to argue for a different approach to error-driven learning research, making it less prescriptive and more enabling self-reflection and self-correction. With this in mind, we argue for a shift of focus in the research of academic writing teaching on critical analysis and editing to make it more sensitive to visualizing errors. In error-driven learning machines enhance human work instead of replacing it. Interaction with artificial tools creates a sense of competition and stimulates further inquiry into where the truth actually lies. This seems to be a good example of how inconsistency and ambiguity may be turned to advantage.
Supplement. Guidelines for evaluation the text appropriacy
Lexico-Semantic Level | Phrasal verbs - Are there direct verbs that could be used to increase clarity and precision? Synonyms appropriateness - Are the synonyms used appropriatly? - Do the meanings of synonyms maintain the equal emotional stress? Spelling (Hyphenated/Solid) - Have the spelling of hyphenated versus solid compound words been checked? - Are all compound words written correctly according to style guidelines? |
Lexico-Grammatical Level | Definiteness/Indefiniteness in Article Usage - Are articles (a, an, the) used correctly to denote definiteness or indefiniteness? - Have noun phrases been reviewed to ensure proper article usage? Nouns Capitalization - Have capitalized proper nouns and other significant terms been used correctly? - Is there consistency with capitalization of terms used throughout the document? Presence of Linking Words - Are logical indicators present to guide the reader through the argument? Sentence Structure - Does the sentence structure predominantly follow the Subject-Predicate-Object sequence? |
Stylistics | Register - Is the writing style consistent in maintaining an formal register? - Has ambiguity or vagueness in word choice been avoided? Colloquial expressions - Are there any colloquial expressions present? - Has informal language been replaced with formal alternatives? Sentence style - Are the sentences varied in structure while maintaining clarity? |
Source: compiled by Anna V. Zabolotskikh, Elena E. Sokolova, Daria V. Tavberidze
About the authors
Anna V. Zabolotskikh
RUDN University
Author for correspondence.
Email: zabolotskikh-av@rudn.ru
ORCID iD: 0000-0002-5253-2733
SPIN-code: 6520-1445
Scopus Author ID: 57217873153
ResearcherId: AAB-4687-2019
Senior lecturer, Department of Foreign Languages, Faculty of Humanities and Social Sciences, Federal State Autonomous Educational Institution of Higher Education
6 Miklukho-Maklaya str., Moscow, Russian Federation, 117198Elena E. Sokolova
Moscow Institute of Physics and Technology
Email: sokolova.ee@mipt.ru
ORCID iD: 0000-0002-1467-3455
SPIN-code: 2013-1826
Scopus Author ID: 56579818200
ResearcherId: G-4265-2019
PhD in Philology, Associate Professor, Department of Foreign Languages
9 Institutsky lane, Moscow region, Dolgoprudny, Russian Federation, 141701Darya V. Tavberidze
RUDN University
Email: tavberidze-dv@rudn.ru
ORCID iD: 0000-0002-2727-6803
SPIN-code: 3183-1668
Scopus Author ID: 57195529929
ResearcherId: ABW-5226-2022
PhD in Philosophy, Associate Professor, Department of Foreign Languages, Faculty of Humanities and Social Sciences
6 Miklukho-Maklaya str., Moscow, Russian Federation, 117198References
- Brown, T., Mann, B., Ryder, N., et al. (2020). Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems, 33, 1877-1901.
- Koraishi, O. (2023). Teaching English in the Age of AI: Embracing ChatGPT to Optimize EFL Materials and Assessment. Language Education and Technology, 3(1), 55-72.
- Binu, S. K., & Shanthi, C. (2024). Clinical insight: Comparative Analysis of Deep Learning Models for Disease Prediction across Multifaceted Datasets. In: 2024 Third International Conference on Distributed Computing and Electrical Circuits and Electronics (ICDCECE) (pp. 1-8). IEEE.
- Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., et al. (2017). Attention is All You Need. Advances in Neural Information Processing Systems, 30.
- Tran, K.T., Dao, D., Nguyen, M.D., Pham, Q.V., O’Sullivan, B., & Nguyen, H.D. (2025). Multi-Agent Collaboration Mechanisms: A Survey of LLMS. arXiv, preprint: 2501.06322.
- Song, C., & Song, Y. (2023). Enhancing Academic Writing Skills and Motivation: Assessing the Efficacy of ChatGPT in AI-Assisted Language Learning for EFL Students. Frontiers in Psychology, 14, 1260843. https://doi.org/10.3389/fpsyg.2023.1260843 EDN: YPEEQR
- Schmohl, T., Watanabe, A., Fröhlich, N., & Herzberg, D. (2020). How Artificial Intelligence can Improve the Academic Writing of Students. In: Proceedings of 10th International Conference the Future of Education (pp. 168-171).
- Khabib, S. (2022). Introducing Artificial Intelligence (AI)-based Digital Writing Assistants for Teachers in Writing Scientific Articles. Teaching English as a Foreign Language Journal, 1(2), 114-124. https://doi.org/10.12928/tefl.v1i2.249 EDN: PFNAFP
- Storey, V.A., & Cunningham, B. (2024). Artificial Intelligence, Educational Leadership, Knowledge Production, and Transfer: A New Infrastructure. Education Leadership Review, 25(2), 641-655. https://doi.org/10.46932/sfjdv4n2-001
- Anson, C.M., & Straume, I. (2022). Amazement and Trepidation: Implications of AI-based Natural Language Production for the Teaching of Writing. Journal of Academic Writing, 12(1), 1-9. https://doi.org/10.18552/joaw.v12i1.820 EDN: ИГЗЕКН
- Floridi, L., & Chiriatti, M. (2020). GPT-3: Its Nature, Scope, Limits, and Consequences. Minds and Machines, 30(4), 681-694. https://doi.org/10.1007/s11023-020-09548-1 EDN: УУЛУЛА
- Safire, W. (2008). Safire’s Political Dictionary. Oxford University Press.
- Bender, E. M., & Koller, A. (2020, July). Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 5185-5198).
- Gee, L., Rigutini, L., Ernandes, M., & Zugarini, A. (2023, December). Multi-Word Tokenization for Sequence Compression. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track (pp. 612-621). https://doi.org/10.18653/v1/2023.emnlp-industry.58
- Ji, Ziwei, Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., et al. (2023). Survey of Hallucination in Natural Language Generation. ACM Computing Surveys, 55(12), 1-38. https://doi.org/10.1145/3571700
- Peng, B., Narayanan, S., & Papadimitriou, C. (2024). On Limitations of the Transformer Architecture. Corr, Abs/2402.08164.
- Bang, J., Lee, J., Lee, G.G., & Chung, M. (2014). Pronunciation Variants Prediction Method to Detect Mispronunciations by Korean Learners of English. ACM Transactions on Asian Language Information Processing (TALIP), 13(4), 1-21. https://doi.org/10.1145/2629545
- Gu, Y., Dong, L., Wei, F., & Huang, M. (2023). Minillm: Knowledge Distillation of Large Language Models. Arxiv, preprint: 2306.08543.
- Lengo, N. (1995). What is an Error? English Teaching Forum, 33(3), 20-24.
- Lüdeling, A., Walter, M., Kroymann, E., & Adolphs, P. (2005). Multi-Level Error Annotation in Learner Corpora. In: Proceedings of Corpus Linguistics (Vol. 1. pp. 14-17).
- Liu, Y., & Reichle, E.D. (2015). An Evolutionary Algorithm for Error-Driven Learning via Reinforcement. arXiv preprint:1503.07609.
- Mollick, E.R., & Mollick, L. (2023). Using AI to Implement Effective Teaching Strategies in Classrooms: Five Strategies, Including Prompts. URL: https://teaching.uoregon.edu/ (accessed: 12.01.2025).
- Kuleto, V., Ilić, M., Dumangiu, M., Ranković, M., Martins, O.M., Păun, D., & Mihoreanu, L. (2021). Exploring Opportunities and Challenges of Artificial Intelligence and Machine Learning in Higher Education Institutions. Sustainability, 13(18), 10424. https://doi.org/10.3390/Su131810424 EDN: QRTPOF
- Berrón, M. (2020). Aristotle’s Politics I and the Method of Theanalytics. Rhizomata, 8(1), 83-106. https://doi.org/10.1515/Rhiz-2020-0006 EDN: SFWMBH
- Küng, L., Kröll, A. M., Ripken, B., & Walker, M. (1999). Impact of the Digital Revolution on the Media and Communications Industries. Javnost-The Public, 6(3), 29-47. https://doi.org/10.1080/13183222.1999.11008717
- De Saussure, F. (1916). Nature of the Linguistic Sign. In: Course on General Linguistics (pp. 65-70). New York : Philosophical Library.
- Siertsema, B. (1955). The Linguistic Sign: the Sign in Itself. In: A Study of Glossematics: Critical Survey of Its Fundamental Concepts (pp. 128-145). Dordrecht: Springer Netherlands.
- Church, A. (1951, July). The Need for Abstract Entities in Semantic Analysis. In: Proceedings of the American Academy of Arts and Sciences (Vol. 80. no. 1. pp. 100-112). American Academy of Arts & Sciences.
- Harris, Z. (1976). On A Theory of Language. The Journal of Philosophy, 73(10), 253-276. https://doi.org/10.2307/2025530
- Chomsky, N. (1957). Logical Structure in Language. Journal of The American Society for Information Science, 8(4), 284. https://doi.org/10.1002/asi.5090080406
- Abbot-Smith, K., & Tomasello, M. (2006). Exemplar-Learning and Schematization in a Usage-Based Account of Syntactic Acquisition. Linguistic Review, 23(3), 275-290. https://doi.org/10.1515/TLR.2006.011
- Abusch, D. (1993). The Scope of Indefinites. Natural Language Semantics, 2(2), 83-135. https://doi.org/10.1007/bf01250400 EDN: KQJLYA
- Carlson, G.N. (1989). On the Semantic Composition of English Generic Sentences. In: Properties, Types and Meaning: Vol. II: Semantic Issues (pp. 167-192). Dordrecht: Springer Netherlands.
- Grice, H.P. (1978). Further Notes on Logic and Conversation. Syntax and Semantics, 9, 113-127. https://doi.org/10.1163/9789004368873_006
- Hiemstra, D. (2018). Language models. In: Encyclopedia of Database Systems (pp. 2061-2065). Springer.
- Grice, H.P. (1975). Grice-Logic and Conversation. London: University College London Press.
- Mirzadeh, I., Alizadeh, K., Shahrokhi, H., Tuzel, O., Bengio, S., & Farajtabar, M. (2024). GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models. arXiv, preprint:2410.05229.
- Abendschein, B., Lin, X., Edwards, C., Edwards, A., & Rijhwani, V. (2024). Credibility and altered communication styles of AI graders in the classroom. Journal of Computer Assisted Learning, 40(4), 1766-1776. https://doi.org/10.1111/jcal.12979 EDN: GVGEBJ
- Galperin, I.R. (1973). About the Concepts of “Style” and “Stylistics”. Voprosy Jazykoznanija, 3, 14-25. (In Russ.). Гальперин И.Р. О понятиях «стиль» и «стилистика» // Вопросы языкознания. 1973. № 3. С. 14-25.
- Shpetnyi, C. I. (2022). School of Linguistic Stylistics in Russia and Vista of Its Advance in the Paradigm of Modern Cognitive Knowledge. Vestnik of Moscow State Linguistic University. Humanities, 8(863), 119-127. https://doi.org/10.52070/2542-2197_2022_8_863_119
- Bennett, K. (2009). English Academic Style Manuals: a Survey. Journal of English for Academic Purposes, 8(1), 43-54. https://doi.org/10.1016/j.jeap.2008.12.003
- Pagliawan, D.L. (2017). Feature Style for Academic and Scholarly Writing. Academic Journal of Interdisciplinary Studies, 6(2), 35-41. https://doi.org/10.1515/ajis-2017-0004
- Hayot, E. (2014). The Elements of Academic Style: Writing for the Humanities. Columbia University Press.
- Rose, D., Rose, M., Farrington, S., & Page, S. (2008). Scaffolding Academic Literacy with Indigenous Health Sciences Students: an Evaluative Study. Journal of English for Academic Purposes, 7(3), 165-179.
- Kneale, P. (2008). A Handbook for Teaching and Learning in Higher Education. Routledge. https://doi.org/10.1016/j.jeap.2008.05.004
- Jordan, R.R. (1997). English for Academic Purposes. Cambridge University Press.
- Willems, S., Bollen, H., Van Der Veen, J., Sterpin, E., Crijns, W., Nuyts, S., & Maes, F. (2021). Learning from Mistakes: an Error-Driven Mechanism to Improve Segmentation Performance Based on Expert Feedback. In: MICCAI Workshop on Distributed and Collaborative Learning (pp. 68-77). Cham: Springer.
Supplementary files
Source: compiled by Anna V. Zabolotskikh, Elena E. Sokolova, Daria V. Tavberidze
Source: compiled by Anna V. Zabolotskikh, Elena E. Sokolova, Daria V. Tavberidze
Source: compiled by Anna V. Zabolotskikh, Elena E. Sokolova, Daria V. Tavberidze
Source: compiled by Anna V. Zabolotskikh, Elena E. Sokolova, Daria V. Tavberidze











