AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

AI-generated feedback on writing: insights into efficacy and ENL student preference

Design
Quasi-experimental study · randomized at none level · 48 participants
Subject and task
Writing language
Academic English paragraph writing (300-word source-integrated paragraphs) by university English-as-a-new-language (ENL) learners; weekly draft-feedback-revise cycles
Exposure: 6 weeks
Population match
Secondary students: Different · University students: Direct
Attrition
Study 1: "The scores of one student who only completed the posttest were not included in the analysis." No other Study 1 attrition reported. Study 2: weekly survey completion varied from 32 to 41 of 43 (average 37.7); fewer students completed the survey in the final week.
Study tier
Study tier 3Quasi-experimental two-group pretest-posttest design; the paper says participants "were divided into two groups" with no randomization described and self-selected enrollment, so baseline equivalence is not guaranteed (pretest difference favored EG, d=0.520, p=0.078). The comparison arm is genuinely active (weekly 30-min one-on-one sessions with a trained, CRLA-certified human tutor), which strengthens interpretability of the contrast, but modality differs (face-to-face dialogue vs emailed written feedback). The headline result is a nonsignificant group-by-time interaction at N=48 (23 vs 25); at this sample size a null difference is weak evidence of equivalence between AI and human tutor feedback, and the interaction was near-threshold (p=0.085) with the descriptive gain favoring the human-tutor group (6.86 vs 5.848).
Assessment
Version 2 · AI: two passes agreed · Sep 18, 2026

Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.

University ENL students in an academic writing course got weekly feedback on 300-word paragraph drafts for six weeks, either from ChatGPT (GPT-4, via a teaching-assistant-run engineered prompt, emailed to students) or from a trained human tutor in weekly 30-minute one-on-one sessions. Both groups improved substantially from a diagnostic pretest to a final writing exam scored by two blinded-to-nothing-stated human raters on a 40-point rubric, and neither the group-by-time interaction nor the between-group contrasts were statistically significant, so AI feedback was neither better nor detectably worse than human tutor feedback at this small sample size. A separate group of 43 students who received both kinds of feedback split roughly evenly on which they preferred, rating both highly.

01Findings

What the study measured

Arms

ArmWhat it got
A · Experimental group: AI-generated feedback (ChatGPT/GPT-4) on weekly drafts
n = 23
Intervention · teacher supervised
ChatGPT · GPT-4 · Prompted pedagogy · feedback interface · homework · used Shortened Spring 2023 semester; six weekly cycles
B · Control group: human tutor feedback (weekly 30-min one-on-one session)
n = 25
Active comparison · teacher led
Weekly 30-minute one-on-one tutoring session with the same paid, trained English language tutor from the university's ENL Tutor Program, certified by the College Reading & Learning Association; same weekly write-feedback-revise-resubmit cycle as the AI arm.

Findings

A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.

Finding and conditionsResult
F1 · Writing score gain pretest (diagnostic writing test, week 1) to posttest (final writing exam, week 8): 2x2 mixed RM-ANOVA group-by-time interaction. Rubric: analytic, four areas (content, coherence, language use, sources and evidence), max 40, two independent raters (ICC 0.816-0.932). Assessed writing was produced under exam/diagnostic conditions WITHOUT the feedback tool: students never had direct access to GPT-4 at any point (a TA generated feedback and emailed it on weekly drafts, not on the tests).
Experimental group: AI-generated feedback (ChatGPT/GPT-4) on weekly drafts vs Control group: human tutor feedback (weekly 30-min one-on-one session)Measured without AIImmediate post-testSame contentResearcher-developed testPrespecification unclear
Null
F = 3.094, eta-p2 = 0.063; descriptive gains EG +5.848 (27.522 to 33.370) vs CG +6.86 (26.820 to 33.680), trend favoring the human-tutor CG
not_reported
0.085
Researcher/institution-developed analytic rubric scored independently by two experienced academic English instructors; excellent pretest inter-rater reliability (ICC 0.932 CG, 0.919 EG) and good posttest reliability (0.877 CG, 0.816 EG). Posttest CG distribution non-normal (Shapiro-Wilk 0.833, p < 0.001); Greenhouse-Geisser correction applied. Rater blinding to group not stated.
F2 · Posttest (T2, final writing exam, week 8) between-group contrast: independent samples t-test of EG vs CG writing scores (same 40-point rubric). Writing produced without the feedback tool available.
Experimental group: AI-generated feedback (ChatGPT/GPT-4) on weekly drafts vs Control group: human tutor feedback (weekly 30-min one-on-one session)Measured without AIImmediate post-testSame contentResearcher-developed testPrespecification unclear
Null
EG M=33.370 (SD 1.908) vs CG M=33.680 (SD 2.066); t(46) = -0.539, Cohen's d = -0.156
not_reported
0.591
Same rubric/raters as F1. Baseline (T1) t-test also nonsignificant but near-threshold favoring EG (t(46)=1.801, p=0.078, d=0.520), so the groups were not perfectly equated at pretest despite the null.
F3 · Study 2 (separate sample, N=43, both feedback types received): weekly Qualtrics preference questionnaire over six weeks - eight 5-point Likert items in four AI-vs-human pairs (satisfaction, clarity, helpfulness, overall preference), a forced-choice item 9 ("which kind of feedback would you prefer?"), and an open-ended explanation.
Experimental group: AI-generated feedback (ChatGPT/GPT-4) on weekly drafts vs Control group: human tutor feedback (weekly 30-min one-on-one session)AI availability at assessment not reported — cannot carry a learning claimDuring interventionTransfer not assessedSelf-reportPrespecification unclear
Unclear
Descriptive near-even split: six-week average 18 students preferred human tutor feedback vs 19.667 preferred AI feedback; six-week Likert means slightly higher for human feedback in every pair (satisfaction H=4.277 vs AI=4.262; clarity H=4.319 vs AI=4.236; helpfulness H=4.305 vs AI=4.228; preference H=4.2 vs AI=4.126). No inferential test of the H-vs-AI item means is reported; the only t-tests reported compare preference subgroups (grouped by week-6 Q9 answer) on each item, all nonsignificant.
not_reported
not_reported
Researcher-developed questionnaire; no validation or reliability statistics reported. Weekly completion varied (32-41 of 43); the final week's lower completion may bias the forced-choice averages. Tutors in Study 2 were unaware students also received AI feedback. Direction left unclear per convention: the H-vs-AI comparison itself is descriptive only (per convention 6, direction requires an inferential test).
Limitations
No dedicated limitations section. Author-acknowledged constraints: students could not ask follow-up questions of the AI (access was deliberately limited; a TA generated and emailed the feedback), and GPT-4 is "not optimized for AWE purposes" (e.g., no text annotation). Additional reviewer-noted limitations: non-random assignment and self-selection; small N; posttest CG scores non-normal (Shapiro-Wilk p<0.001); rubric was researcher/institution-developed; feedback modality confounded with source (spoken interactive tutoring vs written emailed AI feedback); single site and single shortened semester; per-protocol handling of the one incomplete case.

Who was studied

Level
university
Ages
Study 1: 20-30; Study 2: 19-36
Country
not named; small liberal arts university in the Asia-Pacific region (authors affiliated with Brigham Young University-Hawaii)
Prior knowledge
At least CEFR B1 English proficiency per the institution's English Language Admission Test; enrolled in an academic reading and writing language course
Selection
Non-probability self-selection recruitment of 91 participants across both studies; voluntary consent, no academic credit offered

Methodological notes

Extracted from: https://link.springer.com/article/10.1186/s41239-023-00425-2 Access: Open-access full text (CC BY) read via the Springer Nature Link HTML page in a browser session (direct WebFetch/curl blocked by Springer's cookie/bot challenge). Full article body, funding, and competing-interests sections captured; Tables 1 and 3 read from their dedicated table pages (per-group Ns, means, SDs, t-tests). Table 2 (RM-ANOVA) statistics taken from the results text. Appendices A-C (survey items, example prompt/feedback) not retrieved. Sample-size note: Study 1 (learning outcomes): N=48 analyzed (EG/AI n=23, CG/human tutor n=25, per Tables 1 and 3). Study 2 (preference, different participants): N=43, with complete weekly questionnaire responses ranging 32-41 (mean 37.7) across the six weeks. 91 recruited across both studies.

01bRisk of bias

How much each result can be relied on

Risk of bias is judged per result, not per paper: the same study can report one outcome at low risk and another at high. Two reviewers assess each result independently against a written guide, and every domain judgment below carries its reasoning and the sentences it rests on.

Risk of bias is not the same as whether the numbers work. A trial can be well conducted and still print a standard deviation that is arithmetically impossible, or a table whose sign contradicts its own text. Risk of bias has no question for that, so where it was found it is recorded separately below, under Source integrity.

These words are not quality scores. On the RoB 2 scale, High risk of bias is the worst rating a result can receive and Low risk of bias is the best — the opposite of how the same words read on some other scales. Levels are always written out in full here for that reason.

WRITING

Serious risk of bias · ROBINS-I · Two reviewers agreed on every domain

How the overall was reached: Worst-domain rule: D1 (confounding) is Serious because the allocation mechanism for this quasi-experimental design is entirely unreported, the groups differed at baseline by a moderate margin (d = 0.520, p = 0.078) despite the authors' equivalence framing, and no confounder or baseline adjustment was made beyond the implicit conditioning of the repeated-measures design. D4, D6, and D7 are Moderate (unmonitored contamination/adherence, no stated rater blinding, no preregistered analysis plan); D2, D3, and D5 are Low. The study is at serious risk of bias for the WRITING growth result, which is particularly consequential given the result is a null finding used to claim equivalence of AI and human feedback.

DomainJudgment and reasoning
D1 · Confounding
Serious risk of bias
Non-randomized allocation with no description of how the 48 students were "divided into two groups" (mechanism, unit of allocation, and whether groups were intact classes are all unreported), so confounding by proficiency, motivation, or class membership cannot be ruled out. There was a non-trivial baseline imbalance favoring the EG (Table 3: T1 EG M = 27.522 vs CG M = 26.820; t = 1.801, df = 46, p = 0.078, Cohen's d = 0.520), and the analysis made no adjustment for baseline or any other confounder (RM-ANOVA with no covariates; the authors treat the nonsignificant t-test as evidence of equivalence). The repeated-measures interaction conditions on baseline level via within-subject change, which partially mitigates the level imbalance, but with posttest means near 33.5 on a 40-point rubric a higher-scoring EG had less room to grow, and confounders of growth rate remain uncontrolled. Prognostically important baseline difference plus no adjustment and an unknown allocation mechanism warrants Serious.
They were divided into two groups: a control group (CG) that received feedback on their assignments from a human tutor, and an experimental group (EG) that was given feedback generated by AI (GPT-4).
Lastly, two independent samples t-tests were conducted to explore potential differences in means at T1 and T2 to see if there were any significant differences in proficiency levels between the two groups before and after the treatment.
D2 · Selection of participants
Low risk of bias
Participants were recruited by self-selection from one course before the intervention began, and follow-up starts with the intervention (pretest in week 1, weekly feedback cycle thereafter, posttest in week 8). Selection into the study does not appear to be related jointly to intervention and outcome; self-selection sampling threatens generalizability rather than ROBINS-I selection bias. The exclusion of one student without a pretest is handled under missing data (D5).
A non-probability self-selection method was used to recruit 91 participants who were ENL students enrolled in an academic reading and writing language course.
In study 1, students completed the pretest in the first week of the semester. For the next six weeks these students completed a weekly writing assignment.
D3 · Classification of interventions
Low risk of bias
Intervention status (GPT-4-generated feedback vs human tutor feedback) was defined at the start, applied group-wide, and recorded prospectively by the researchers, who themselves generated and emailed the AI feedback; misclassification of intervention is implausible.
For each week students wrote a 300 word paragraph on a provided topic related to the material in class, received feedback either from a human tutor (CG) or AI (EG), revised, and submitted a final draft.
A teaching assistant utilized GPT-4 and the prompt described above to generate feedback for each student.
D4 · Deviations from intended interventions
Moderate risk of bias
For the effect of assignment, delivery differences (CG: weekly 30-minute face-to-face session with a certified tutor; EG: written feedback emailed within two days) are part of the interventions as defined rather than deviations. However, the paper reports nothing on adherence (revision was optional), on whether CG students independently used ChatGPT or other GenAI tools during Spring 2023 (contamination was plausible given the study's timing), or on co-interventions outside the feedback cycle. Absent any reported monitoring, some imbalance in unintended exposures cannot be excluded, but there is no positive evidence of important deviations; analysis was by assigned group.
Participants in the CG held one 30-min one-on-one tutoring session per week with the same tutor for the duration of the study.
Feedback was emailed to students within two working days to ensure students had ample time to read the feedback and incorporate it into their revisions, should they desire to.
D5 · Missing data
Low risk of bias
Outcome data were nearly complete: all 48 analyzed participants (EG n = 23, CG n = 25 in Tables 1-3) had both pretest and posttest scores, and only one additional student, who lacked a pretest, was excluded. A single exclusion of this kind is very unlikely to bias the group x time comparison.
Missing data, outliers, and possible statistical assumption violations were examined. The scores of one student who only completed the posttest were not included in the analysis.
D6 · Measurement of outcomes
Moderate risk of bias
The outcome was human-rated writing quality scored with an analytic rubric by two experienced instructors, with good-to-excellent inter-rater reliability (ICCs 0.816-0.932), and the pre/post essays were produced without access to the tool, so the measure itself is comparable across groups. However, the paper never states that raters were blinded to group assignment (identifying information was removed only from EG drafts sent to GPT-4, not from the rated tests) or to timepoint, and subjective human scoring by assessors potentially aware of condition leaves room for modest bias, though there is no indication scoring could have been differentially influenced.
Both pre- and posttests were independently rated by two experienced academic English language instructors. An analytic rubric assessing four key writing areas, namely content, coherence, language use, and sources and evidence, was used to assess students' writing in the pre- and post test.
These values indicated excellent inter-rater reliability for the CG and EG pretests (0.932, p < 0.001; 0.919, p < 0.001), as well as good reliability for the CG and EG posttests (0.877, p < 0.001; 816, p < 0.001).
D7 · Selection of the reported result
Moderate risk of bias
No preregistration, protocol, or analysis plan is mentioned anywhere in the article, so the analyses cannot be verified as pre-specified. Bias risk is tempered because there was a single writing outcome measured at two fixed timepoints, the RM-ANOVA plus baseline/posttest t-tests are the conventional analyses for this design, and all of them are reported (including the nonsignificant interaction); there is no sign of selective reporting from multiple measures, timepoints, or analyses, but the absence of a registered plan precludes a Low rating.
To explore the first research question in study 1, a repeated-measure analysis of variance (RM-ANOVA) was conducted using SPSS 28.
Using a general linear model, the time of the tests (pretest [T1] and posttest [T2]) served as a within-subjects variable, with the group serving as a between-subjects variable (experimental group or EG, and control group or CG).
02Funding

Funding and conflicts

Funding
None: no specific grant from any public, commercial, or not-for-profit funding agency
Vendor funded
No
Vendor
None identified
Notes
Authors declared no competing interests. Authors are university faculty (BYU-Hawaii, Florida State); no AI-industry ties disclosed.

Funding is shown on every study and never used to score it.

03Claims

Claims this study bears on

04Source

The source, as retrieved

Abstract

No abstract retrieved.

Where this record came from

SourceRetrievedIdentifier
seedSep 18, 2026link
first ingestion