AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance

Design
Randomized trial · randomized at student level · 117 participants
Subject and task
Writing language
English (L2) reading-and-writing task: read materials on three topics (AI, differentiated teaching, scaffolding teaching) and write, then revise, an essay envisioning the future of education in 2035, scored against a 25-point rubric.
Exposure: Single lab session: 2-hour reading and writing task (stage 1, no differentiated support) followed by a 1-hour revising task (stage 2, in which the group-specific support was available); post-test completed within one day.
Population match
Secondary students: Different · University students: Direct
Attrition
not_reported as such; analysis Ns vary by outcome (117 for essay scores, 114 for IMI, 97-107 for knowledge/transfer tests) with no explanation of missing data or exclusions in the arXiv text.
Study tier
Study tier 2Randomised experiment with an unusually informative set of comparison conditions (ChatGPT vs. human expert vs. writing-analytics checklist vs. no support), objective trace data, and blinded-rubric essay scoring with high inter-rater reliability on a subsample. However, it is a single ~3-hour lab session with small per-arm samples (24-35), researcher-developed knowledge tests, unexplained variation in analysis Ns and internally inconsistent group sizes across tables, and the headline "metacognitive laziness" claim rests partly on descriptive process-mining contrasts without inferential statistics. Null findings on knowledge gain and transfer are imprecise given the sample size.
Assessment
Version 2 · AI: two passes agreed · Sep 18, 2026

Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.

117 university students wrote an English essay in a 3-hour lab session and were randomly given one of four kinds of help for the 1-hour revision stage: a task-restricted ChatGPT-4, a human writing expert, an AI-powered checklist/writing-analytics toolkit, or nothing. The ChatGPT group improved their essays significantly more than every other group - including the human-expert group - but on tests of what they actually learned (knowledge of the topic and transfer to a new domain, taken after the session without the tool) they did no better than anyone else, and motivation did not differ. Behaviour traces showed the ChatGPT group's revising revolved around the chatbot, with relatively fewer self-checking (metacognitive) moves than the human-expert and checklist groups, which the authors call potential "metacognitive laziness". In short: the AI made the assisted product better without making the learner better.

01Findings

What the study measured

Arms

ArmWhat it got
A · ChatGPT group (AI group)
n = 35
Intervention · teacher independent
ChatGPT (embedded in study platform via OpenAI API) · GPT-4 · Guardrailed tutor · chat interface · in class · used July to September 2023
B · Human Expert group (HE group)
n = 25
Active comparison · teacher independent
One-on-one real-time chat (unrestricted conversations) with a human expert - a professional researcher, editor, and academic writing teacher who had taught academic writing for four semesters and knew the task - who received essays in real time via an "Ask Teacher Tool" during the revision stage.
C · Checklist group (CL group)
n = 27
Active comparison · teacher independent
On-demand feedback (button press) from a writing-analytics "Checklist Tools" kit: (1) basic writing tool for spelling/grammar (feedback generated by GPT-4 via API), (2) academic writing tool from an experienced teacher's materials, (3) originality tool flagging 7-word overlaps with the readings, (4) integration/elaboration tool using a GPT-based rhetorical classifier. Note: although labelled a comparison tool, parts of its feedback are GPT-4-generated.
D · Control group (CN group)
n = 30
Control · teacher independent
Same base learning environment (timer, planner, highlight/note-taking, search, dictionary tools) in both stages with no additional revision support; training video reminded learners to use task instructions and rubric to revise for a higher score.

Findings

A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.

Finding and conditionsResult
F1 · Essay score improvement (rubric-scored essay after revision minus before revision; 25-point researcher rubric)
ChatGPT group (AI group) vs Control group (CN group)With AI at assessment — assisted performance, not learning evidenceDuring interventionSame contentGraded artifactPrespecification unclear
Benefit
Mean difference = 1.970 points (omnibus ANOVA on improvements F=4.549, p=0.005, eta2=0.108; AI M=3.600 vs CN M=1.630)
95% CI [0.083, 3.858] (Tukey table reports CN-AI: -1.970 [-3.858, -0.083])
p-adjusted = 0.037 (Tukey HSD)
Researcher-developed rubric; two raters scored 12 essays (all ICCs > 0.85), remainder scored by a single researcher. Essay was revised while the support tool was available, so this measures assisted performance; authors observed some learners copy-pasting ChatGPT-generated sentences to cater to the rubric.
F2 · Essay score improvement (as F1)
ChatGPT group (AI group) vs Human Expert group (HE group)With AI at assessment — assisted performance, not learning evidenceDuring interventionSame contentGraded artifactPrespecification unclear
Benefit
Mean difference = 2.120 points (AI M=3.600 vs HE M=1.480)
95% CI [0.191, 4.049] (Tukey table reports HE-AI: -2.120 [-4.049, -0.191])
p-adjusted = 0.025 (Tukey HSD)
Same measure and caveats as F1.
F3 · Essay score improvement (as F1)
ChatGPT group (AI group) vs Checklist group (CL group)With AI at assessment — assisted performance, not learning evidenceDuring interventionSame contentGraded artifactPrespecification unclear
Benefit
Mean difference = 2.200 points (AI M=3.600 vs CL M=1.400)
95% CI [0.367, 4.033] (Tukey table reports CL-AI: -2.200 [-4.033, -0.367])
p-adjusted = 0.012 (Tukey HSD)
Same measure and caveats as F1.
F4 · Knowledge gain: pre- to post-test difference on a 10-item researcher knowledge test on AI in education
ChatGPT group (AI group) vs Measured without AIImmediate post-testSame contentResearcher-developed testPrespecification unclear
Null
No significant group differences in pre-test (F=1.294, p=0.281, eta2=0.036) or post-test (F=0.913, p=0.438, eta2=0.030) scores, reported as no significant differences in knowledge gain
not reported
pre-test p=0.281; post-test p=0.438
10 single/multiple-choice items; authors state test reliability was examined in prior (anonymised self-cited) studies. Post-test completed within one day after the lab session, i.e., after tool access ended; administration/supervision conditions not described. Analysis Ns shrink to 97-107 without explanation.
F5 · Knowledge transfer: 10-item researcher test on AI in healthcare (post-task only)
ChatGPT group (AI group) vs Measured without AIImmediate post-testNear transferResearcher-developed testPrespecification unclear
Null
F=0.019, eta2=0.000 (group means 0.771-0.785)
not reported
p=0.996
Same caveats as F4 (post-test within one day, no tool access, unexplained N=107). Transfer domain (AI in healthcare) is adjacent to the studied domain (AI in education).
F6 · Post-task intrinsic motivation: Intrinsic Motivation Inventory (IMI), four subscales
ChatGPT group (AI group) vs Measured without AIImmediate post-testTransfer not assessedSelf-reportPrespecification unclear
Null
No significant group differences on any subscale: Interest/Enjoyment F=1.087, eta2=0.029; Perceived Competence F=0.453, eta2=0.012; Effort/Importance F=1.152, eta2=0.030; Pressure/Tension F=0.546, eta2=0.015
not reported
p=0.358 / 0.716 / 0.332 / 0.652 respectively
Established instrument (IMI); overall Cronbach's alpha 0.82, subscales 0.86-0.94. Administered in the post-task questionnaire (N=114); measures motivation toward the experimental task only.
F7 · Frequency of trace-parsed SRL processes in the revising stage (orientation, planning, monitoring, evaluation, reading, elaboration/organisation, other)
ChatGPT group (AI group) vs Control group (CN group)With AI at assessment — assisted performance, not learning evidenceDuring interventionTransfer not assessedProcess / behaviorPrespecification unclear
Mixed
Significant Kruskal-Wallis differences in revising-stage process frequencies (exact statistics shown only as significance stars in Figure 3): AI, HE and CL groups showed more Elaboration/Organisation and more Orientation than CN; AI and HE groups showed less Reading; no group differences in Monitoring or Planning; only CL showed increased Evaluation
not reported
Reported only as figure significance levels (p<0.05 to p<0.0001 stars); no exact values in text
Behavioural trace data (navigation logs, clicks, mouse, keystrokes) parsed to processes via a published trace-parser action/process library; Kruskal-Wallis with Mann-Whitney post hocs. Direction coded mixed: more writing/orientation activity but less reading for the AI group, and no evaluation increase (unlike CL).
F8 · Temporal SRL process models (first-order Markov process mining of revising-stage transitions) - basis of the "metacognitive laziness" claim
ChatGPT group (AI group) vs With AI at assessment — assisted performance, not learning evidenceDuring interventionTransfer not assessedProcess / behaviorPrespecification unclear
Mixed
Descriptive only: transitions differing by >10% probability highlighted. AI group showed a dominant loop between essay revision and ChatGPT interaction and relatively fewer metacognitive processes (evaluation, orientation) than HE and CL groups; HE group showed stronger revising-reading and orientation-evaluation links; AI group showed more planning-before-reading transitions
not reported
not reported (no inferential statistics for process maps)
pMineR first-order Markov models with a 10% transition-probability difference threshold; interpretive, not hypothesis-tested. Authors acknowledge the study lacked a targeted, mature measure of metacognitive laziness.
Limitations
Source-stated: possible lack of power from task duration and sample size; 70% female sample limiting representativeness; single reading/writing task; no long-term follow-up; no targeted, mature measure of metacognitive laziness. Observed: group Ns inconsistent between methods text, essay table, IMI table, and knowledge-test tables with no attrition accounting; the appendix "Essay Score after Revision" row for CL duplicates its before-revision values (14.367, SD 3.882), an apparent table error; the essay improvement (the only significant performance benefit) was produced while the tool was in use, and authors themselves observed learners copy-pasting ChatGPT-generated sentences to cater to the rubric; knowledge/transfer post-test was completed "within one day" after the session, with supervision conditions unstated; process-mining comparisons use a 10% transition-probability threshold, not significance tests; ChatGPT was restricted to task topics and advice-giving, so results may not generalise to unrestricted chatbot use.

Who was studied

Level
university
Ages
mean 22.61 (SD=3.39); 70% female; 55% undergraduates
Country
Not explicitly stated in the arXiv version (participants' first language is redacted as "[disclosed]"); lead authors are at Peking University, Beijing, China, and all participants were English-as-second-language speakers.
Prior knowledge
Not directly reported beyond a pre-test knowledge score (no significant group differences, F=1.294, p=0.281); participants came from "a diverse array of disciplines".
Selection
117 university students recruited July-September 2023; recruitment method not reported.

Methodological notes

Extracted from: https://arxiv.org/pdf/2412.09315 Access: Full text extracted from the arXiv version (arXiv:2412.09315, PDF converted to text locally), not the Wiley version of record (DOI 10.1111/bjet.13544). The arXiv copy is partially anonymised: participants' first language is "[disclosed]" and some self-citations appear as "Author (2023)" / "Authors, 2022, 2023". Sample-size note: Methods text: CN (control) 30, AI (ChatGPT) 35, HE (human expert) 25, CL (checklist) 27. Inconsistent with appendix essay-score table (CN 27, AI 35, HE 25, CL 30, total 117) and with the IMI motivation table (CN 27, AI 35, HE 24, CL 28, total 114). Knowledge-test analysis Ns are smaller still (pretest 107, posttest 105, gain 97, transfer 107).

01bRisk of bias

How much each result can be relied on

Risk of bias is judged per result, not per paper: the same study can report one outcome at low risk and another at high. Two reviewers assess each result independently against a written guide, and every domain judgment below carries its reasoning and the sentences it rests on.

Risk of bias is not the same as whether the numbers work. A trial can be well conducted and still print a standard deviation that is arithmetically impossible, or a table whose sign contradicts its own text. Risk of bias has no question for that, so where it was found it is recorded separately below, under Source integrity.

These words are not quality scores. On the RoB 2 scale, High risk of bias is the worst rating a result can receive and Low risk of bias is the best — the opposite of how the same words read on some other scales. Levels are always written out in full here for that reason.

KT

High risk of bias · RoB 2 · Two reviewers agreed on every domain

How the overall was reached: Worst-domain rule: D3 is High because 9-17% of randomized participants are missing from the knowledge-gain and transfer analyses with no explanation, no flow accounting, and no sensitivity analysis, the reported arm Ns are internally inconsistent (score-improvement arm Ns sum to 105 but the total is printed as 97), and the unsupervised within-one-day post-test makes outcome-dependent missingness plausible. D1, D4, and D5 each carry Some concerns (no randomization-procedure or concealment detail; researcher-developed tests self-administered without stated supervision; no preregistration), reinforcing the overall High judgment for this null four-group result.

DomainJudgment and reasoning
D1 · Randomisation process
Some concerns about risk of bias
The study is described as randomized and participants were "randomly assigned" to the four arms, but no information is given on the randomization method (sequence generation) or on allocation concealment. Arm sizes are unequal (30/35/25/27) with no stated allocation ratio, which is compatible with simple randomization but cannot be verified. Baseline knowledge (pre-test) did not differ significantly between groups (F=1.294, p=0.281) and pre-test means are similar across arms (0.459-0.533 in Appendix Table 2), so there is no positive evidence of a problem with the randomization process; however, the baseline comparison itself is computed on a reduced N (107 of 117) and essay-score baseline group Ns in Appendix Table 1 (CN 27, CL 30) contradict the text's randomized arm sizes (CN 30, CL 27), leaving some uncertainty. Absent any description of sequence generation or concealment, Some concerns is appropriate.
The participants were randomly assigned to four experimental groups: one group did not have any support and finished the task by themselves (CN group, 30 participants); one group of learners were supported by ChatGPT 4.0 (AI group, 35 participants); one group of learners were supported by a human expert (HE group, 25 participants); and one group of learners had the support of the writing analytics toolkit named Checklist Tools (CL group, 27 participants).
The ANOVA results indicated no significant differences between the groups in terms of the pre-test score (F=1.294, p=0.281, η²=0.036) and post-test score (F=0.913, p=0.438, η²=0.030), which means no significant differences in knowledge gain.
D2 · Deviations from intended interventions
Low risk of bias
Participants and personnel were necessarily aware of assigned support (a chat with ChatGPT, a human expert, checklist tools, or nothing), but for the effect of assignment the question is whether deviations from intended interventions arose because of the trial context. The intervention was delivered inside a controlled single-session lab platform with scripted training videos, and tool access was technically embedded and restricted (ChatGPT constrained to the task via API); learning trace data confirm participants interacted with their assigned tools. No cross-arm contamination, co-intervention, or context-driven deviation is reported or plausible in a supervised lab session. Available-case rather than ITT analysis is a missing data issue handled under D3. Low risk for this domain.
As shown in Figure 1, we conducted our experiment in the lab, where participants were required to complete the task following six steps: pre-task, stage 1 training, stage 1 reading and writing, stage 2 training, stage 2 revising, and post-task.
the AI group received assistance from ChatGPT 4.0 (embedded in our platform user interfaces), which was trained and restricted to the content covered by our learning task (only conversations based on the learning task are allowed).
D3 · Missing outcome data
High risk of bias
117 students were randomized, but the knowledge outcomes are analyzed on substantially fewer: Appendix Table 2 reports pre-test N=107 (CN 27, AI 32, HE 21, CL 27), post-test N=105 (25/33/23/24), and a "Score Improvement" total printed as N=97 even though its own arm Ns (25/33/23/24) sum to 105 - an internal inconsistency the paper never addresses; Appendix Table 3 reports transfer N=107 (26/34/21/26). That is 9-17% missing outcome data with no explanation anywhere in the paper (no participant flow, no reasons for exclusion, no sensitivity analysis), and attrition is uneven across arms (HE loses 4 of 25 at pre-test alone). Because the post-test was completed "within one day" after participants finished the lab tasks rather than under direct supervision at a fixed point, non-completion is plausibly related to participants' engagement, motivation, or knowledge - the true value of the outcome - and could plausibly differ by assigned support. With non-trivial, unexplained, internally inconsistent missingness, no evidence that the result is unbiased, and a plausible outcome-dependent missingness mechanism, this domain is High risk.
A total of 117 university students (average age 22.61, SD=3.39, with 70% identifying as female and 55% undergraduates) participated in the experiment from July to September 2023.
Once these tasks were completed, participants were asked to complete the post-test within one day.
D4 · Measurement of the outcome
Some concerns about risk of bias
The knowledge and transfer tests are objective (10 single- or multiple-choice items each), so scorer subjectivity and assessor blinding are minor issues, and the instruments had reliability examined in the authors' previous studies. However, the tests were researcher-developed (self-citations "Authors, 2022, 2023"), and the post-test was completed "within one day" of the session rather than under stated supervision, with no description of proctoring or of measures preventing use of outside resources (including AI tools) during the unsupervised window; the paper never states how or where the post-test was administered. Measurement methods and timing were the same across arms, so any resulting error is not clearly differential, but unsupervised self-administration by outcome-aware participants up to a day later leaves Some concerns.
The knowledge test on AI in education (10 items of single or multiple-choice questions) and the transfer test on AI in healthcare (10 items of single or multiple-choice questions) were developed and examined the reliability in previous studies (Authors, 2022, 2023).
Once these tasks were completed, participants were asked to complete the post-test within one day.
D5 · Selection of the reported result
Some concerns about risk of bias
No preregistration, trial registration, protocol, or pre-specified statistical analysis plan is mentioned anywhere in the arXiv version, so the reported analyses cannot be compared against prior intentions. The analysis itself is simple and consistently reported (ANOVA with Tukey's HSD across the three pre-declared performance dimensions, with descriptive tables for all arms in the Appendix), and both null knowledge results are reported alongside the significant essay result, which argues against aggressive selective reporting. Under RoB 2, absence of a pre-specified plan without evidence of selection yields Some concerns.
To answer RQ3, we evaluated learners' learning performance across three dimensions: 1) essay score improvement (difference in essay scores before and after revising), 2) knowledge gain (difference between pre- and post-test scores on the same knowledge test on AI in education), and 3) knowledge transfer (knowledge test score on AI in healthcare).
We compared the transfer test scores between the four groups, and ANOVA results showed that there were no significant differences between the four groups (F=0.019, p=0.996,η²=0.000).
02Funding

Funding and conflicts

Funding
National Natural Science Foundation of China (Grant 62407001); Society for Learning Analytics Research (ECR Research Grant, 2023).
Vendor funded
No
Vendor
None identified
Notes
not_reported (no conflict-of-interest statement in the arXiv text)

Funding is shown on every study and never used to score it.

03Claims

Claims this study bears on

04Source

The source, as retrieved

Abstract

No abstract retrieved.

Where this record came from

SourceRetrievedIdentifier
seedSep 18, 2026link
first ingestion