AI-assisted learning, measured without the AI
Generative AI is moving into classrooms faster than the evidence about it. This collection records what published studies actually show about one question: does AI used during instruction or practice improve what learners can do on their own, afterward?
- Studies collected
- 14Every count on this page is of this collection, not of the literature.
- Claims assessed
- 9 of 9Each claim is a single testable statement, assessed per population, with the evidence for and against it.
- Findings recorded
- 99Of the 99 outcome contrasts recorded here, 39 were measured without AI at assessment, and 14 of those on a delayed test.
- Open questions
- 5What is not known, and what study would settle it.
- Snapshot
- Sep 18, 2026
Counts are from the snapshot date above. Assessments carry a review-status label saying who has stood behind them.
The question every entry answers
Does this study measure what learners can do without the AI, after using it? Most studies of AI in education measure performance while the tool is still available. That is assisted performance — a fact about the learner-plus-tool system — and it is never counted here as evidence of learning. Every finding in this collection is badged with whether AI was available at assessment, when the outcome was measured, and how far it sits from the practiced material.
The distinction is not pedantry: the studies that separate the two keep finding large gains with the tool alongside flat or negative results without it.
Claims, and where the evidence stands
- Unrestricted use of a general-purpose LLM chatbot during practice improves secondary students' subsequent performance on assessments taken without AI.Not supported · Secondary students · Evidence stable
- Unrestricted use of a general-purpose LLM chatbot during practice improves university students' subsequent performance on assessments taken without AI.Insufficient evidence · University students · Evidence unclear
- Guardrailed AI tutors (hints, no direct answers, pedagogical prompting) improve secondary students' performance on assessments taken without AI, compared with business-as-usual instruction.Contradictory · Secondary students · Evidence unclear
- Guardrails prevent the unassisted-performance harm observed when secondary students practice with answer-giving chatbots.Preliminary signal · Secondary students · Evidence stable
- AI feedback on writing produces unassisted writing growth noninferior to human feedback — no worse than a stated margin.Insufficient evidence · University students · Evidence stable
- Learning gains from AI-assisted practice persist on delayed (one week or longer) retention tests taken without AI.Insufficient evidence · Secondary students · Evidence stable
Recent changes in the evidence
- Claim assessed: Preliminary SignalClaim status changed · Sep 18, 2026
"Secondary students with low prior knowledge gain less, or are harmed more, by AI assistance during practice than high-prior-knowledge peers." received its first assessed status: PRELIMINARY_SIGNAL (unclear). Two independent weak analyses point the claim's way (retention gains confined to high-prior learners; a positive treatment-by-baseline interaction), while the strongest trial's preregistered heterogeneity analysis — gated for stance because its assessment conditions are not reported — found no detectable ability moderation. An early, internally inconsistent signal.
- Claim assessed: Insufficient EvidenceClaim status changed · Sep 18, 2026
"Learning gains from AI-assisted practice persist on delayed (one week or longer) retention tests taken without AI." received its first assessed status: INSUFFICIENT_EVIDENCE (stable). One underpowered null (N=33, ~43% power). The claim is effectively untested at meaningful precision.
- Claim assessed: Insufficient EvidenceClaim status changed · Sep 18, 2026
"Learning gains from AI-assisted practice persist on delayed (one week or longer) retention tests taken without AI." received its first assessed status: INSUFFICIENT_EVIDENCE (stable). One study measured delayed unassisted retention: three concordant nulls with favorable trends at N=69, plus an unprespecified positive high-prior subgroup. Neither persistence nor decay is demonstrated.
- Claim assessed: Insufficient EvidenceClaim status changed · Sep 18, 2026
"AI feedback on writing produces unassisted writing growth noninferior to human feedback — no worse than a stated margin." received its first assessed status: INSUFFICIENT_EVIDENCE (stable). No noninferiority evidence exists: the one head-to-head comparison is small, non-randomized (Serious ROBINS-I), and directionally favors the human tutor without significance. A nonsignificant difference does not establish the claim; as written it is untested.
- Claim assessed: Proxy OnlyClaim status changed · Sep 18, 2026
"University students with low prior knowledge gain less, or are harmed more, by AI assistance during practice than high-prior-knowledge peers." received its first assessed status: PROXY_ONLY (unclear). The only relevant evidence is qualitative and process-tracing (struggling CS1 students' difficulties compounding, with an illusion of competence): consistent with differential harm, but entirely proxy measures with no unassisted outcome and no comparison arm.
- Claim assessed: Preliminary SignalClaim status changed · Sep 18, 2026
"Guardrails prevent the unassisted-performance harm observed when secondary students practice with answer-giving chatbots." received its first assessed status: PRELIMINARY_SIGNAL (stable). One strong three-arm trial shows the vanilla-chat harm eliminated (not reversed) by teacher-designed guardrails — the claim's exact contrast, within a single study, site, subject, and product. Pass B read this as PROBABLE on the strength of the internal contrast; the single-study rule holds it at PRELIMINARY_SIGNAL until replicated.
What the research has not answered
- Does generative AI use during learning impair or improve transfer to unfamiliar problem types?Open
- Does sustained reliance on generative AI erode previously acquired skills?Open
- Does teacher supervision change what learners retain from AI use, holding the tool constant?Open
- Do findings from university students generalize to secondary students, and vice versa?Open
- Does offering AI tools change participation and engagement, separately from any effect on learning?Open