AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

AI-assisted learning, measured without the AI

Generative AI is moving into classrooms faster than the evidence about it. This collection records what published studies actually show about one question: does AI used during instruction or practice improve what learners can do on their own, afterward?

Studies collected
14Every count on this page is of this collection, not of the literature.
Claims assessed
9 of 9Each claim is a single testable statement, assessed per population, with the evidence for and against it.
Findings recorded
99Of the 99 outcome contrasts recorded here, 39 were measured without AI at assessment, and 14 of those on a delayed test.
Open questions
5What is not known, and what study would settle it.
Snapshot
Sep 18, 2026

Counts are from the snapshot date above. Assessments carry a review-status label saying who has stood behind them.

01The question

The question every entry answers

Does this study measure what learners can do without the AI, after using it? Most studies of AI in education measure performance while the tool is still available. That is assisted performance — a fact about the learner-plus-tool system — and it is never counted here as evidence of learning. Every finding in this collection is badged with whether AI was available at assessment, when the outcome was measured, and how far it sits from the practiced material.

The distinction is not pedantry: the studies that separate the two keep finding large gains with the tool alongside flat or negative results without it.

02Claims

Claims, and where the evidence stands

All 9 assessed claims

03What changed

Recent changes in the evidence

  • Claim assessed: Preliminary SignalClaim status changed · Sep 18, 2026

    "Secondary students with low prior knowledge gain less, or are harmed more, by AI assistance during practice than high-prior-knowledge peers." received its first assessed status: PRELIMINARY_SIGNAL (unclear). Two independent weak analyses point the claim's way (retention gains confined to high-prior learners; a positive treatment-by-baseline interaction), while the strongest trial's preregistered heterogeneity analysis — gated for stance because its assessment conditions are not reported — found no detectable ability moderation. An early, internally inconsistent signal.

  • Claim assessed: Insufficient EvidenceClaim status changed · Sep 18, 2026

    "Learning gains from AI-assisted practice persist on delayed (one week or longer) retention tests taken without AI." received its first assessed status: INSUFFICIENT_EVIDENCE (stable). One underpowered null (N=33, ~43% power). The claim is effectively untested at meaningful precision.

  • Claim assessed: Insufficient EvidenceClaim status changed · Sep 18, 2026

    "Learning gains from AI-assisted practice persist on delayed (one week or longer) retention tests taken without AI." received its first assessed status: INSUFFICIENT_EVIDENCE (stable). One study measured delayed unassisted retention: three concordant nulls with favorable trends at N=69, plus an unprespecified positive high-prior subgroup. Neither persistence nor decay is demonstrated.

  • Claim assessed: Insufficient EvidenceClaim status changed · Sep 18, 2026

    "AI feedback on writing produces unassisted writing growth noninferior to human feedback — no worse than a stated margin." received its first assessed status: INSUFFICIENT_EVIDENCE (stable). No noninferiority evidence exists: the one head-to-head comparison is small, non-randomized (Serious ROBINS-I), and directionally favors the human tutor without significance. A nonsignificant difference does not establish the claim; as written it is untested.

  • Claim assessed: Proxy OnlyClaim status changed · Sep 18, 2026

    "University students with low prior knowledge gain less, or are harmed more, by AI assistance during practice than high-prior-knowledge peers." received its first assessed status: PROXY_ONLY (unclear). The only relevant evidence is qualitative and process-tracing (struggling CS1 students' difficulties compounding, with an illusion of competence): consistent with differential harm, but entirely proxy measures with no unassisted outcome and no comparison arm.

  • Claim assessed: Preliminary SignalClaim status changed · Sep 18, 2026

    "Guardrails prevent the unassisted-performance harm observed when secondary students practice with answer-giving chatbots." received its first assessed status: PRELIMINARY_SIGNAL (stable). One strong three-arm trial shows the vanilla-chat harm eliminated (not reversed) by teacher-designed guardrails — the claim's exact contrast, within a single study, site, subject, and product. Pass B read this as PROBABLE on the strength of the internal contrast; the single-study rule holds it at PRELIMINARY_SIGNAL until replicated.

The full change record · RSS

04Open questions

What the research has not answered

The open questions, with the closest evidence so far