AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

Learning gains from AI-assisted practice persist on delayed (one week or longer) retention tests taken without AI.

Current status
Insufficient evidenceEvidence stable · what the nine statuses mean
Why
One study measured delayed unassisted retention: three concordant nulls with favorable trends at N=69, plus an unprespecified positive high-prior subgroup. Neither persistence nor decay is demonstrated.
What would change it
Nearly any delayed unassisted retention evidence — the emptiest cell in the literature.
Linked evidence
1 link · 1 mixed
Last updated
Sep 18, 2026

The status is the collection's own judgment on its nine-label scale — not a certainty rating, not a GRADE level, and not advice. Every status change is dated, reasoned, and kept below under History.

01The evidence

The evidence, sorted by what it shows

Mixed (1)

  • Studying the effect of AI Code Generators on Supporting Novice Learners in Introductory ProgrammingRandomized trial · 2023 · Study tier 2 · Secondary direct · University different

    One-week unassisted retention (authoring F7; also modification F8, MCQ F9): no significant differences after large assisted gains, with trends favoring the AI-trained group (d 0.28-0.41, ns, n=69) and a significantly positive high-prior subgroup. No demonstrated persistence, no demonstrated decay. Rests on finding F7, judged at assessment v2.

02Certainty

Certainty by outcome

A claim can be broken into separate bodies of evidence — one per outcome, split by whether AI was available at assessment and when the outcome was measured. Each body carries four separate judgments: how confident we are (certainty), what the evidence points to (the conclusion), how directly it speaks to this claim (applicability), and who has stood behind the judgment. Confidence and conclusion are never merged into one word.

UA-DEL — Unassisted retention at >=1 week · AI at assessment: no · delayed post

Field Assembly certainty: Weak (not a GRADE rating — what our scale means) · Conclusion: Not estimable · Applicability: Partial · AI: two passes agreed

Rated against: persistent advantage for the AI-trained group

No demonstrated persistence and no demonstrated decay; the emptiest cell in the literature remains near-empty

DomainJudgment and reasoning
Risk of biasserious · Some concerns with an unsettled D2; completers-only analysis
Inconsistencynot serious · three concordant nulls within one study
Indirectnessserious · code generator, not chat; camp setting
Imprecisionvery serious · N=69; trends (d 0.28-0.41) compatible with real persistence or none
Reporting and publication biasnot serious · not assessable

Population: secondary-age novice programmers (10-17)
Comparator: same training without the generator
Assessed by software, two independent passes · search: seed/corpus.yaml (verified inventory, proposal section 7) + data/searches/ · method: EVIDENCE-MODEL.md v2 + fa-certainty-scale v2

1 study in this body.

03History

How this assessment has changed

  • Insufficient evidenceSep 18, 2026 · Evidence stable

    initial assessed status (Q-009, owner-accepted IN-018)

  • Not yet assessedSep 18, 2026 · Evidence unclear

    initial curated status