AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

Guardrailed AI tutors (hints, no direct answers, pedagogical prompting) improve secondary students' performance on assessments taken without AI, compared with business-as-usual instruction.

Current status
ContradictoryEvidence unclear · what the nine statuses mean
Why
A well-powered trial shows no unassisted gain when guardrailed tutoring replaces equivalent practice time; a preliminary supervised after-school program shows +0.24 SD when the program adds instruction. Candidate explanation: what the comparator holds constant (time and materials). No design yet isolates it.
What would change it
Randomized trials with active non-AI comparators matched for time and materials.

Delayed unassisted outcomes rather than same-session exams.
Linked evidence
3 links · 1 supports · 1 contradicts · 1 consistent, but doesn't test the claim
Last updated
Sep 18, 2026

The status is the collection's own judgment on its nine-label scale — not a certainty rating, not a GRADE level, and not advice. Every status change is dated, reasoned, and kept below under History.

01The evidence

The evidence, sorted by what it shows

Supports (1)

  • From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in NigeriaRandomized trial · 2025 · Preliminary · Secondary direct · University different · Working paper

    English endline (pencil-and-paper, no AI): +0.238 SD for the supervised after-school Copilot program vs no program. Caveats: the comparator bundles no extra instruction time (the program adds it); PRELIMINARY working paper; attrition handling assessed-not-settled (D3). Rests on finding F3, judged at assessment v2.

Contradicts (1)

  • Generative AI without guardrails can harm learning: Evidence from high school mathematicsCluster-randomized trial · 2025 · Study tier 1 · Secondary direct · University different

    Guardrailed tutor indistinguishable from business-as-usual on the unassisted same-session exam (-0.004) despite +127% assisted practice: a fairly precise null in a large preregistered trial. One site, one subject, immediate only. Rests on finding F4, judged at assessment v2.

Consistent, but doesn't test the claim (1)

  • Effective and Scalable Math Support: Experimental Evidence on the Impact of an AI Math Tutor in GhanaCluster-randomized trial · 2024 · Study tier 3 · Secondary partial · University different · Partly vendor funded

    Positive growth effect consistent with the claim, but assessment AI-availability is not explicit (gate), population mostly primary (partial), RoB High (differential attrition; unclustered analysis over 11 schools), vendor-affiliated authors. Rests on finding F1, judged at assessment v2.

02Certainty

Certainty by outcome

A claim can be broken into separate bodies of evidence — one per outcome, split by whether AI was available at assessment and when the outcome was measured. Each body carries four separate judgments: how confident we are (certainty), what the evidence points to (the conclusion), how directly it speaks to this claim (applicability), and who has stood behind the judgment. Confidence and conclusion are never merged into one word.

UA-IMM — Unassisted performance after guardrailed tutoring · AI at assessment: no · immediate post

Field Assembly certainty: Weak (not a GRADE rating — what our scale means) · Conclusion: Varies by intervention design · Applicability: Partial · AI: two passes agreed

Rated against: improvement vs business-as-usual

Null when the tutor replaces equivalent practice time; +0.24 SD when a supervised program adds instruction — what the comparator holds constant is the live question

DomainJudgment and reasoning
Risk of biasserious · one HIGH trial; one PRELIMINARY working paper with unsettled attrition handling (D3)
Inconsistencyserious · a precise null against a matched-time comparator vs a positive effect when the program adds instruction — heterogeneity with a design-shaped candidate explanation, not settled
Indirectnessserious · comparator structures differ from the claim's framing; same-session vs six-week outcomes
Imprecisionserious · two studies, different constructs
Reporting and publication biasnot serious · not assessable

Population: secondary students (Turkey; Nigeria)
Comparator: business-as-usual (matched-time practice; no program)
Assessed by software, two independent passes · search: seed/corpus.yaml (verified inventory, proposal section 7) + data/searches/ · method: EVIDENCE-MODEL.md v2 + fa-certainty-scale v2

2 studies in this body. 1 further record was considered and left out:

03History

How this assessment has changed

  • ContradictorySep 18, 2026 · Evidence unclear

    initial assessed status (Q-009, owner-accepted IN-018)

  • Not yet assessedSep 18, 2026 · Evidence unclear

    initial curated status