AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

Changes

Every change in the evidence, typed and dated: new results, status changes, corrections, retractions. Each entry says what the new evidence is and whether any claim status changed because of it — or why not.

Also available as an RSS feed. Claim pages carry their own full status history.

01The record
  • Claim assessed: Preliminary SignalClaim status changed · Sep 18, 2026 · Notable

    "Secondary students with low prior knowledge gain less, or are harmed more, by AI assistance during practice than high-prior-knowledge peers." received its first assessed status: PRELIMINARY_SIGNAL (unclear). Two independent weak analyses point the claim's way (retention gains confined to high-prior learners; a positive treatment-by-baseline interaction), while the strongest trial's preregistered heterogeneity analysis — gated for stance because its assessment conditions are not reported — found no detectable ability moderation. An early, internally inconsistent signal.

  • Claim assessed: Insufficient EvidenceClaim status changed · Sep 18, 2026 · Notable

    "Learning gains from AI-assisted practice persist on delayed (one week or longer) retention tests taken without AI." received its first assessed status: INSUFFICIENT_EVIDENCE (stable). One underpowered null (N=33, ~43% power). The claim is effectively untested at meaningful precision.

  • Claim assessed: Insufficient EvidenceClaim status changed · Sep 18, 2026 · Notable

    "Learning gains from AI-assisted practice persist on delayed (one week or longer) retention tests taken without AI." received its first assessed status: INSUFFICIENT_EVIDENCE (stable). One study measured delayed unassisted retention: three concordant nulls with favorable trends at N=69, plus an unprespecified positive high-prior subgroup. Neither persistence nor decay is demonstrated.

  • Claim assessed: Insufficient EvidenceClaim status changed · Sep 18, 2026 · Notable

    "AI feedback on writing produces unassisted writing growth noninferior to human feedback — no worse than a stated margin." received its first assessed status: INSUFFICIENT_EVIDENCE (stable). No noninferiority evidence exists: the one head-to-head comparison is small, non-randomized (Serious ROBINS-I), and directionally favors the human tutor without significance. A nonsignificant difference does not establish the claim; as written it is untested.

  • Claim assessed: Proxy OnlyClaim status changed · Sep 18, 2026 · Notable

    "University students with low prior knowledge gain less, or are harmed more, by AI assistance during practice than high-prior-knowledge peers." received its first assessed status: PROXY_ONLY (unclear). The only relevant evidence is qualitative and process-tracing (struggling CS1 students' difficulties compounding, with an illusion of competence): consistent with differential harm, but entirely proxy measures with no unassisted outcome and no comparison arm.

  • Claim assessed: Preliminary SignalClaim status changed · Sep 18, 2026 · Notable

    "Guardrails prevent the unassisted-performance harm observed when secondary students practice with answer-giving chatbots." received its first assessed status: PRELIMINARY_SIGNAL (stable). One strong three-arm trial shows the vanilla-chat harm eliminated (not reversed) by teacher-designed guardrails — the claim's exact contrast, within a single study, site, subject, and product. Pass B read this as PROBABLE on the strength of the internal contrast; the single-study rule holds it at PRELIMINARY_SIGNAL until replicated.

  • Claim assessed: ContradictoryClaim status changed · Sep 18, 2026 · Notable

    "Guardrailed AI tutors (hints, no direct answers, pedagogical prompting) improve secondary students' performance on assessments taken without AI, compared with business-as-usual instruction." received its first assessed status: CONTRADICTORY (unclear). A well-powered trial shows no unassisted gain when guardrailed tutoring replaces equivalent practice time; a preliminary supervised after-school program shows +0.24 SD when the program adds instruction. Candidate explanation: what the comparator holds constant (time and materials). No design yet isolates it.

  • Claim assessed: Insufficient EvidenceClaim status changed · Sep 18, 2026 · Notable

    "Unrestricted use of a general-purpose LLM chatbot during practice improves university students' subsequent performance on assessments taken without AI." received its first assessed status: INSUFFICIENT_EVIDENCE (unclear). Evidence leans against the claim — a significant post-use quality deficit vs web search, and an underpowered retention null — but nothing adequately powered tests it directly, and the largest randomized offer cannot (partial AI at its exam). Pass B read the same evidence as NOT_SUPPORTED; the conservative status stands until something powered exists.

  • Claim assessed: Not SupportedClaim status changed · Sep 18, 2026 · Major

    "Unrestricted use of a general-purpose LLM chatbot during practice improves secondary students' subsequent performance on assessments taken without AI." received its first assessed status: NOT_SUPPORTED (stable). The claim asserts improvement. The one direct, gate-passing test (preregistered cluster RCT, HIGH) found a 17% unassisted-exam deficit; nothing gate-passing supports the claim. A single setting and product generation keep this NOT_SUPPORTED rather than a harm claim of its own; replication would move it.