AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

Guardrails prevent the unassisted-performance harm observed when secondary students practice with answer-giving chatbots.

Current status
Preliminary signalEvidence stable · what the nine statuses mean
Why
One strong three-arm trial shows the vanilla-chat harm eliminated (not reversed) by teacher-designed guardrails — the claim's exact contrast, within a single study, site, subject, and product. Pass B read this as PROBABLE on the strength of the internal contrast; the single-study rule holds it at PRELIMINARY_SIGNAL until replicated.
What would change it
More three-arm (tutor vs vanilla vs control) designs with unassisted outcomes; there is currently essentially one.
Linked evidence
2 links · 2 supports
Last updated
Sep 18, 2026

The status is the collection's own judgment on its nine-label scale — not a certainty rating, not a GRADE level, and not advice. Every status change is dated, reasoned, and kept below under History.

01The evidence

The evidence, sorted by what it shows

Supports (2)

02Certainty

Certainty by outcome

A claim can be broken into separate bodies of evidence — one per outcome, split by whether AI was available at assessment and when the outcome was measured. Each body carries four separate judgments: how confident we are (certainty), what the evidence points to (the conclusion), how directly it speaks to this claim (applicability), and who has stood behind the judgment. Confidence and conclusion are never merged into one word.

UA-IMM — Harm prevention, unassisted immediate · AI at assessment: no · immediate post

Field Assembly certainty: Supported (not a GRADE rating — what our scale means) · Conclusion: Varies by intervention design · Applicability: Partial · AI: two passes agreed

Rated against: guardrailed arm avoids the vanilla arm's deficit

-17% (vanilla) vs -0.4% ns (guardrailed) on the same unassisted exam: design determines the outcome

DomainJudgment and reasoning
Risk of biasnot serious · HIGH trial; D1 reporting split published as unsettled
Inconsistencynot serious · single trial, internally consistent contrast
Indirectnessserious · one product, one guardrail design, one setting; prevention of harm, not creation of gains
Imprecisionserious · essentially one three-arm trial in the literature
Reporting and publication biasnot serious · not assessable

Population: secondary students (grades 9-11, one Turkish school network)
Comparator: within-trial three-arm contrast
Assessed by software, two independent passes · search: seed/corpus.yaml (verified inventory, proposal section 7) + data/searches/ · method: EVIDENCE-MODEL.md v2 + fa-certainty-scale v2

1 study in this body.

03History

How this assessment has changed

  • Preliminary signalSep 18, 2026 · Evidence stable

    initial assessed status (Q-009, owner-accepted IN-018)

  • Not yet assessedSep 18, 2026 · Evidence unclear

    initial curated status