AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

Unrestricted use of a general-purpose LLM chatbot during practice improves secondary students' subsequent performance on assessments taken without AI.

Current status
Not supportedEvidence stable · what the nine statuses mean
Why
The claim asserts improvement. The one direct, gate-passing test (preregistered cluster RCT, HIGH) found a 17% unassisted-exam deficit; nothing gate-passing supports the claim. A single setting and product generation keep this NOT_SUPPORTED rather than a harm claim of its own; replication would move it.
What would change it
Adequately sized randomized studies with unassisted post-tests, in either direction.

Independent replication or contradiction of the existing large three-arm RCT's unassisted-exam result.
Linked evidence
3 links · 1 contradicts · 1 consistent, but doesn't test the claim · 1 no direct evidence
Last updated
Sep 18, 2026

The status is the collection's own judgment on its nine-label scale — not a certainty rating, not a GRADE level, and not advice. Every status change is dated, reasoned, and kept below under History.

01The evidence

The evidence, sorted by what it shows

Contradicts (1)

  • Generative AI without guardrails can harm learning: Evidence from high school mathematicsCluster-randomized trial · 2025 · Study tier 1 · Secondary direct · University different

    Preregistered primary outcome: vanilla GPT-4 chat practice reduced unassisted exam performance 17% vs business-as-usual (cluster RCT, HIGH). Direct test; direction opposite to the claim. RoB for this result is assessed-not-settled on D1 (reporting detail). Rests on finding F3, judged at assessment v2.

Consistent, but doesn't test the claim (1)

No direct evidence (1)

02Certainty

Certainty by outcome

A claim can be broken into separate bodies of evidence — one per outcome, split by whether AI was available at assessment and when the outcome was measured. Each body carries four separate judgments: how confident we are (certainty), what the evidence points to (the conclusion), how directly it speaks to this claim (applicability), and who has stood behind the judgment. Confidence and conclusion are never merged into one word.

UA-IMM — Unassisted performance, immediate · AI at assessment: no · immediate post

Field Assembly certainty: Supported (not a GRADE rating — what our scale means) · Conclusion: Harm · Applicability: Partial · AI: two passes agreed

Rated against: any improvement vs comparator (claim direction); observed direction is the reverse

-17% on the unassisted exam (GPT Base vs control), P<0.05; the claim's direction is unsupported and the observed direction is harm

DomainJudgment and reasoning
Risk of biasnot serious · single HIGH cluster RCT; D1 two-pass split (reporting detail) published as unsettled; both readings leave the design strong
Inconsistencynot serious · single study; no conflicting unassisted secondary evidence in the collection
Indirectnessserious · one school network, one country, Fall-2023 GPT-4, same-session outcome only
Imprecisionserious · one trial, one setting; precise within it (SE 0.022) but unreplicated
Reporting and publication biasnot serious · not assessable for a single preregistered trial

Population: secondary students (grades 9-11, one Turkish school network)
Comparator: business-as-usual practice (textbooks/notes)
Assessed by software, two independent passes · search: seed/corpus.yaml (verified inventory, proposal section 7) + data/searches/ · method: EVIDENCE-MODEL.md v2 + fa-certainty-scale v2

1 study in this body. 2 further records were considered and left out:

03History

How this assessment has changed

  • Not supportedSep 18, 2026 · Evidence stable

    initial assessed status (Q-009, owner-accepted IN-018)

  • Not yet assessedSep 18, 2026 · Evidence unclear

    initial curated status