AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

Unrestricted use of a general-purpose LLM chatbot during practice improves university students' subsequent performance on assessments taken without AI.

Current status
Insufficient evidenceEvidence unclear · what the nine statuses mean
Why
Evidence leans against the claim — a significant post-use quality deficit vs web search, and an underpowered retention null — but nothing adequately powered tests it directly, and the largest randomized offer cannot (partial AI at its exam). Pass B read the same evidence as NOT_SUPPORTED; the conservative status stands until something powered exists.
What would change it
Any adequately sized university RCT with unassisted post-tests and an active comparison.
Linked evidence
5 links · 2 contradicts · 2 consistent, but doesn't test the claim · 1 no direct evidence
Last updated
Sep 18, 2026

The status is the collection's own judgment on its nine-label scale — not a certainty rating, not a GRADE level, and not advice. Every status change is dated, reasoned, and kept below under History.

01The evidence

The evidence, sorted by what it shows

Contradicts (2)

Consistent, but doesn't test the claim (2)

No direct evidence (1)

02Certainty

Certainty by outcome

A claim can be broken into separate bodies of evidence — one per outcome, split by whether AI was available at assessment and when the outcome was measured. Each body carries four separate judgments: how confident we are (certainty), what the evidence points to (the conclusion), how directly it speaks to this claim (applicability), and who has stood behind the judgment. Confidence and conclusion are never merged into one word.

UA-ANY — Unassisted performance, immediate and delayed · AI at assessment: no · immediate post

Field Assembly certainty: Weak (not a GRADE rating — what our scale means) · Conclusion: Inconsistent · Applicability: Partial · AI: two passes agreed

Rated against: any improvement vs comparator

No benefit demonstrated; one significant post-use quality deficit and one underpowered retention null trending pro-ChatGPT

DomainJudgment and reasoning
Risk of biasserious · a LOW trial (N=33, unsettled D4) and a Some-concerns lab experiment
Inconsistencyserious · a significant quality deficit vs an underpowered null trending the other way — the two results disagree in kind
Indirectnessserious · single-session/single-quiz designs; neither is sustained practice
Imprecisionvery serious · N=33 and N=91; the null is underpowered by the authors' own analysis
Reporting and publication biasnot serious · not assessable at two studies

Population: university students (medical; general German university)
Comparator: non-AI resources (institutional materials; Google search)
Assessed by software, two independent passes · search: seed/corpus.yaml (verified inventory, proposal section 7) + data/searches/ · method: EVIDENCE-MODEL.md v2 + fa-certainty-scale v2

2 studies in this body. 3 further records were considered and left out:

03History

How this assessment has changed

  • Insufficient evidenceSep 18, 2026 · Evidence unclear

    initial assessed status (Q-009, owner-accepted IN-018)

  • Not yet assessedSep 18, 2026 · Evidence unclear

    initial curated status