Unrestricted use of a general-purpose LLM chatbot during practice improves secondary students' subsequent performance on assessments taken without AI.
- Current status
- Not supportedEvidence stable · what the nine statuses mean
- Why
- The claim asserts improvement. The one direct, gate-passing test (preregistered cluster RCT, HIGH) found a 17% unassisted-exam deficit; nothing gate-passing supports the claim. A single setting and product generation keep this NOT_SUPPORTED rather than a harm claim of its own; replication would move it.
- What would change it
- Adequately sized randomized studies with unassisted post-tests, in either direction.
Independent replication or contradiction of the existing large three-arm RCT's unassisted-exam result. - Linked evidence
- 3 links · 1 contradicts · 1 consistent, but doesn't test the claim · 1 no direct evidence
- Last updated
- Sep 18, 2026
The status is the collection's own judgment on its nine-label scale — not a certainty rating, not a GRADE level, and not advice. Every status change is dated, reasoned, and kept below under History.
The evidence, sorted by what it shows
Contradicts (1)
- Generative AI without guardrails can harm learning: Evidence from high school mathematicsCluster-randomized trial · 2025 · Study tier 1 · Secondary direct · University different
Preregistered primary outcome: vanilla GPT-4 chat practice reduced unassisted exam performance 17% vs business-as-usual (cluster RCT, HIGH). Direct test; direction opposite to the claim. RoB for this result is assessed-not-settled on D1 (reporting detail). Rests on finding F3, judged at assessment v2.
Consistent, but doesn't test the claim (1)
- ChatGPT's impact on student learning outcomes: a meta-analysis of 35 experimental studiesMeta-analysis · 2026 · Study tier 3 · Secondary partial · University partial
Pooled g=0.670 does not separate assisted from unassisted outcomes and spans intervention classes; AMSTAR 2 Critically low. Rests on finding F1, judged at assessment v2.
No direct evidence (1)
- From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in NigeriaRandomized trial · 2025 · Preliminary · Secondary direct · University different · Working paper
Positive unassisted endline, but teacher-supervised prompted-pedagogy Copilot use is not unrestricted chatbot use (intervention-class lint). Whole-study judgment at assessment v2.
Certainty by outcome
A claim can be broken into separate bodies of evidence — one per outcome, split by whether AI was available at assessment and when the outcome was measured. Each body carries four separate judgments: how confident we are (certainty), what the evidence points to (the conclusion), how directly it speaks to this claim (applicability), and who has stood behind the judgment. Confidence and conclusion are never merged into one word.
UA-IMM — Unassisted performance, immediate · AI at assessment: no · immediate post
Field Assembly certainty: Supported (not a GRADE rating — what our scale means) · Conclusion: Harm · Applicability: Partial · AI: two passes agreed
Rated against: any improvement vs comparator (claim direction); observed direction is the reverse
-17% on the unassisted exam (GPT Base vs control), P<0.05; the claim's direction is unsupported and the observed direction is harm
| Domain | Judgment and reasoning |
|---|---|
| Risk of bias | not serious · single HIGH cluster RCT; D1 two-pass split (reporting detail) published as unsettled; both readings leave the design strong |
| Inconsistency | not serious · single study; no conflicting unassisted secondary evidence in the collection |
| Indirectness | serious · one school network, one country, Fall-2023 GPT-4, same-session outcome only |
| Imprecision | serious · one trial, one setting; precise within it (SE 0.022) but unreplicated |
| Reporting and publication bias | not serious · not assessable for a single preregistered trial |
Population: secondary students (grades 9-11, one Turkish school network)
Comparator: business-as-usual practice (textbooks/notes)
Assessed by software, two independent passes · search: seed/corpus.yaml (verified inventory, proposal section 7) + data/searches/ · method: EVIDENCE-MODEL.md v2 + fa-certainty-scale v2
1 study in this body. 2 further records were considered and left out:
- ChatGPT's impact on student learning outcomes: a meta-analysis of 35 experimental studies
pools assisted and unassisted outcomes
- Studying the effect of AI Code Generators on Supporting Novice Learners in Introductory Programming
intervention class mismatch (code generator, not chatbot)
How this assessment has changed
- Not supportedSep 18, 2026 · Evidence stable
initial assessed status (Q-009, owner-accepted IN-018)
- Not yet assessedSep 18, 2026 · Evidence unclear
initial curated status