Unrestricted use of a general-purpose LLM chatbot during practice improves university students' subsequent performance on assessments taken without AI.
- Current status
- Insufficient evidenceEvidence unclear · what the nine statuses mean
- Why
- Evidence leans against the claim — a significant post-use quality deficit vs web search, and an underpowered retention null — but nothing adequately powered tests it directly, and the largest randomized offer cannot (partial AI at its exam). Pass B read the same evidence as NOT_SUPPORTED; the conservative status stands until something powered exists.
- What would change it
- Any adequately sized university RCT with unassisted post-tests and an active comparison.
- Linked evidence
- 5 links · 2 contradicts · 2 consistent, but doesn't test the claim · 1 no direct evidence
- Last updated
- Sep 18, 2026
The status is the collection's own judgment on its nine-label scale — not a certainty rating, not a GRADE level, and not advice. Every status change is dated, reasoned, and kept below under History.
The evidence, sorted by what it shows
Contradicts (2)
- ChatGPT as a Learning Tool for Medical Students: Results From a Randomized Controlled TrialRandomized trial · 2025 · Study tier 3 · Secondary different · University partial
Closed-book retention one week after open-resource use: no significant differences (p=0.118). Weak evidence — n=33, ~43% power by the authors' own analysis, trend numerically favored ChatGPT; weighted by precision, never discarded. Rests on finding F5, judged at assessment v2.
- Cognitive ease at a cost: LLMs reduce mental effort but compromise depth in student scientific inquiryLaboratory experiment · 2024 · Study tier 2 · Secondary different · University direct
Unassisted justification quality significantly worse after 20 minutes of ChatGPT-3.5 inquiry than after Google-search inquiry (p=.001, eta2=.11). Active non-AI comparator; single session; immediate. Rests on finding F4, judged at assessment v2.
Consistent, but doesn't test the claim (2)
- ChatGPT's impact on student learning outcomes: a meta-analysis of 35 experimental studiesMeta-analysis · 2026 · Study tier 3 · Secondary partial · University partial
Assisted/unassisted pooling plus Critically-low appraisal confidence; higher-education subgroup shares the same problems. Rests on finding F1, judged at assessment v2.
- The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement But May Increase Adopters' Exam PerformancesRandomized trial · 2025 · Study tier 2 · Secondary different · University partial · Partly vendor funded
The largest randomized offer cannot test the claim: the unproctored exam had partial AI availability and High-RoB intervention-affected missingness; population match partial (adult global MOOC). Rests on finding F4, judged at assessment v2.
No direct evidence (1)
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceRandomized trial · 2025 · Study tier 2 · Secondary different · University direct
Thematically the closest university evidence (assisted essay gains, unassisted knowledge/transfer nulls), but the chat was task-restricted — not unrestricted general-purpose use (intervention-class lint). Whole-study judgment at assessment v2.
Certainty by outcome
A claim can be broken into separate bodies of evidence — one per outcome, split by whether AI was available at assessment and when the outcome was measured. Each body carries four separate judgments: how confident we are (certainty), what the evidence points to (the conclusion), how directly it speaks to this claim (applicability), and who has stood behind the judgment. Confidence and conclusion are never merged into one word.
UA-ANY — Unassisted performance, immediate and delayed · AI at assessment: no · immediate post
Field Assembly certainty: Weak (not a GRADE rating — what our scale means) · Conclusion: Inconsistent · Applicability: Partial · AI: two passes agreed
Rated against: any improvement vs comparator
No benefit demonstrated; one significant post-use quality deficit and one underpowered retention null trending pro-ChatGPT
| Domain | Judgment and reasoning |
|---|---|
| Risk of bias | serious · a LOW trial (N=33, unsettled D4) and a Some-concerns lab experiment |
| Inconsistency | serious · a significant quality deficit vs an underpowered null trending the other way — the two results disagree in kind |
| Indirectness | serious · single-session/single-quiz designs; neither is sustained practice |
| Imprecision | very serious · N=33 and N=91; the null is underpowered by the authors' own analysis |
| Reporting and publication bias | not serious · not assessable at two studies |
Population: university students (medical; general German university)
Comparator: non-AI resources (institutional materials; Google search)
Assessed by software, two independent passes · search: seed/corpus.yaml (verified inventory, proposal section 7) + data/searches/ · method: EVIDENCE-MODEL.md v2 + fa-certainty-scale v2
2 studies in this body. 3 further records were considered and left out:
- ChatGPT's impact on student learning outcomes: a meta-analysis of 35 experimental studies
pools assisted and unassisted outcomes
- The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement But May Increase Adopters' Exam Performances
partial AI availability at the exam; High-RoB missingness
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance
task-restricted chat; intervention class mismatch
How this assessment has changed
- Insufficient evidenceSep 18, 2026 · Evidence unclear
initial assessed status (Q-009, owner-accepted IN-018)
- Not yet assessedSep 18, 2026 · Evidence unclear
initial curated status