Learning gains from AI-assisted practice persist on delayed (one week or longer) retention tests taken without AI.
- Current status
- Insufficient evidenceEvidence stable · what the nine statuses mean
- Why
- One underpowered null (N=33, ~43% power). The claim is effectively untested at meaningful precision.
- What would change it
- Delayed unassisted retention arms in university studies at meaningful sample sizes.
- Linked evidence
- 1 link · 1 contradicts
- Last updated
- Sep 18, 2026
The status is the collection's own judgment on its nine-label scale — not a certainty rating, not a GRADE level, and not advice. Every status change is dated, reasoned, and kept below under History.
The evidence, sorted by what it shows
Contradicts (1)
- ChatGPT as a Learning Tool for Medical Students: Results From a Randomized Controlled TrialRandomized trial · 2025 · Study tier 3 · Secondary different · University partial
The only delayed unassisted university measure: no demonstrated persistence at one week; severely underpowered. Rests on finding F5, judged at assessment v2.
Certainty by outcome
A claim can be broken into separate bodies of evidence — one per outcome, split by whether AI was available at assessment and when the outcome was measured. Each body carries four separate judgments: how confident we are (certainty), what the evidence points to (the conclusion), how directly it speaks to this claim (applicability), and who has stood behind the judgment. Confidence and conclusion are never merged into one word.
UA-DEL — Unassisted retention at >=1 week · AI at assessment: no · delayed post
Field Assembly certainty: Weak (not a GRADE rating — what our scale means) · Conclusion: Not estimable · Applicability: Partial · AI: two passes agreed
Rated against: persistent advantage
One underpowered null; nothing estimable
| Domain | Judgment and reasoning |
|---|---|
| Risk of bias | serious · LOW study; unsettled D4 |
| Inconsistency | not serious · single result |
| Indirectness | serious · lookup-style episode, not sustained practice |
| Imprecision | very serious · N=33 vs 72-132 needed for 80% power |
| Reporting and publication bias | not serious · not assessable |
Population: first-year medical students
Comparator: non-AI resources
Assessed by software, two independent passes · search: seed/corpus.yaml (verified inventory, proposal section 7) + data/searches/ · method: EVIDENCE-MODEL.md v2 + fa-certainty-scale v2
1 study in this body.
How this assessment has changed
- Insufficient evidenceSep 18, 2026 · Evidence stable
initial assessed status (Q-009, owner-accepted IN-018)
- Not yet assessedSep 18, 2026 · Evidence unclear
initial curated status