Guardrailed AI tutors (hints, no direct answers, pedagogical prompting) improve secondary students' performance on assessments taken without AI, compared with business-as-usual instruction.
- Current status
- ContradictoryEvidence unclear · what the nine statuses mean
- Why
- A well-powered trial shows no unassisted gain when guardrailed tutoring replaces equivalent practice time; a preliminary supervised after-school program shows +0.24 SD when the program adds instruction. Candidate explanation: what the comparator holds constant (time and materials). No design yet isolates it.
- What would change it
- Randomized trials with active non-AI comparators matched for time and materials.
Delayed unassisted outcomes rather than same-session exams. - Linked evidence
- 3 links · 1 supports · 1 contradicts · 1 consistent, but doesn't test the claim
- Last updated
- Sep 18, 2026
The status is the collection's own judgment on its nine-label scale — not a certainty rating, not a GRADE level, and not advice. Every status change is dated, reasoned, and kept below under History.
The evidence, sorted by what it shows
Supports (1)
- From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in NigeriaRandomized trial · 2025 · Preliminary · Secondary direct · University different · Working paper
English endline (pencil-and-paper, no AI): +0.238 SD for the supervised after-school Copilot program vs no program. Caveats: the comparator bundles no extra instruction time (the program adds it); PRELIMINARY working paper; attrition handling assessed-not-settled (D3). Rests on finding F3, judged at assessment v2.
Contradicts (1)
- Generative AI without guardrails can harm learning: Evidence from high school mathematicsCluster-randomized trial · 2025 · Study tier 1 · Secondary direct · University different
Guardrailed tutor indistinguishable from business-as-usual on the unassisted same-session exam (-0.004) despite +127% assisted practice: a fairly precise null in a large preregistered trial. One site, one subject, immediate only. Rests on finding F4, judged at assessment v2.
Consistent, but doesn't test the claim (1)
- Effective and Scalable Math Support: Experimental Evidence on the Impact of an AI Math Tutor in GhanaCluster-randomized trial · 2024 · Study tier 3 · Secondary partial · University different · Partly vendor funded
Positive growth effect consistent with the claim, but assessment AI-availability is not explicit (gate), population mostly primary (partial), RoB High (differential attrition; unclustered analysis over 11 schools), vendor-affiliated authors. Rests on finding F1, judged at assessment v2.
Certainty by outcome
A claim can be broken into separate bodies of evidence — one per outcome, split by whether AI was available at assessment and when the outcome was measured. Each body carries four separate judgments: how confident we are (certainty), what the evidence points to (the conclusion), how directly it speaks to this claim (applicability), and who has stood behind the judgment. Confidence and conclusion are never merged into one word.
UA-IMM — Unassisted performance after guardrailed tutoring · AI at assessment: no · immediate post
Field Assembly certainty: Weak (not a GRADE rating — what our scale means) · Conclusion: Varies by intervention design · Applicability: Partial · AI: two passes agreed
Rated against: improvement vs business-as-usual
Null when the tutor replaces equivalent practice time; +0.24 SD when a supervised program adds instruction — what the comparator holds constant is the live question
| Domain | Judgment and reasoning |
|---|---|
| Risk of bias | serious · one HIGH trial; one PRELIMINARY working paper with unsettled attrition handling (D3) |
| Inconsistency | serious · a precise null against a matched-time comparator vs a positive effect when the program adds instruction — heterogeneity with a design-shaped candidate explanation, not settled |
| Indirectness | serious · comparator structures differ from the claim's framing; same-session vs six-week outcomes |
| Imprecision | serious · two studies, different constructs |
| Reporting and publication bias | not serious · not assessable |
Population: secondary students (Turkey; Nigeria)
Comparator: business-as-usual (matched-time practice; no program)
Assessed by software, two independent passes · search: seed/corpus.yaml (verified inventory, proposal section 7) + data/searches/ · method: EVIDENCE-MODEL.md v2 + fa-certainty-scale v2
2 studies in this body. 1 further record was considered and left out:
- Effective and Scalable Math Support: Experimental Evidence on the Impact of an AI Math Tutor in Ghana
assessment AI-availability not explicit; High RoB
How this assessment has changed
- ContradictorySep 18, 2026 · Evidence unclear
initial assessed status (Q-009, owner-accepted IN-018)
- Not yet assessedSep 18, 2026 · Evidence unclear
initial curated status