Open questions
What the research has not yet answered, why not, and which studies come closest. A question stays open until evidence capable of answering it exists — thin evidence gets a description here, not an unearned verdict.
Does generative AI use during learning impair or improve transfer to unfamiliar problem types?
Most studies do not test transfer at all; the few that do report no significant differences. Absence of transfer testing is not absence of effect.
Closest evidence so far
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceRandomized trial · 2025 · Study tier 2 · Secondary different · University direct
The only transfer test in the collection: 10-item near-transfer, no group differences (p=.996) at N~107; High RoB from missing-data accounting.
Does sustained reliance on generative AI erode previously acquired skills?
Longitudinal within-person evidence is absent; existing signals are cross-sectional correlations and scaffold-withdrawal designs.
Closest evidence so far
- Impact of AI assistance on student agencyRandomized trial · 2024 · Study tier 2 · Secondary different · University direct
Scaffold withdrawal reduced peer-review quality on the intervention's own metrics (High RoB: circular measurement); pre-LLM system.
Does teacher supervision change what learners retain from AI use, holding the tool constant?
Positive supervised programs exist, but supervision is confounded with structure and materials; no study randomizes supervision itself.
Closest evidence so far
- From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in NigeriaRandomized trial · 2025 · Preliminary · Secondary direct · University different · Working paper
Teacher-supervised program, positive endline — supervision confounded with added instruction time.
- Effective and Scalable Math Support: Experimental Evidence on the Impact of an AI Math Tutor in GhanaCluster-randomized trial · 2024 · Study tier 3 · Secondary partial · University different · Partly vendor funded
Monitored study-hall tutoring, positive growth — High RoB; assessment conditions not explicit.
Do findings from university students generalize to secondary students, and vice versa?
Meta-analytic moderator tests find no education-level difference, but over mostly-university samples; paired replications are missing.
Closest evidence so far
- ChatGPT's impact on student learning outcomes: a meta-analysis of 35 experimental studiesMeta-analysis · 2026 · Study tier 3 · Secondary partial · University partial
Education level not a significant moderator (p=0.527) — fragile (one-study exclusion flips it), over mostly-university samples pooling assisted/unassisted outcomes.
Does offering AI tools change participation and engagement, separately from any effect on learning?
One MOOC-scale randomized offer reduced exam participation on average while adopters differed; mechanisms unresolved.
Closest evidence so far
- The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement But May Increase Adopters' Exam PerformancesRandomized trial · 2025 · Study tier 2 · Secondary different · University partial · Partly vendor funded
Randomized offer reduced exam participation (-4.3pp, survives correction) and week-6 homework; adopter gains are self-selected.