AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

ChatGPT's impact on student learning outcomes: a meta-analysis of 35 experimental studies

Design
Meta-analysis · 4193 participants
Subject and task
Other
Pools student learning outcomes from 35 experimental and quasi-experimental ChatGPT studies across many subjects (physics, chemistry, English, mathematics, computer science, teaching skills, literature, science, interdisciplinary/STEM, and others). Outcomes span a cognitive dimension (learning achievement, critical thinking, problem-solving, creative thinking, social skills) and a non-cognitive dimension (learning interest, self-efficacy, learning engagement, learning motivation), the latter largely self-report constructs.
Exposure: Included interventions coded into three bands: less than one month (n=7), one to three months (n=18), and more than three months (n=7). Individual studies cited in the review ranged from ~10 days to 15-16 weeks; the paper notes definitions of "long-term" varied from 3 to 12 months across studies.
Population match
Secondary students: Partial · University students: Partial
Attrition
not_reported
Study tier
Study tier 3Multi-database search (Web of Science, Wiley, SpringerLink, ProQuest Education, Elsevier ScienceDirect, CNKI) with PRISMA screening, dual independent coding (kappa = 0.851), a 7-point quality checklist (mean 6.11/7), sensitivity analyses, and a three-method publication-bias assessment are genuine strengths. However, heterogeneity is very high (I2 = 91.4%, Q = 409.067, p < 0.01) and largely unexplained; 134 effect sizes from 35 studies are analyzed in CMA 3.0 without meta-regression or three-level modeling (the authors themselves flag this), so effect-size dependency is not handled. Critically for our question, the meta-analysis pools post-test outcomes without ever distinguishing whether ChatGPT was available to students during outcome assessment, and it mixes objective achievement measures with self-report non-cognitive scales in the overall g = 0.670. Quasi-experiments are pooled with experiments, and the search window (Nov 2022 - Jun 2024) captures only early, mostly short studies.
Assessment
Version 2 · AI: two passes agreed · Sep 18, 2026

Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.

This meta-analysis pooled 35 experimental and quasi-experimental studies (4193 students, 134 effect sizes) published between November 2022 and June 2024 and found a moderate overall benefit of ChatGPT on student learning outcomes (g = 0.670, 95% CI 0.495-0.844), with larger effects on cognitive outcomes (g = 0.872) than non-cognitive ones (g = 0.539). Effects were stronger for interventions longer than three months, in traditional (teacher-centred) instruction, and in some subjects (physics, chemistry, English), while education level and knowledge type did not significantly moderate results. Heterogeneity was very high (I2 = 91.4%) and the paper does not distinguish whether outcomes were measured while students still had ChatGPT available, so assisted and unassisted performance appear to be pooled together.

01Findings

What the study measured

Findings

A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.

Finding and conditionsResult
F1 · pooled effect on overall student learning outcomes (35 studies, 134 effect sizes)
vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear
Benefit
g=0.670 (random-effects)
95% CI [0.495, 0.844]
p=0.000
Pooled across objective and self-report outcomes and across cognitive and non-cognitive dimensions; computed from post-test scores of experimental vs control groups with no report of whether ChatGPT was available at assessment. No preregistration/protocol registration reported (PRISMA followed for selection only).
F10 · pooled effect on self-efficacy (n=9 effects)
vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear
Benefit
g=0.467
not reported in text
p<0.05
Self-report questionnaire construct.
F11 · pooled effect on learning engagement (n=7 effects)
vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear
Benefit
g=0.660
not reported in text
p<0.05
Self-report questionnaire construct.
F12 · pooled effect on learning motivation (n=10 effects)
vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear
Benefit
g=0.409
not reported in text
p<0.05
Self-report questionnaire construct.
F13 · moderator test: academic subject (subgroup g from physics 1.951 down to science 0.169)
vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear
Mixed
between-group Q=93.933; physics g=1.951, chemistry g=1.276, English g=0.994, mathematics g=0.655, computer science g=0.436, teaching skills g=0.421, literature g=0.376 (all p<0.05); interdisciplinary g=0.16 and science g=0.169 (both p>0.05)
not reported in text
p=0.000 (between-group)
Subject was one of five moderators named in RQ2; significant between-subject differences. Interdisciplinary and science subgroups nonsignificant with very small k.
F15 · moderator test: educational level (primary, secondary, higher education)
vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear
Null
between-group Q=0.399; primary g=0.824 (p=0.114, k=2, ns), secondary g=0.847 (p<0.05), higher education g=0.744 (p<0.05)
primary subgroup 95% CI [-0.198, 1.846] (from sensitivity analysis); others not reported in text
p=0.527 (between-group)
No significant between-level differences, but the primary subgroup has only k=2 and low power. Sensitivity analysis: removing 'Study B' made between-group differences significant (g rose to 1.329, 95% CI [0.896, 1.762], Qbet=7.884, p=0.005), so this null moderator result is fragile.
F17 · moderator test: type of knowledge (declarative vs procedural)
vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear
Null
between-group Q=1.599; declarative knowledge g=0.903 (n=13), procedural knowledge g=0.705 (n=22), both p<0.05; no significant between-type difference
not reported in text
p=0.206 (between-group)
Both subgroup effects significant and positive; declarative numerically larger but the moderator test itself is nonsignificant.
F2 · pooled effect on cognitive learning outcomes (creativity, social skills, problem-solving, achievement, critical thinking)
vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear
Benefit
g=0.872
not reported in text
p<0.05
Cognitive vs non-cognitive difference significant (Q=7.337, p=0.007). Subgroup CIs reported only in Table 2 (not rendered in accessible text).
F3 · pooled effect on non-cognitive learning outcomes (interest, engagement, motivation, self-efficacy)
vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear
Benefit
g=0.539
not reported in text
p<0.05
Non-cognitive constructs (interest, engagement, motivation, self-efficacy) are questionnaire/self-report measures in the primary studies; treat as self-report-derived pooled outcome.
F4 · pooled effect on learning achievement (n=29 effects)
vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear
Benefit
g=0.876
not reported in text
p<0.05
Closest sub-outcome to objective test performance; paper does not say whether achievement post-tests were taken with or without ChatGPT access.
F5 · pooled effect on critical thinking (n=8 effects)
vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear
Benefit
g=1.008
not reported in text
p<0.05
Largest cognitive sub-dimension effect; measurement instruments in primary studies not characterized in the meta text.
F6 · pooled effect on problem-solving ability (n=10 effects)
vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear
Benefit
g=0.933
not reported in text
p<0.05
Instruments not characterized at meta level.
F7 · pooled effect on creative thinking (n=7 effects)
vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear
Benefit
g=0.633
not reported in text
p<0.05
Instruments not characterized at meta level.
F8 · pooled effect on social skills (n=2 effects)
vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear
Benefit
g=0.481
not reported in text
p<0.05
Only 2 effect sizes; interpret cautiously.
F9 · pooled effect on learning interest (n=3 effects)
vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear
Null
g=0.915
not reported in text
p=0.085
Large point estimate but nonsignificant (only 3 effects); self-report questionnaire construct.
Limitations
Source-stated: only published peer-reviewed literature (possible publication bias toward positive findings); moderator coverage incomplete (intervention setting, ChatGPT's role, teacher's role not examined); primary school subgroup has only k = 2 studies with limited statistical power; limited search timeframe; focus on basic cognitive/non-cognitive outcomes with little attention to AI literacy, computational thinking, or ethics; authors recommend meta-regression or three-level meta-analysis for future rigor. Observed: the paper never reports whether pooled post-test outcomes were measured with or without ChatGPT access, so assisted and unassisted performance are presumably pooled together; the overall estimate mixes objective test scores with self-report questionnaire outcomes (motivation, engagement, self-efficacy, interest); Begg's test is borderline (p = 0.079); one sensitivity exclusion ("Study B") changed a subgroup effect substantially (to g = 1.329 with Qbet = 7.884, p = 0.005), suggesting fragility in the educational-level subgroup conclusion; quasi-experimental designs are pooled with randomized ones without subgroup separation by design.

Who was studied

Level
mixed
Ages
Not reported as ages; grade bands: primary education (k=2), secondary education (k=7), higher education (k=26).
Country
Multi-country; not itemized. 28 English-language studies (80%) and 7 Chinese-language studies (20%); Chinese literature deliberately added to cover ChatGPT-restricted contexts.
Prior knowledge
Not reported at the pooled level; discussion notes inexperienced students struggle to prompt ChatGPT effectively.
Selection
Included studies had to (1) use ChatGPT as a direct or supported learning tool with focus on student learning outcomes; (2) use experimental or quasi-experimental designs; (3) have experimental and control groups or pre-/post-test assessments; (4) sample mainstream primary, secondary, or college students (special education students, adult learners excluded); (5) report sufficient data to compute effect sizes. Search window November 2022 - June 2024; 2038 deduplicated records screened to 35 included articles.

Methodological notes

Extracted from: https://www.nature.com/articles/s41599-026-07019-z Access: Open-access full text retrieved from Nature (Humanities and Social Sciences Communications 13:684, published 26 March 2026). Full HTML article text was accessible, including Methods, Results (overall effect, sensitivity analysis, publication bias, heterogeneity, moderator analyses), Discussion, Limitations, and funding/competing-interest statements. Tables 1-4 and Supplementary Tables S1-S5 render as links only; all statistics below are taken from the running text, which reports the key values. Sample-size note: k = 35 studies, 134 effect sizes, 4193 pooled participants.

01bRisk of bias

How much each result can be relied on

Risk of bias is judged per result, not per paper: the same study can report one outcome at low risk and another at high. Two reviewers assess each result independently against a written guide, and every domain judgment below carries its reasoning and the sentences it rests on.

Risk of bias is not the same as whether the numbers work. A trial can be well conducted and still print a standard deviation that is arithmetically impossible, or a table whose sign contradicts its own text. Risk of bias has no question for that, so where it was found it is recorded separately below, under Source integrity.

These words are not quality scores. On the RoB 2 scale, High risk of bias is the worst rating a result can receive and Low risk of bias is the best — the opposite of how the same words read on some other scales. Levels are always written out in full here for that reason.

POOLED

Critically low confidence in the review · AMSTAR 2 · Two reviewers agreed on every domain

How the overall was reached: Five critical flaws under AMSTAR 2: no registered or a-priori protocol (Item 2); no list of excluded studies with justifications in article or supplement (Item 7); risk of bias in included studies not satisfactorily assessed — a 7-item summary quality score omitting allocation concealment, blinding, attrition, and selective reporting (Item 9); inappropriate meta-analytic handling of 134 dependent effect sizes from 35 studies pooled as independent in CMA 3.0, inflating the precision of g=0.670 and distorting moderator tests, a dependency problem the authors themselves defer to future three-level meta-analysis (Item 11); and no consideration of study-level RoB when interpreting the pooled result (Item 13). More than one critical flaw mandates Critically low; publication bias (Item 15, Yes) and heterogeneity handling (Item 14, Yes) do not offset this.

DomainJudgment and reasoning
overall_confidence ·
Critically low confidence in the review
Item 2 (protocol): No — fails by ABSENCE: no protocol, PROSPERO/registry entry, or a-priori methods commitment is mentioned anywhere in the article or supplementary materials (the word 'protocol' appears only in Nature's footer link to protocols.io); PRISMA is cited only for study selection reporting. CRITICAL FLAW. Item 4 (comprehensive search): Partial Yes — six databases including CNKI were searched with named keywords and a stated date window (Nov 2022-Jun 2024), and inclusion of Chinese-language literature is justified; but no full per-database Boolean strategy is reproduced, no reference-list checking, trial-registry, grey-literature, or expert consultation is reported, and the search 'prioritised papers published in core journals' — an unjustified quality restriction that compounds the published-only limitation the authors themselves concede. Item 7 (excluded-studies list): No — fails by ABSENCE: only PRISMA flow counts are given (2038 deduplicated records → 463 → 35 included); neither the article nor the supplementary DOCX contains a list of studies excluded at full-text review with per-study justifications (supplement holds only Tables S1-S5: prior meta-analyses comparison, quality checklist, included-study characteristics, forest plot, sensitivity analyses). CRITICAL FLAW. Item 9 (RoB of included studies): No — a 7-item summary quality-score checklist adapted from Montuori et al. (2023) was used (true RCT or matched design, active control, pre-post measures, tool reliability, pre-test equivalence, method clarity, intervention clarity; verified in Supplementary Table S2), reported only as a mean score of 6.11/7. It is a quality scale, not a risk-of-bias assessment: allocation concealment, blinding of participants/assessors, attrition, and selective outcome reporting are not assessed at all, which AMSTAR 2 requires for RCTs/quasi-experiments even at Partial Yes. CRITICAL FLAW. Item 11 (meta-analytical methods, dependency handling): No — random-effects Hedges' g in CMA 3.0 is defensible in itself and the RE model is justified by I²=91.4%, but the analysis pools 134 effect sizes drawn from 35 studies with no handling of statistical dependency whatsoever: no three-level/multilevel model, no robust variance estimation, no within-study averaging is described, and moderator subgroup n's count effect sizes, not studies. Treating dependent effects as independent inflates precision of g=0.670 [0.495, 0.844] and distorts the Q-based moderator tests; the authors effectively concede the flaw by recommending 'three-level meta-analysis' as future work in Limitations. CRITICAL FLAW. Item 13 (RoB in interpretation): No — the Discussion and Conclusion interpret the pooled effect and moderators without reference to risk of bias in the primary studies; the only quality claim ('relatively high quality', mean 6.11/7) rests on the unsatisfactory tool from Item 9, and 'Bias tests found no significant bias' in the Discussion refers to publication bias, not study-level RoB. CRITICAL FLAW. Item 15 (publication bias): Yes — investigated with three methods (funnel plot inspection, classic fail-safe N = 3522 against a 5n+10 criterion, Begg's test z=1.757, p=0.079) and the published-only-literature risk is explicitly discussed in Limitations; a minor caveat is reliance on the low-powered Begg test and the deprecated fail-safe N rather than Egger/trim-and-fill. Non-critical items materially relevant: Item 1 (PICO): Yes. Item 3 (design selection explained): Yes — restriction to experimental/quasi-experimental designs with control groups or pre-post assessment is stated and reasoned. Item 5 (duplicate selection): Yes — two researchers screened independently with consensus resolution. Item 6 (duplicate extraction): Yes — dual independent coding, kappa 0.851. Item 8 (included-study detail): Partial Yes — Supplementary Table S3 tabulates study characteristics. Item 10 (funding of included studies): No — not reported (non-critical weakness). Item 12 (impact of RoB on synthesis): No — leave-one-out and subgroup sensitivity analyses were run, but no analysis of whether study quality/RoB affects the pooled estimate (non-critical weakness). Item 14 (heterogeneity discussed): Yes — Q=409.067, I²=91.444% reported, RE model adopted, moderators investigated and discussed. Item 16 (review-author COI): Yes — 'The authors declare no competing interests.'
This study employed a multi-database comprehensive retrieval strategy utilising authoritative literature databases, including the Web of Science, Wiley Online Library, SpringerLink, ProQuest Education, Elsevier Science Direct, and CNKI.
To ensure the quality and reliability of the literature, this study prioritised papers published in core journals. The literature search was limited to articles published between November 2022 and June 2024.
02Funding

Funding and conflicts

Funding
General Project of the National Social Science Fund (Education), Grant No. BIA250124, and the Scientific Research Fund of Hunan Provincial Education Department, Grant No. 25A0349 (Chinese government/academic funding).
Vendor funded
No
Vendor
None identified
Notes
The authors declare no competing interests. No AI-industry involvement disclosed.

Funding is shown on every study and never used to score it.

03Claims

Claims this study bears on

04Source

The source, as retrieved

Abstract

No abstract retrieved.

Where this record came from

SourceRetrievedIdentifier
seedSep 18, 2026link
first ingestion