ChatGPT's impact on student learning outcomes: a meta-analysis of 35 experimental studies
- Design
- Meta-analysis · 4193 participants
- Subject and task
- Other
Pools student learning outcomes from 35 experimental and quasi-experimental ChatGPT studies across many subjects (physics, chemistry, English, mathematics, computer science, teaching skills, literature, science, interdisciplinary/STEM, and others). Outcomes span a cognitive dimension (learning achievement, critical thinking, problem-solving, creative thinking, social skills) and a non-cognitive dimension (learning interest, self-efficacy, learning engagement, learning motivation), the latter largely self-report constructs.
Exposure: Included interventions coded into three bands: less than one month (n=7), one to three months (n=18), and more than three months (n=7). Individual studies cited in the review ranged from ~10 days to 15-16 weeks; the paper notes definitions of "long-term" varied from 3 to 12 months across studies. - Population match
- Secondary students: Partial · University students: Partial
- Attrition
- not_reported
- Study tier
- Study tier 3Multi-database search (Web of Science, Wiley, SpringerLink, ProQuest Education, Elsevier ScienceDirect, CNKI) with PRISMA screening, dual independent coding (kappa = 0.851), a 7-point quality checklist (mean 6.11/7), sensitivity analyses, and a three-method publication-bias assessment are genuine strengths. However, heterogeneity is very high (I2 = 91.4%, Q = 409.067, p < 0.01) and largely unexplained; 134 effect sizes from 35 studies are analyzed in CMA 3.0 without meta-regression or three-level modeling (the authors themselves flag this), so effect-size dependency is not handled. Critically for our question, the meta-analysis pools post-test outcomes without ever distinguishing whether ChatGPT was available to students during outcome assessment, and it mixes objective achievement measures with self-report non-cognitive scales in the overall g = 0.670. Quasi-experiments are pooled with experiments, and the search window (Nov 2022 - Jun 2024) captures only early, mostly short studies.
- Assessment
- Version 2 · AI: two passes agreed · Sep 18, 2026
Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.
This meta-analysis pooled 35 experimental and quasi-experimental studies (4193 students, 134 effect sizes) published between November 2022 and June 2024 and found a moderate overall benefit of ChatGPT on student learning outcomes (g = 0.670, 95% CI 0.495-0.844), with larger effects on cognitive outcomes (g = 0.872) than non-cognitive ones (g = 0.539). Effects were stronger for interventions longer than three months, in traditional (teacher-centred) instruction, and in some subjects (physics, chemistry, English), while education level and knowledge type did not significantly moderate results. Heterogeneity was very high (I2 = 91.4%) and the paper does not distinguish whether outcomes were measured while students still had ChatGPT available, so assisted and unassisted performance appear to be pooled together.
What the study measured
Findings
A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.
| Finding and conditions | Result |
|---|---|
| F1 · pooled effect on overall student learning outcomes (35 studies, 134 effect sizes) vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear | Benefit g=0.670 (random-effects) 95% CI [0.495, 0.844] p=0.000 Pooled across objective and self-report outcomes and across cognitive and non-cognitive dimensions; computed from post-test scores of experimental vs control groups with no report of whether ChatGPT was available at assessment. No preregistration/protocol registration reported (PRISMA followed for selection only). |
| F10 · pooled effect on self-efficacy (n=9 effects) vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear | Benefit g=0.467 not reported in text p<0.05 Self-report questionnaire construct. |
| F11 · pooled effect on learning engagement (n=7 effects) vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear | Benefit g=0.660 not reported in text p<0.05 Self-report questionnaire construct. |
| F12 · pooled effect on learning motivation (n=10 effects) vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear | Benefit g=0.409 not reported in text p<0.05 Self-report questionnaire construct. |
| F13 · moderator test: academic subject (subgroup g from physics 1.951 down to science 0.169) vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear | Mixed between-group Q=93.933; physics g=1.951, chemistry g=1.276, English g=0.994, mathematics g=0.655, computer science g=0.436, teaching skills g=0.421, literature g=0.376 (all p<0.05); interdisciplinary g=0.16 and science g=0.169 (both p>0.05) not reported in text p=0.000 (between-group) Subject was one of five moderators named in RQ2; significant between-subject differences. Interdisciplinary and science subgroups nonsignificant with very small k. |
| F15 · moderator test: educational level (primary, secondary, higher education) vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear | Null between-group Q=0.399; primary g=0.824 (p=0.114, k=2, ns), secondary g=0.847 (p<0.05), higher education g=0.744 (p<0.05) primary subgroup 95% CI [-0.198, 1.846] (from sensitivity analysis); others not reported in text p=0.527 (between-group) No significant between-level differences, but the primary subgroup has only k=2 and low power. Sensitivity analysis: removing 'Study B' made between-group differences significant (g rose to 1.329, 95% CI [0.896, 1.762], Qbet=7.884, p=0.005), so this null moderator result is fragile. |
| F17 · moderator test: type of knowledge (declarative vs procedural) vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear | Null between-group Q=1.599; declarative knowledge g=0.903 (n=13), procedural knowledge g=0.705 (n=22), both p<0.05; no significant between-type difference not reported in text p=0.206 (between-group) Both subgroup effects significant and positive; declarative numerically larger but the moderator test itself is nonsignificant. |
| F2 · pooled effect on cognitive learning outcomes (creativity, social skills, problem-solving, achievement, critical thinking) vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear | Benefit g=0.872 not reported in text p<0.05 Cognitive vs non-cognitive difference significant (Q=7.337, p=0.007). Subgroup CIs reported only in Table 2 (not rendered in accessible text). |
| F3 · pooled effect on non-cognitive learning outcomes (interest, engagement, motivation, self-efficacy) vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear | Benefit g=0.539 not reported in text p<0.05 Non-cognitive constructs (interest, engagement, motivation, self-efficacy) are questionnaire/self-report measures in the primary studies; treat as self-report-derived pooled outcome. |
| F4 · pooled effect on learning achievement (n=29 effects) vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear | Benefit g=0.876 not reported in text p<0.05 Closest sub-outcome to objective test performance; paper does not say whether achievement post-tests were taken with or without ChatGPT access. |
| F5 · pooled effect on critical thinking (n=8 effects) vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear | Benefit g=1.008 not reported in text p<0.05 Largest cognitive sub-dimension effect; measurement instruments in primary studies not characterized in the meta text. |
| F6 · pooled effect on problem-solving ability (n=10 effects) vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear | Benefit g=0.933 not reported in text p<0.05 Instruments not characterized at meta level. |
| F7 · pooled effect on creative thinking (n=7 effects) vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear | Benefit g=0.633 not reported in text p<0.05 Instruments not characterized at meta level. |
| F8 · pooled effect on social skills (n=2 effects) vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear | Benefit g=0.481 not reported in text p<0.05 Only 2 effect sizes; interpret cautiously. |
| F9 · pooled effect on learning interest (n=3 effects) vs AI availability at assessment not reported — cannot carry a learning claimTransfer not assessedPrespecification unclear | Null g=0.915 not reported in text p=0.085 Large point estimate but nonsignificant (only 3 effects); self-report questionnaire construct. |
- Limitations
- Source-stated: only published peer-reviewed literature (possible publication bias toward positive findings); moderator coverage incomplete (intervention setting, ChatGPT's role, teacher's role not examined); primary school subgroup has only k = 2 studies with limited statistical power; limited search timeframe; focus on basic cognitive/non-cognitive outcomes with little attention to AI literacy, computational thinking, or ethics; authors recommend meta-regression or three-level meta-analysis for future rigor. Observed: the paper never reports whether pooled post-test outcomes were measured with or without ChatGPT access, so assisted and unassisted performance are presumably pooled together; the overall estimate mixes objective test scores with self-report questionnaire outcomes (motivation, engagement, self-efficacy, interest); Begg's test is borderline (p = 0.079); one sensitivity exclusion ("Study B") changed a subgroup effect substantially (to g = 1.329 with Qbet = 7.884, p = 0.005), suggesting fragility in the educational-level subgroup conclusion; quasi-experimental designs are pooled with randomized ones without subgroup separation by design.
Who was studied
- Level
- mixed
- Ages
- Not reported as ages; grade bands: primary education (k=2), secondary education (k=7), higher education (k=26).
- Country
- Multi-country; not itemized. 28 English-language studies (80%) and 7 Chinese-language studies (20%); Chinese literature deliberately added to cover ChatGPT-restricted contexts.
- Prior knowledge
- Not reported at the pooled level; discussion notes inexperienced students struggle to prompt ChatGPT effectively.
- Selection
- Included studies had to (1) use ChatGPT as a direct or supported learning tool with focus on student learning outcomes; (2) use experimental or quasi-experimental designs; (3) have experimental and control groups or pre-/post-test assessments; (4) sample mainstream primary, secondary, or college students (special education students, adult learners excluded); (5) report sufficient data to compute effect sizes. Search window November 2022 - June 2024; 2038 deduplicated records screened to 35 included articles.
Methodological notes
Extracted from: https://www.nature.com/articles/s41599-026-07019-z Access: Open-access full text retrieved from Nature (Humanities and Social Sciences Communications 13:684, published 26 March 2026). Full HTML article text was accessible, including Methods, Results (overall effect, sensitivity analysis, publication bias, heterogeneity, moderator analyses), Discussion, Limitations, and funding/competing-interest statements. Tables 1-4 and Supplementary Tables S1-S5 render as links only; all statistics below are taken from the running text, which reports the key values. Sample-size note: k = 35 studies, 134 effect sizes, 4193 pooled participants.
How much each result can be relied on
Risk of bias is judged per result, not per paper: the same study can report one outcome at low risk and another at high. Two reviewers assess each result independently against a written guide, and every domain judgment below carries its reasoning and the sentences it rests on.
Risk of bias is not the same as whether the numbers work. A trial can be well conducted and still print a standard deviation that is arithmetically impossible, or a table whose sign contradicts its own text. Risk of bias has no question for that, so where it was found it is recorded separately below, under Source integrity.
These words are not quality scores. On the RoB 2 scale, High risk of bias is the worst rating a result can receive and Low risk of bias is the best — the opposite of how the same words read on some other scales. Levels are always written out in full here for that reason.
POOLED
Critically low confidence in the review · AMSTAR 2 · Two reviewers agreed on every domain
How the overall was reached: Five critical flaws under AMSTAR 2: no registered or a-priori protocol (Item 2); no list of excluded studies with justifications in article or supplement (Item 7); risk of bias in included studies not satisfactorily assessed — a 7-item summary quality score omitting allocation concealment, blinding, attrition, and selective reporting (Item 9); inappropriate meta-analytic handling of 134 dependent effect sizes from 35 studies pooled as independent in CMA 3.0, inflating the precision of g=0.670 and distorting moderator tests, a dependency problem the authors themselves defer to future three-level meta-analysis (Item 11); and no consideration of study-level RoB when interpreting the pooled result (Item 13). More than one critical flaw mandates Critically low; publication bias (Item 15, Yes) and heterogeneity handling (Item 14, Yes) do not offset this.
| Domain | Judgment and reasoning |
|---|---|
| overall_confidence · Critically low confidence in the review | Item 2 (protocol): No — fails by ABSENCE: no protocol, PROSPERO/registry entry, or a-priori methods commitment is mentioned anywhere in the article or supplementary materials (the word 'protocol' appears only in Nature's footer link to protocols.io); PRISMA is cited only for study selection reporting. CRITICAL FLAW. Item 4 (comprehensive search): Partial Yes — six databases including CNKI were searched with named keywords and a stated date window (Nov 2022-Jun 2024), and inclusion of Chinese-language literature is justified; but no full per-database Boolean strategy is reproduced, no reference-list checking, trial-registry, grey-literature, or expert consultation is reported, and the search 'prioritised papers published in core journals' — an unjustified quality restriction that compounds the published-only limitation the authors themselves concede. Item 7 (excluded-studies list): No — fails by ABSENCE: only PRISMA flow counts are given (2038 deduplicated records → 463 → 35 included); neither the article nor the supplementary DOCX contains a list of studies excluded at full-text review with per-study justifications (supplement holds only Tables S1-S5: prior meta-analyses comparison, quality checklist, included-study characteristics, forest plot, sensitivity analyses). CRITICAL FLAW. Item 9 (RoB of included studies): No — a 7-item summary quality-score checklist adapted from Montuori et al. (2023) was used (true RCT or matched design, active control, pre-post measures, tool reliability, pre-test equivalence, method clarity, intervention clarity; verified in Supplementary Table S2), reported only as a mean score of 6.11/7. It is a quality scale, not a risk-of-bias assessment: allocation concealment, blinding of participants/assessors, attrition, and selective outcome reporting are not assessed at all, which AMSTAR 2 requires for RCTs/quasi-experiments even at Partial Yes. CRITICAL FLAW. Item 11 (meta-analytical methods, dependency handling): No — random-effects Hedges' g in CMA 3.0 is defensible in itself and the RE model is justified by I²=91.4%, but the analysis pools 134 effect sizes drawn from 35 studies with no handling of statistical dependency whatsoever: no three-level/multilevel model, no robust variance estimation, no within-study averaging is described, and moderator subgroup n's count effect sizes, not studies. Treating dependent effects as independent inflates precision of g=0.670 [0.495, 0.844] and distorts the Q-based moderator tests; the authors effectively concede the flaw by recommending 'three-level meta-analysis' as future work in Limitations. CRITICAL FLAW. Item 13 (RoB in interpretation): No — the Discussion and Conclusion interpret the pooled effect and moderators without reference to risk of bias in the primary studies; the only quality claim ('relatively high quality', mean 6.11/7) rests on the unsatisfactory tool from Item 9, and 'Bias tests found no significant bias' in the Discussion refers to publication bias, not study-level RoB. CRITICAL FLAW. Item 15 (publication bias): Yes — investigated with three methods (funnel plot inspection, classic fail-safe N = 3522 against a 5n+10 criterion, Begg's test z=1.757, p=0.079) and the published-only-literature risk is explicitly discussed in Limitations; a minor caveat is reliance on the low-powered Begg test and the deprecated fail-safe N rather than Egger/trim-and-fill. Non-critical items materially relevant: Item 1 (PICO): Yes. Item 3 (design selection explained): Yes — restriction to experimental/quasi-experimental designs with control groups or pre-post assessment is stated and reasoned. Item 5 (duplicate selection): Yes — two researchers screened independently with consensus resolution. Item 6 (duplicate extraction): Yes — dual independent coding, kappa 0.851. Item 8 (included-study detail): Partial Yes — Supplementary Table S3 tabulates study characteristics. Item 10 (funding of included studies): No — not reported (non-critical weakness). Item 12 (impact of RoB on synthesis): No — leave-one-out and subgroup sensitivity analyses were run, but no analysis of whether study quality/RoB affects the pooled estimate (non-critical weakness). Item 14 (heterogeneity discussed): Yes — Q=409.067, I²=91.444% reported, RE model adopted, moderators investigated and discussed. Item 16 (review-author COI): Yes — 'The authors declare no competing interests.'This study employed a multi-database comprehensive retrieval strategy utilising authoritative literature databases, including the Web of Science, Wiley Online Library, SpringerLink, ProQuest Education, Elsevier Science Direct, and CNKI. To ensure the quality and reliability of the literature, this study prioritised papers published in core journals. The literature search was limited to articles published between November 2022 and June 2024. |
Funding and conflicts
- Funding
- General Project of the National Social Science Fund (Education), Grant No. BIA250124, and the Scientific Research Fund of Hunan Provincial Education Department, Grant No. 25A0349 (Chinese government/academic funding).
- Vendor funded
- No
- Vendor
- None identified
- Notes
- The authors declare no competing interests. No AI-industry involvement disclosed.
Funding is shown on every study and never used to score it.
Claims this study bears on
- Unrestricted use of a general-purpose LLM chatbot during practice improves secondary students' subsequent performance on assessments taken without AI.Consistent, but doesn't test the claim · rests on finding F1
Pooled g=0.670 does not separate assisted from unassisted outcomes and spans intervention classes; AMSTAR 2 Critically low.
- Unrestricted use of a general-purpose LLM chatbot during practice improves university students' subsequent performance on assessments taken without AI.Consistent, but doesn't test the claim · rests on finding F1
Assisted/unassisted pooling plus Critically-low appraisal confidence; higher-education subgroup shares the same problems.
The source, as retrieved
Abstract
No abstract retrieved.
Where this record came from
| Source | Retrieved | Identifier |
|---|---|---|
| seed | Sep 18, 2026 | link first ingestion |