AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

How Generative AI Influences Learning Outcomes in Programming Education: A Three-level Bayesian Meta-analysis

Design
Meta-analysis
Subject and task
Programming
GenAI use in programming education; three-level Bayesian meta-analysis of 35 empirical studies with 131 effect sizes, published 2022-2025
Exposure: not_reported
Population match
Secondary students: Unclear · University students: Unclear
Study tier
Not ratedNOT_RATED (owner ruling IN-021, 2026-09-30). Rating a synthesis requires appraising its search comprehensiveness, inclusion criteria, dependence handling, primary-study risk-of-bias integration, and publication-bias assessment; all of that is in the paywalled body, so a rating would imply an appraisal that could not be performed. The accessible facts cut both ways: a rigorous venue (Educational Psychology Review), a three-level Bayesian model that is an appropriate device for the 131 dependent effect sizes, deposited analysis code, no funding, no conflicts - and, against that, no visible registration and no verifiable methods. Revisable on legitimate full-text access (NEW_EVIDENCE).
Assessment
Version 1 · AI: two passes agreed · Oct 1, 2026

Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.

A peer-reviewed three-level Bayesian meta-analysis (Wu & Ouyang 2026, Educational Psychology Review) of 35 empirical studies with 131 effect sizes on generative AI in programming education, published 2022-2025. The abstract reports credible small-to-medium positive effects of GenAI on all four outcome categories — AI-assisted programming outcomes, independent programming outcomes, higher-order skills, and motivational-emotional outcomes — with no credible difference among categories and no credible moderation by educational context, instructional design, or GenAI system design; a few exploratory pairwise differences were credible for motivational-emotional outcomes only. The authors conclude the conditions that strengthen or weaken these effects remain undetermined. Only the abstract and declarations are publicly accessible, so methods quality and effect sizes could not be verified; this assessment is labeled abstract-based. This is the second meta-analysis in this corpus's domain (Rule 13 context); two-pass extraction here checks reliability of the reading, not the correctness of the meta-analysis itself.

01Findings

What the study measured

Findings

A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.

Finding and conditionsResult
F1 · pooled effect of GenAI on AI-assisted programming outcomes (AIPO)
vs With AI at assessment — assisted performance, not learning evidencePrespecification unclear
Benefit
credible small-to-medium positive effects (abstract's phrase for all four categories; no numbers in the accessible text)
Coded "yes" from how the paper defines the category: AIPO is named "AI-assisted programming outcomes", so by definition the pooled measures assess the learner-plus-tool system. What the abstract establishes: the category label only. What it does not establish: which instruments feed the pool, their outcome types, timing, or how primary studies verified assistance conditions. Under the gating rule this pooled finding can speak only to assisted performance, never to a learning claim.
F2 · pooled effect of GenAI on independent programming outcomes (IPO)
vs Measured without AIPrespecification unclear
Benefit
credible small-to-medium positive effects (abstract's phrase for all four categories; no numbers in the accessible text)
Coded "no" from how the paper defines the category: IPO is named "independent programming outcomes", i.e. performance without the AI — the category most relevant to this site's question. What the abstract establishes: the category label only. What it does not establish: how primary studies ensured AI was unavailable at assessment, the outcome types, timing (immediate vs delayed), or transfer distance of the pooled measures.
F3 · pooled effect of GenAI on higher-order skills (HOS)
vs AI availability at assessment not reported — cannot carry a learning claimPrespecification unclear
Benefit
credible small-to-medium positive effects (abstract's phrase for all four categories; no numbers in the accessible text)
The category name "higher-order skills" does not state whether the constituent measures (e.g. computational thinking, problem-solving, critical thinking per the starred reference titles) were assessed with or without AI access, so availability is not_reported rather than guessed (Rule 1).
F4 · pooled effect of GenAI on motivational-emotional outcomes (MEO)
vs AI availability at assessment not reported — cannot carry a learning claimPrespecification unclear
Benefit
credible small-to-medium positive effects (abstract's phrase for all four categories; no numbers in the accessible text)
Motivational-emotional outcomes are proxy outcomes under Rule 3 and can never carry a learning claim. They are typically self-report instruments, but the abstract does not state the instruments or survey conditions, so availability defaults to not_reported (IN-006 spirit).
F5 · between-category comparison of pooled effects (AIPO vs IPO vs HOS vs MEO)
vs AI partly available at assessment — cannot carry a learning claimPrespecification unclear
Null
no credible difference among categories was detected (no numbers in the accessible text)
Coded "partial" because this contrast spans categories whose availability codings differ (AIPO yes, IPO no, HOS/MEO not_reported) — it aggregates assisted and unassisted components (IN-005 by analogy). Note this is an absence-of-detected-difference result, not established equivalence (Rule 6): no equivalence margin is stated in the abstract, and the authors themselves warn the corpus may carry limited statistical information.
F6 · moderator analyses: educational contexts, instructional designs, and GenAI system designs, across all four outcome categories
vs AI availability at assessment not reported — cannot carry a learning claimPrespecification unclear
Null
did not detect credible moderation in any outcome category (no numbers in the accessible text)
Recorded as the paper reports it: no credible moderation detected. This is "no demonstrated moderation", not "evidence of no moderation" (Rule 11): the authors state the corpus is small and unevenly distributed across moderator levels, so the absence of detected moderation may reflect limited statistical information rather than equivalence across conditions. Moderator level definitions are in the paywalled body and unverifiable.
F7 · exploratory pairwise moderator contrasts for MEO under specific educational contexts and strategy-training conditions
vs AI availability at assessment not reported — cannot carry a learning claimNot prespecified
Unclear
a few exploratory pairwise differences were credible for MEO (which conditions were favored, and by how much, is not stated in the accessible text)
Prespecified coded "no" because the abstract itself labels these contrasts "exploratory". Direction is unclear: the abstract says differences were credible but not which levels were favored. A credible exploratory subgroup contrast inside an otherwise no-moderation result is recorded as exactly that (Rule 7). MEO is a proxy outcome category in any case.
Limitations
Abstract-based assessment; the article body is paywalled. Unverifiable from the accessible text: search strategy and databases; inclusion and exclusion criteria; how the four outcome categories (AIPO, IPO, HOS, MEO) are operationally defined beyond their names, and which primary measures feed each; risk-of-bias handling; publication-bias methods; effect magnitudes (only the qualitative phrase "credible small-to-medium positive effects" is available — no numbers, no credible intervals); moderator level definitions; primary-study designs, populations, and assessment conditions. The paper's own stated limitation: the corpus is small and unevenly distributed across moderator levels, so the absence of detected moderation may reflect limited statistical information rather than equivalence across conditions. Capture-level discrepancy: 34 starred entries were transcribed from the public reference list against the stated 35 reviewed studies; one starred entry may have been missed in extraction or is listed only in the supplement (flagged in the capture for the review pass).

Who was studied

Level
not_reported
Ages
not_reported
Country
not_reported
Prior knowledge
not_reported

Methodological notes

Extracted from: https://link.springer.com/article/10.1007/s10648-026-10211-x Access: public page only (abstract, references, declarations); article body paywalled

02Funding

Funding and conflicts

Funding
No funding was received (stated).
Vendor funded
No
Vendor
None identified
Notes
Authors declare no competing interests (College of Education, Zhejiang University); no vendor funding or affiliation apparent in the accessible declarations.

Funding is shown on every study and never used to score it.

03Claims

Claims this study bears on

04Source

The source, as retrieved

Abstract

Generative AI (GenAI) has shown promise in programming education, yet empirical evidence remains inconsistent due to varied pedagogical and technical contexts. This research addressed this inconsistency through a three-level Bayesian meta-analysis of 35 empirical studies with 131 effect sizes published between 2022 and 2025. Learning outcomes were classified into AI-assisted programming outcomes (AIPO), independent programming outcomes (IPO), higher-order skills (HOS), and motivational-emotional outcomes (MEO). Results demonstrated credible small-to-medium positive effects of GenAI on all four outcome categories, and no credible difference among categories was detected. The moderator analyses did not detect credible moderation by educational contexts, instructional designs, or GenAI system designs in any outcome category. A few exploratory pairwise differences were credible for MEO under specific educational contexts and strategy-training conditions. Within the current evidence, GenAI shows consistent positive average effects, whereas the conditions that strengthen or weaken these effects remain undetermined. Because the corpus is small and unevenly distributed across moderator levels, the absence of detected moderation may reflect limited statistical information rather than equivalence across conditions. Larger and more balanced studies with more complete reporting of implementation details are needed to identify these conditions.

Where this record came from

SourceRetrievedIdentifier
monthly-sweepOct 1, 2026doi:10.1007/s10648-026-10211-x
Screened in 2026-09-30; metadata verified on the publisher page the same day.