AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

The effectiveness of ChatGPT in assisting high school students in programming learning: evidence from a quasi-experimental research

Design
Quasi-experimental study · randomized at class level · 153 participants
Subject and task
Programming
High school programming course in C++; lecture-based instruction with exercises and projects; experimental group additionally used ChatGPT as a virtual teaching assistant (syntax clarification, debugging, conceptual explanations).
Exposure: Six-week experiment, two instructional sessions per week; ChatGPT available to the experimental group only in the second three-week phase (pretests at end of phase 1, posttests after phase 2).
Population match
Secondary students: Direct · University students: Different
Attrition
not_reported
Study tier
Study tier 3Quasi-experimental allocation of six intact classes (three per condition) at a single single-sex school. Although the class allocation is described as random, there are only six clusters and every analysis (MANCOVA/ANCOVA with pretest covariates) is at student level with no cluster-appropriate adjustment, so precision is overstated and the achievement result's p = .048 is fragile (Rule 4: a student-level analysis of a handful of randomized classrooms is a strength problem). Comparison is business-as-usual with the same instructor and identical materials, so the experimental group received ChatGPT as an addition; note the observed direction is harm, which bundled-extra-resources confounding does not explain in the usual direction. Baseline handling is by covariate adjustment; attrition is unreported. The achievement test is researcher/teacher-developed and course-aligned. Peer-reviewed journal article. LOW.
Assessment
Version 3 · AI: two passes agreed · Sep 30, 2026

Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.

In a six-week quasi-experiment in six classes at an all-girls senior high school in northern Taiwan, 77 students used ChatGPT as a programming study aid in C++ for three weeks while 76 continued conventional lessons with the same teacher and materials. Adjusted for pretests, the ChatGPT group came out significantly lower on all three measured outcomes: flow experience, self-efficacy, and learning achievement (the achievement gap was small and at the edge of significance). Interviewed students were split on whether ChatGPT helped. The study's class-level allocation with only six classrooms, analyzed at student level, means its precision is overstated.

01Findings

What the study measured

Arms

ArmWhat it got
A · ChatGPT-assisted programming learning
n = 77
Intervention · teacher supervised
ChatGPT · GPT · Unrestricted chat used as a virtual teaching assistant during classroom practice: syntax clarification, debugging, example provision, code verification, conceptual explanations. · chat interface · not reported · used not_reported (before 2023-09-13, the manuscript receipt date)
B · Conventional lecture-based instruction
n = 76
Control · teacher led
Same instructor, identical materials and weekly assignments within the same curriculum; no ChatGPT.

Findings

A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.

Finding and conditionsResult
F1 · Programming learning achievement (three-task C++ test, 0-100)
ChatGPT-assisted programming learning vs Conventional lecture-based instructionAI availability at assessment not reported — cannot carry a learning claimImmediate post-testNear transferResearcher-developed testPrespecification unclear
Harm
ANCOVA adjusted for pretest: F = 3.96, partial eta-squared = 0.03 (EG lower)
not_reported
.048
Pre/post tests of three tasks each, designed by a course-experienced high school teacher and a professor; scored for accuracy and completeness; course-aligned, so alignment inflation is possible and transfer is not assessed. The paper never states whether ChatGPT was available during the posttest; standard exam conditions are implied but not reported, so the gating dimension stays not_reported. Both groups scored lower at post than pre (difficulty not equated), which the ANCOVA design accommodates for between-group comparison only.
F2 · Flow experience (adapted Pearce et al. 2005 questionnaire)
ChatGPT-assisted programming learning vs Conventional lecture-based instructionAI availability at assessment not reported — cannot carry a learning claimImmediate post-testTransfer not assessedSelf-reportPrespecification unclear
Harm
ANCOVA adjusted for pretest: F = 9.20, partial eta-squared = 0.06 (EG lower)
not_reported
<.01
Self-report scale (Cronbach's alpha 0.88), translated to Chinese and vetted by HS teachers; proxy outcome — cannot carry a learning claim (Rule 3).
F3 · Programming self-efficacy (Pintrich scale as modified by Wang & Lin 2007)
ChatGPT-assisted programming learning vs Conventional lecture-based instructionAI availability at assessment not reported — cannot carry a learning claimImmediate post-testTransfer not assessedSelf-reportPrespecification unclear
Harm
ANCOVA adjusted for pretest: F = 15.78, partial eta-squared = 0.10 (EG lower)
not_reported
<0.01
Self-report scale (Cronbach's alpha 0.86, six-point); proxy outcome — cannot carry a learning claim (Rule 3).
Limitations
Single site, all-girls school (authors note this and frame it as future work on girls' programming education); six clusters analyzed at student level; flow and self-efficacy are self-report; both groups scored lower at posttest than pretest (test difficulty not equated across waves, acknowledged by the authors); ChatGPT model version and dates of use not reported (manuscript received 2023-09-13, so use predates that); whether ChatGPT was available during the posttest is not stated; no registration mentioned.

Who was studied

Level
secondary
Ages
not_reported (second-year senior high students; grade stated, ages not)
Country
Taiwan (northern Taiwan)
Prior knowledge
All participants taught basic programming concepts by the same teacher before the experiment; basic computer competency (tablet, browser, Internet).
Selection
Six general-education classes at a senior high school for girls (single-sex, all female).

Methodological notes

Extracted from: https://www.tandfonline.com/doi/full/10.1080/10494820.2025.2450659 Access: Full text (open access, CC-BY), retrieved 2026-09-30 from the publisher page; saved as data/fulltext-cache/yang-2025-fulltext.txt. Supersedes the abstract-only pass A of 2026-09-18; the HOLD's full-text verification is performed in this pass. Sample-size note: N=153 across six classes; experimental n=77 (three classes), control n=76 (three classes). Analysis ns per outcome are not restated; no participant flow or attrition reporting.

02Funding

Funding and conflicts

Funding
National Science and Technology Council, Taiwan (grants 112-2628-H-A49-001-MY2; 111-2410-H-A49-066-MY3; inter-discipline empowerment project) — public funder.
Vendor funded
No
Vendor
None identified
Notes
Disclosure statement present: "The author declares that they have no conflict of interest." (sic — singular, on a three-author paper).

Funding is shown on every study and never used to score it.

03Claims

Claims this study bears on

  • Unrestricted use of a general-purpose LLM chatbot during practice improves secondary students' subsequent performance on assessments taken without AI.Consistent, but doesn't test the claim · rests on finding F1

    Exact population (secondary), intervention class (unrestricted ChatGPT chat during practice), and outcome domain (researcher-developed achievement test) of the claim, with an adverse direction: the ChatGPT classes scored lower adjusted for pretest (F=3.96, p=.048, partial eta-squared .03). But the paper never states whether ChatGPT was available during the posttest, so under the availability gate (EVIDENCE-MODEL Rule 2) the finding cannot support or contradict the claim — the nie-2025 F4 precedent for gate-failing direct-outcome findings. The class-level allocation of six classrooms analyzed at student level makes the p=.048 fragile besides. If a later source establishes the posttest was unassisted, this link is re-issued as contradicts against that assessment version.

04Source

The source, as retrieved

Abstract

No abstract retrieved.

Where this record came from

SourceRetrievedIdentifier
seedSep 18, 2026link
first ingestion