The effectiveness of ChatGPT in assisting high school students in programming learning: evidence from a quasi-experimental research
- Design
- Quasi-experimental study · randomized at class level · 153 participants
- Subject and task
- Programming
High school programming course in C++; lecture-based instruction with exercises and projects; experimental group additionally used ChatGPT as a virtual teaching assistant (syntax clarification, debugging, conceptual explanations).
Exposure: Six-week experiment, two instructional sessions per week; ChatGPT available to the experimental group only in the second three-week phase (pretests at end of phase 1, posttests after phase 2). - Population match
- Secondary students: Direct · University students: Different
- Attrition
- not_reported
- Study tier
- Study tier 3Quasi-experimental allocation of six intact classes (three per condition) at a single single-sex school. Although the class allocation is described as random, there are only six clusters and every analysis (MANCOVA/ANCOVA with pretest covariates) is at student level with no cluster-appropriate adjustment, so precision is overstated and the achievement result's p = .048 is fragile (Rule 4: a student-level analysis of a handful of randomized classrooms is a strength problem). Comparison is business-as-usual with the same instructor and identical materials, so the experimental group received ChatGPT as an addition; note the observed direction is harm, which bundled-extra-resources confounding does not explain in the usual direction. Baseline handling is by covariate adjustment; attrition is unreported. The achievement test is researcher/teacher-developed and course-aligned. Peer-reviewed journal article. LOW.
- Assessment
- Version 3 · AI: two passes agreed · Sep 30, 2026
Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.
In a six-week quasi-experiment in six classes at an all-girls senior high school in northern Taiwan, 77 students used ChatGPT as a programming study aid in C++ for three weeks while 76 continued conventional lessons with the same teacher and materials. Adjusted for pretests, the ChatGPT group came out significantly lower on all three measured outcomes: flow experience, self-efficacy, and learning achievement (the achievement gap was small and at the edge of significance). Interviewed students were split on whether ChatGPT helped. The study's class-level allocation with only six classrooms, analyzed at student level, means its precision is overstated.
What the study measured
Arms
| Arm | What it got |
|---|---|
| A · ChatGPT-assisted programming learning n = 77 | Intervention · teacher supervised ChatGPT · GPT · Unrestricted chat used as a virtual teaching assistant during classroom practice: syntax clarification, debugging, example provision, code verification, conceptual explanations. · chat interface · not reported · used not_reported (before 2023-09-13, the manuscript receipt date) |
| B · Conventional lecture-based instruction n = 76 | Control · teacher led Same instructor, identical materials and weekly assignments within the same curriculum; no ChatGPT. |
Findings
A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.
| Finding and conditions | Result |
|---|---|
| F1 · Programming learning achievement (three-task C++ test, 0-100) ChatGPT-assisted programming learning vs Conventional lecture-based instructionAI availability at assessment not reported — cannot carry a learning claimImmediate post-testNear transferResearcher-developed testPrespecification unclear | Harm ANCOVA adjusted for pretest: F = 3.96, partial eta-squared = 0.03 (EG lower) not_reported .048 Pre/post tests of three tasks each, designed by a course-experienced high school teacher and a professor; scored for accuracy and completeness; course-aligned, so alignment inflation is possible and transfer is not assessed. The paper never states whether ChatGPT was available during the posttest; standard exam conditions are implied but not reported, so the gating dimension stays not_reported. Both groups scored lower at post than pre (difficulty not equated), which the ANCOVA design accommodates for between-group comparison only. |
| F2 · Flow experience (adapted Pearce et al. 2005 questionnaire) ChatGPT-assisted programming learning vs Conventional lecture-based instructionAI availability at assessment not reported — cannot carry a learning claimImmediate post-testTransfer not assessedSelf-reportPrespecification unclear | Harm ANCOVA adjusted for pretest: F = 9.20, partial eta-squared = 0.06 (EG lower) not_reported <.01 Self-report scale (Cronbach's alpha 0.88), translated to Chinese and vetted by HS teachers; proxy outcome — cannot carry a learning claim (Rule 3). |
| F3 · Programming self-efficacy (Pintrich scale as modified by Wang & Lin 2007) ChatGPT-assisted programming learning vs Conventional lecture-based instructionAI availability at assessment not reported — cannot carry a learning claimImmediate post-testTransfer not assessedSelf-reportPrespecification unclear | Harm ANCOVA adjusted for pretest: F = 15.78, partial eta-squared = 0.10 (EG lower) not_reported <0.01 Self-report scale (Cronbach's alpha 0.86, six-point); proxy outcome — cannot carry a learning claim (Rule 3). |
- Limitations
- Single site, all-girls school (authors note this and frame it as future work on girls' programming education); six clusters analyzed at student level; flow and self-efficacy are self-report; both groups scored lower at posttest than pretest (test difficulty not equated across waves, acknowledged by the authors); ChatGPT model version and dates of use not reported (manuscript received 2023-09-13, so use predates that); whether ChatGPT was available during the posttest is not stated; no registration mentioned.
Who was studied
- Level
- secondary
- Ages
- not_reported (second-year senior high students; grade stated, ages not)
- Country
- Taiwan (northern Taiwan)
- Prior knowledge
- All participants taught basic programming concepts by the same teacher before the experiment; basic computer competency (tablet, browser, Internet).
- Selection
- Six general-education classes at a senior high school for girls (single-sex, all female).
Methodological notes
Extracted from: https://www.tandfonline.com/doi/full/10.1080/10494820.2025.2450659 Access: Full text (open access, CC-BY), retrieved 2026-09-30 from the publisher page; saved as data/fulltext-cache/yang-2025-fulltext.txt. Supersedes the abstract-only pass A of 2026-09-18; the HOLD's full-text verification is performed in this pass. Sample-size note: N=153 across six classes; experimental n=77 (three classes), control n=76 (three classes). Analysis ns per outcome are not restated; no participant flow or attrition reporting.
Funding and conflicts
- Funding
- National Science and Technology Council, Taiwan (grants 112-2628-H-A49-001-MY2; 111-2410-H-A49-066-MY3; inter-discipline empowerment project) — public funder.
- Vendor funded
- No
- Vendor
- None identified
- Notes
- Disclosure statement present: "The author declares that they have no conflict of interest." (sic — singular, on a three-author paper).
Funding is shown on every study and never used to score it.
Claims this study bears on
- Unrestricted use of a general-purpose LLM chatbot during practice improves secondary students' subsequent performance on assessments taken without AI.Consistent, but doesn't test the claim · rests on finding F1
Exact population (secondary), intervention class (unrestricted ChatGPT chat during practice), and outcome domain (researcher-developed achievement test) of the claim, with an adverse direction: the ChatGPT classes scored lower adjusted for pretest (F=3.96, p=.048, partial eta-squared .03). But the paper never states whether ChatGPT was available during the posttest, so under the availability gate (EVIDENCE-MODEL Rule 2) the finding cannot support or contradict the claim — the nie-2025 F4 precedent for gate-failing direct-outcome findings. The class-level allocation of six classrooms analyzed at student level makes the p=.048 fragile besides. If a later source establishes the posttest was unassisted, this link is re-issued as contradicts against that assessment version.
The source, as retrieved
Abstract
No abstract retrieved.
Where this record came from
| Source | Retrieved | Identifier |
|---|---|---|
| seed | Sep 18, 2026 | link first ingestion |