AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

Studying the effect of AI Code Generators on Supporting Novice Learners in Introductory Programming

Design
Randomized trial · randomized at student level · 69 participants
Subject and task
Programming
Introductory Python programming in a self-paced web environment (Coding Steps): 45 two-part tasks (code-authoring followed by code-modification) plus 40 multiple-choice questions across basics, data-types, conditionals, loops, and arrays
Exposure: Ten 90-minute sessions over three consecutive weeks (1 introduction session, 7 training sessions during which the Codex group had AI access on authoring tasks, 2 evaluation sessions); first session was 2 hours
Population match
Secondary students: Direct · University students: Different
Attrition
21 of 90 starters (23%) did not complete: 11 participated only in the first session, 4 in less than half of the sessions, and 6 missed one of the last two evaluation sessions. Paper states no common dropout factors were identified (disability, native language, computer/internet access) and that completers in the two groups had similar Scratch pre-test means (Codex M=62.7%, Baseline M=60%, t(67)=0.54, p=.67). Attrition counts per assigned condition are not reported.
Study tier
Study tier 2True experiment with student-level random assignment within pre-test-matched pairs and a well-controlled active learning environment identical across arms except for the code generator. However, the sample is small (n=69, 33/36 per arm), outcome measures are researcher-developed and unvalidated, post-tests stayed within trained content (no far transfer), the key retention contrasts are nonsignificant and likely underpowered, and 23% attrition from randomization to analysis is not broken down by condition.
Assessment
Version 2 · AI: two passes agreed · Sep 18, 2026

Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.

69 novices aged 10-17 with no prior text-based coding experience learned Python over ten sessions; half could use an OpenAI Codex-based code generator on the code-writing (authoring) tasks during training. With the AI available, that group completed more tasks, wrote more correct code, made fewer errors, and worked faster, and did just as well as the control group on manual code-modification tasks done without the AI. On post-tests taken without any AI (one day and one week after training), the two groups performed similarly overall — the Codex group scored somewhat higher at one-week retention, but not statistically significantly. Learners with higher Scratch pre-test scores who had trained with the AI did significantly better on the retention test than comparable controls.

01Findings

What the study measured

Arms

ArmWhat it got
A · Codex group (AI code generator available during training authoring tasks)
n = 33
Intervention · teacher supervised
Coding Steps with OpenAI Codex code generator · OpenAI Codex (GPT-3 code model) · in class · used not_reported
B · Baseline group (same environment, code generator removed)
n = 36
Control · teacher supervised
Identical Coding Steps environment, same 45 authoring/modifying tasks and 40 MCQs in the same fixed order, same embedded novice-friendly Python documentation, same instructor help and personalized grading feedback — but with the AI code generator removed; learners wrote all code manually.

Findings

A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.

Finding and conditionsResult
F1 · Overall task completion rate during the training phase
Codex group (AI code generator available during training authoring tasks) vs Baseline group (same environment, code generator removed)AI partly available at assessment — cannot carry a learning claimDuring interventionSame contentProcess / behaviorPrespecification unclear
Benefit
Codex M=90.9% vs Baseline M=79.0%; d=0.68 (abstract: 1.15x increased completion rate)
not_reported
p=.006 (t(67)=2.8)
Progress through fixed task sequence regardless of correctness or skips; computed from logs. Rate spans the whole training phase: AI was available to arm A only on the authoring half of each task pair, not on modifying tasks or MCQs.
F10 · Prior-competency moderation: retention post-test performance among learners with high Scratch pre-test scores (median-style split into High/Low)
Codex group (AI code generator available during training authoring tasks) vs Baseline group (same environment, code generator removed)Measured without AIDelayed post-test · 1 week after the immediate post-testSame contentResearcher-developed testPrespecification unclear
Benefit
Codex-High vs Baseline-High on retention (Figure 10): authoring M=86% vs M=63%, d=0.95; modifying M=71% vs M=49%, d=0.90; MCQs M=65% vs M=44%, d=0.93. Codex-Low vs Baseline-Low retention contrasts were all nonsignificant (d=0.06-0.20).
not_reported
authoring p=.011; modifying p=.015; MCQs p=.014 (Figure 10)
Post-hoc subgroup analysis (RQ5) with small cells (n=16-18); numeric values read from the Figure 10 statistics table. Benefit applies to the high-prior-knowledge subgroup only; low-prior subgroups showed near-equal retention performance, and low scorers still benefited during training authoring (p<.001, d=2.47). AI was not available during the retention measurement.
F2 · Correctness score on code-authoring tasks during the training phase (also faster: completion time Codex M=210s vs Baseline M=361s, t(67)=6.40, p<.001, d=1.56; fewer errors per task, d=0.74)
Codex group (AI code generator available during training authoring tasks) vs Baseline group (same environment, code generator removed)With AI at assessment — assisted performance, not learning evidenceDuring interventionSame contentGraded artifactPrespecification unclear
Benefit
Codex M=80.1% (SD=14.5%) vs Baseline M=44.4% (SD=26.5%); d=1.67 (abstract: 1.8x higher scores)
not_reported
p<.001 (t(67)=6.92)
Researcher-graded final submissions, two independent graders, simple 25%-deduction rubric; 79% full agreement. Paper states authoring tasks were the only time the Codex group had AI access, so this measures AI-assisted performance, not unaided learning. 49% of AI-used tasks were submitted with the AI-generated code unmodified.
F3 · Correctness score on manual code-modification tasks during the training phase
Codex group (AI code generator available during training authoring tasks) vs Baseline group (same environment, code generator removed)Measured without AIDuring interventionSame contentGraded artifactPrespecification unclear
Null
Codex M=66.2% (SD=25.6%) vs Baseline M=58.4% (SD=23.9%); d=0.31
not_reported
p=.202 (t(67)=1.28)
Both groups modified provided working code with no AI access; tests whether AI use on the preceding authoring task degraded immediate manual ability. No significant differences on any topic after Bonferroni correction (arrays nominal p=.022, d=0.573 favoring Codex).
F4 · Immediate post-test, code-authoring tasks (5 tasks, one day after training)
Codex group (AI code generator available during training authoring tasks) vs Baseline group (same environment, code generator removed)Measured without AIImmediate post-testSame contentResearcher-developed testPrespecification unclear
Null
Codex M=61.3% (SD=30.1%) vs Baseline M=62.9% (SD=32.0%); d=0.05
not_reported
p=.838 (t(67)=0.20)
No code generator, documentation, or feedback available; tasks analogous to training tasks in topics and difficulty (near/same-content, per the stated limitation that post-tests did not leave trained boundaries). Tasks from topics a learner covered <50% of during training were excluded from analysis.
F5 · Immediate post-test, code-modification tasks (5 tasks)
Codex group (AI code generator available during training authoring tasks) vs Baseline group (same environment, code generator removed)Measured without AIImmediate post-testSame contentResearcher-developed testPrespecification unclear
Null
Codex M=59.7% (SD=33.4%) vs Baseline M=59.3% (SD=31.6%); d=0.01
not_reported
p=.953 (t(67)=0.058)
Same evaluation-phase conditions as F4; no meaningful differences on completion times or error rates either.
F6 · Immediate post-test, multiple-choice questions (40 MCQs, overall score)
Codex group (AI code generator available during training authoring tasks) vs Baseline group (same environment, code generator removed)Measured without AIImmediate post-testSame contentResearcher-developed testPrespecification unclear
Null
Codex M=49.3% (SD=27.3%) vs Baseline M=42.0% (SD=21.6%); d=0.30
not_reported
p=.228 (t(67)=1.21)
Overall contrast nonsignificant; on the arrays subtopic the Codex group was significantly higher (p=.007, d=0.68), and 16% higher on loops (p=.041, d=0.51) which did not survive the adjusted alpha.
F7 · Retention post-test, code-authoring tasks (new task set, one week after immediate post-test)
Codex group (AI code generator available during training authoring tasks) vs Baseline group (same environment, code generator removed)Measured without AIDelayed post-test · 1 week after the immediate post-test (~8 days after training ended)Same contentResearcher-developed testPrespecification unclear
Null
Codex 9.0% higher: M=59.1% (SD=34.8%) vs Baseline M=49.8% (SD=32.4%); d=0.28
not_reported
p=.262 (t(67)=1.13)
New set of tasks with similar number/order to the immediate test; no AI, documentation, or feedback. Codex group encountered significantly more errors on retention authoring tasks (t(67)=2.08, p=.042, d=0.51), which the authors attribute to the Codex group skipping fewer tasks and having no starter code.
F8 · Retention post-test, code-modification tasks
Codex group (AI code generator available during training authoring tasks) vs Baseline group (same environment, code generator removed)Measured without AIDelayed post-test · 1 week after the immediate post-testSame contentResearcher-developed testPrespecification unclear
Null
Codex 12.7% higher: M=47.5% (SD=31.6%) vs Baseline M=34.8% (SD=29.0%); d=0.41
not_reported
p=.092 (t(67)=1.70)
Paper explicitly states neither this nor the authoring-task retention difference reached statistical significance.
F9 · Retention post-test, multiple-choice questions (overall score)
Codex group (AI code generator available during training authoring tasks) vs Baseline group (same environment, code generator removed)Measured without AIDelayed post-test · 1 week after the immediate post-testSame contentResearcher-developed testPrespecification unclear
Null
Codex M=44.1% (SD=28.9%) vs Baseline M=34.6% (SD=20.3%); d=0.38
not_reported
p=.129 (t(67)=1.54)
Subtopic advantages for Codex on loops (p=.04, d=0.51) and arrays (p=.047, d=0.49) did not reach the Bonferroni-adjusted significance threshold; paper reports neither difference as statistically significant.
Limitations
Source-stated: correctness scored with the authors' own simple rubric; several between-group differences showed effect-size trends but did not reach significance, possibly due to sample size; post-tests "did not leave the boundaries of what learners were trained on" (no far-transfer assessment); most participants were non-native English speakers; study covered only code generation, not other AI-assistant capabilities; qualitative usage analysis left to future work. Observed: 23% attrition between assignment and analysis without per-condition breakdown; no preregistration mentioned; prior-competency moderation analysis is a post-hoc quartile-style split (n=16-18 per cell); retention interval only one week; heavy AI usage patterns (49% of AI-used tasks submitted with AI-generated code unmodified) complicate interpretation of training-phase "performance."

Who was studied

Level
secondary
Ages
10-17 (M=12.5, SD=1.8)
Country
Canada (recruited through coding camps described as located in two major North American cities; acknowledgments name camps in Ottawa and the Greater Toronto Area)
Prior knowledge
No prior text-based programming experience (self-reported); 64 of 69 had used a block-based environment like Scratch or Code.org, and 27 had taken a programming-related class; balanced on a 25-item Scratch pre-test
Selection
Volunteers from more than 200 coding-camp sign-ups; 90 reporting no prior text-based programming experience were contacted and started; compensated with a $50 gift card; ethics-board approved

Methodological notes

Extracted from: https://arxiv.org/pdf/2302.07427 Access: Full text (arXiv version, arXiv:2302.07427v2, 21 Feb 2023), extracted from the arXiv PDF rather than the ACM version of record. PDF downloaded and read page-by-page; quotes transcribed from the rendered pages. Sample-size note: 69 completers analyzed (21 female/48 male): Codex group n=33, Baseline group n=36. 90 participants started and were pair-matched into groups after the Scratch pre-test.

01bRisk of bias

How much each result can be relied on

Risk of bias is judged per result, not per paper: the same study can report one outcome at low risk and another at high. Two reviewers assess each result independently against a written guide, and every domain judgment below carries its reasoning and the sentences it rests on.

Risk of bias is not the same as whether the numbers work. A trial can be well conducted and still print a standard deviation that is arithmetically impossible, or a table whose sign contradicts its own text. Risk of bias has no question for that, so where it was found it is recorded separately below, under Source integrity.

These words are not quality scores. On the RoB 2 scale, High risk of bias is the worst rating a result can receive and Low risk of bias is the best — the opposite of how the same words read on some other scales. Levels are always written out in full here for that reason.

Assessed, not settled (1)

Recorded rather than omitted. Leaving these out would make the appraisal look cleaner than it is, and a disagreement between two careful readers is itself worth knowing.

EVAL

Assessed, not settled · RoB 2 · Two reviewers reached different judgments and nothing has settled it

This result has no overall judgment, because the two independent reviews did not reach one. Both readings are shown. The domains they did agree on are below.

Where they differThe two readings
D2 · Deviations from intended interventionsOne reviewer: Low risk of bias
The other: Some concerns about risk of bias
DomainJudgment and reasoning
D1 · Randomisation process
Some concerns about risk of bias
Assignment was genuinely randomized within pairs matched on Scratch pre-test scores, and the completing groups were balanced on the pre-test (Codex M=62.7% vs Baseline M=60%, t(67)=0.54, p=.67), with no baseline imbalances suggesting a problem with randomization. However, the paper gives no information about who generated the allocation sequence or whether allocation was concealed from participants and experimenters until assignment, so concealment cannot be confirmed. Random sequence plausibly adequate + no information on concealment + no suggestive imbalance yields Some concerns under the RoB 2 algorithm.
Following the Introduction Phase using Scratch, participants were divided into two groups using a matched-groups design [13, 60].
One learner in each pair was randomly assigned to either the Codex or Baseline group.
D2 · Deviations from intended interventions
Low risk of bias
Participants and instructors were necessarily aware of condition (Codex group used the code generator; Baseline did not), but for the effect of assignment no deviations from the intended interventions arising from the trial context are apparent: both groups used the same Coding Steps environment, worked the same 45 tasks in the same fixed order, and received comparable instructor feedback (feedback counts and lengths were measured and similar across groups). During the evaluation sessions being assessed here, neither group had access to the AI code generator, which is the intended contrast (training with vs without AI, tested unassisted). All completers were analyzed in the groups to which they were assigned; incompleteness of follow-up is addressed under D3.
The tasks and the topics were presented in a fixed order for all students in both groups, as they were designed to gradually increase in complexity.
To ensure both groups received consistent and unbiased feedback on their submissions, the grading dashboard did not display any information that would reveal the identity of learners or their group (Figure 5, Right).
D3 · Missing outcome data
Some concerns about risk of bias
Of 90 randomized participants, only 69 (77%) completed all phases and only completers were analyzed (23% attrition), with no intention-to-treat or sensitivity analysis. The paper reports completer counts per arm (33 Codex, 36 Baseline) but does not report the original per-arm allocation, so differential attrition cannot be fully verified. The authors checked dropout characteristics and found no common factors, and completing groups remained balanced on the pre-test, which argues against strong bias; but dropout in a voluntary out-of-school program could plausibly relate to frustration or struggle, which the intervention itself affected (Codex learners reported less stress), so missingness depending on the true outcome cannot be ruled out. There is, however, no positive evidence that it did, supporting Some concerns rather than High.
From the 90 participants that started the study, 69 learners (21 female/48 male) ages 10-17 (M=12.5; SD=1.8) successfully completed all three phases of the study (11 only participated in the first session, four participated in less than half of the sessions, and six missed one of the last two evaluation sessions).
No common factors were identified among the 21 students who dropped out in terms of disability, native language (six non-English) and computer or internet access.
D4 · Measurement of the outcome
Low risk of bias
Outcome measurement was identical for both groups: the same researcher-designed authoring, modifying, and multiple-choice tasks, administered in the same Coding Steps environment with neither group having access to the code generator or documentation during evaluation. Coding tasks were graded by two independent researchers using a fixed rubric (79% full agreement, disagreements resolved), and the grading dashboard concealed learner identity and group, so assessors were effectively blinded; multiple-choice scoring is objective. Measures were researcher-made and aligned to the training content, but this applies equally to both arms and does not favor one condition at an unassisted test.
The Coding Steps application was still used for the evaluation phase, but learners did not have access to either the code generator or the Python documentation, or any feedback after submitting the tasks.
For calculating correctness scores on the coding tasks, two independent researchers graded each submitted solution using a simple and consistent grading rubric of deducting 25% for each major issue or missing part in their final submission (0%, 25%, 50%, 75%, and 100%).
D5 · Selection of the reported result
Some concerns about risk of bias
No preregistration, registered protocol, or prespecified analysis plan is mentioned anywhere in the paper, so reported results cannot be checked against prior intentions. Many outcomes, topic-level breakdowns, and subgroup analyses are reported, and at least one analytic exclusion for the evaluation phase (dropping tasks on topics where a learner progressed less than 50% in training) is not described as prespecified. Mitigating this, the metrics are defined systematically (Table 2), both significant and non-significant results are reported with means, SDs, CIs, and effect sizes, and Bonferroni corrections are applied to multiplicity, so selective reporting of favorable results is not suggested; but the absence of a prespecified plan leaves Some concerns.
For each reported metric, we report means, standard deviations, and 95-percentile confidence intervals for each condition.
For our analysis of the evaluation phase, we excluded tasks related to topics for which the learner progressed less than 50% during the training phase.
02Funding

Funding and conflicts

Funding
not_reported (no funding statement; acknowledgments thank after-school coding and STEM camps — Coder Sports, CodeZilla, Hive5 Innovative Center, Junior Innovators — for help recruiting)
Vendor funded
Unclear
Vendor
None identified
Notes
No conflict-of-interest statement observed. Authors are university researchers (University of Toronto, University of Michigan, University of Maryland). The paper notes "Coding Steps was approved by the OpenAI App Review team prior to running the study" — API access approval, not stated as funding or employment ties.

Funding is shown on every study and never used to score it.

03Claims

Claims this study bears on

04Source

The source, as retrieved

Abstract

No abstract retrieved.

Where this record came from

SourceRetrievedIdentifier
seedSep 18, 2026link
first ingestion