AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers

Design
Qualitative study · randomized at none level · 21 participants
Subject and task
Programming
Solving one CS1 C++ programming problem ("More Positive or Negative": count whether more positive or negative integers were entered, loop with sentinel 0) with GitHub Copilot and ChatGPT available, submitted to the Athene automated assessment tool; replication of Prather et al.'s pre-GenAI think-aloud study.
Exposure: Single lab session: warm-up "Hello World" task, then up to 35 minutes on the target problem (observed solve times 5 to 35 minutes, mean 17.1, SD 8.1), followed by a post-session interview.
Population match
Secondary students: Different · University students: Direct
Attrition
No dropout reported among the 21 participants; all completed the session (1 of 21 did not produce a working program within the 35-minute limit). 6 of 27 enrolled students chose not to participate, including 1 student who identified as African-American and 2 who identified as Hispanic.
Study tier
Study tier 3Small single-site observational lab study (n=21) with no comparison group and no randomization; the pre-GenAI baseline is a separate prior cohort (n=31), so causal claims about GenAI effects cannot be supported. However, as mechanism evidence it is comparatively strong for its genre: triangulated think-aloud, screen replay, and eye-tracking data; a replication protocol reusing the prior study's task, time limit, and codebook; dual coding with Cohen's Kappa 0.74; and correlational statistics with reported p-values. Findings should be read as hypotheses about how GenAI interacts with novice metacognition, not as causal effect estimates on learning outcomes (the authors state they did not measure learning outcomes).
Assessment
Version 2 · AI: two passes agreed · Sep 18, 2026

Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.

Researchers watched 21 first-semester programming students solve a standard beginner problem while using GitHub Copilot and ChatGPT, with think-aloud, eye tracking, and interviews, replicating an earlier study run before generative AI existed. Almost everyone (20 of 21) finished the problem, versus 20 of 31 in the earlier no-AI study, but about half struggled: all five previously known metacognitive difficulties reappeared and three new AI-driven ones emerged (constant interruption by suggestions, being misled down wrong paths, and being conceptually behind while feeling confident). Students with lower grades and lower self-efficacy showed more of these difficulties and accepted more AI suggestions, while stronger students used the AI to speed up code they already intended to write. The authors conclude that AI let struggling students finish while leaving them with an illusion of competence, potentially widening the gap between well-prepared and under-prepared students.

01Findings

What the study measured

Arms

ArmWhat it got
A · Observed cohort: CS1 students solving a programming problem in VSCode with GitHub Copilot (inline code suggestions) and browser ChatGPT available; no comparison arm (prior n=31 pre-GenAI cohort serves as historical reference only).
n = 21
Intervention · teacher independent
GitHub Copilot and ChatGPT · GPT (OpenAI models underlying ChatGPT and GitHub Copilot) · Vanilla chat · chat interface · in class · used not_reported (single semester at study site; loops introduced two weeks before sessions; preprint posted May 2024)

Findings

A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.

Finding and conditionsResult
F1 · Task completion within 35-minute limit and time-to-solve, judged by the Athene automated assessment tool's test cases (25 test cases)
Observed cohort: CS1 students solving a programming problem in VSCode with GitHub Copilot (inline code suggestions) and browser ChatGPT available; no comparison arm (prior n=31 pre-GenAI cohort serves as historical reference only). vs With AI at assessment — assisted performance, not learning evidenceDuring interventionSame contentGraded artifactPrespecification unclear
Unclear
20 of 21 students completed a working program within the time limit; times 5-35 minutes, mean 17.1 (SD 8.1); Copilot suggestion accept rates 10%-53%. Authors contrast descriptively with the prior pre-GenAI study where 11 of 31 did not finish, but note 9 of the 10 strugglers in this study "utilized GenAI to get to the solution" and their observations suggest those students would not have solved it unaided.
not_reported
not_reported
Assisted measure: AI was fully available during the assessed task, so completion conflates student competence with tool output; the paper itself argues completion overstates struggling students' understanding. No inferential test against the historical cohort.
F2 · Presence of metacognitive difficulties during problem solving: the five previously identified (Forming, Dislodging, Assumption, Location, Achievement) plus three new GenAI-related difficulties (Interruption, Mislead, Progression), coded from observation notes, think-aloud, and eye-tracking replay
Observed cohort: CS1 students solving a programming problem in VSCode with GitHub Copilot (inline code suggestions) and browser ChatGPT available; no comparison arm (prior n=31 pre-GenAI cohort serves as historical reference only). vs With AI at assessment — assisted performance, not learning evidenceDuring interventionSame contentProcess / behaviorPrespecification unclear
Unclear
About half of students were observed struggling; 9 of 21 tagged with previously known metacognitive difficulties, 8 with new GenAI-related difficulties (only 1 had a new label without any old label). All five prior difficulties recurred; Location was most common, "which GenAI unfortunately facilitated for many of our participants, giving them the illusion of progress."
not_reported
not_reported
Qualitative coding with an initial 15-code codebook; two coders reached Cohen's Kappa 0.74 after four jointly discussed sessions, then coded independently. Stated by the authors as a result (difficulties persist and are compounded by GenAI), but no inferential test attaches to the theme itself, so direction coded unclear per convention.
F3 · Correlation between course grade (weekly programming quiz grades) and count of new GenAI-related metacognitive difficulties
Observed cohort: CS1 students solving a programming problem in VSCode with GitHub Copilot (inline code suggestions) and browser ChatGPT available; no comparison arm (prior n=31 pre-GenAI cohort serves as historical reference only). vs AI partly available at assessment — cannot carry a learning claimDuring interventionSame contentProcess / behaviorPrespecification unclear
Harm
r = -0.503 (grade x new metacognitive issues count); also grade x old metacognitive issues count r = -0.482 (Table 1 p=0.02682; running text prints p=0.268, apparent typo), and grade x time-to-solve r = -0.727 (p=0.0002)
not_reported
0.020
Correlational association within a single AI-using cohort, not a causal AI-vs-no-AI contrast; direction coded harm because the paper presents it as evidence that GenAI-specific difficulties concentrate among lower-performing students ("widening the digital divide"). Grades are from unassisted weekly quizzes; difficulty counts from AI-assisted sessions, hence ai_available partial for the aggregated relation. n=21, so estimates are unstable.
F4 · Copilot suggestion acceptance rate by presence of metacognitive difficulties (Table 4)
Observed cohort: CS1 students solving a programming problem in VSCode with GitHub Copilot (inline code suggestions) and browser ChatGPT available; no comparison arm (prior n=31 pre-GenAI cohort serves as historical reference only). vs With AI at assessment — assisted performance, not learning evidenceDuring interventionSame contentProcess / behaviorPrespecification unclear
Unclear
Students with metacognitive difficulties: mean accept rate 34.1% (SD 12.5%, range 17-53%); without: mean 24.5% (SD 6.6%, range 10-33%). "students who experienced metacognitive difficulties tended to accept CoPilot suggestions at higher rates" and "accepted more suggestions that they subsequently reworked or rolled back entirely."
not_reported
not_reported
Descriptive group comparison; no inferential test reported for the Metas vs No-Metas difference, so direction coded unclear. Related significant correlation reported: metacognitive issues count x negative self-efficacy count r=0.552, p=0.009.
F5 · Qualitative theme: divide between students who accelerated and students who struggled ("widening gap"): accelerators (11 students with no difficulties) used GenAI to create code they already intended to make and ignored unhelpful or incorrect suggestions; strugglers had prior difficulties compounded and new ones introduced by GenAI
Observed cohort: CS1 students solving a programming problem in VSCode with GitHub Copilot (inline code suggestions) and browser ChatGPT available; no comparison arm (prior n=31 pre-GenAI cohort serves as historical reference only). vs With AI at assessment — assisted performance, not learning evidenceDuring interventionSame contentProcess / behaviorPrespecification unclear
Mixed
"Eleven students did not display any metacognitive difficulties. Instead, some used AI in positive ways to help them solve the problem quickly." vs. half of students observed struggling; fastest accelerators finished in 5-13 minutes with low accept rates (10-30%).
not_reported
not_reported
Central stated result of the paper (title claim), inherently bidirectional, hence direction mixed. Qualitative characterization triangulated with eye tracking; no inferential test of the divide itself.
F6 · Qualitative theme: illusion of competence / cognitive dissonance among struggling students (post-problem interview statements contradicting observed in-session behavior)
Observed cohort: CS1 students solving a programming problem in VSCode with GitHub Copilot (inline code suggestions) and browser ChatGPT available; no comparison arm (prior n=31 pre-GenAI cohort serves as historical reference only). vs Measured without AIImmediate post-testSame contentProcess / behaviorPrespecification unclear
Unclear
"Post-problem interviews confirmed that many of the participants rationalized their usage of GenAI tools in ways that contradicted their words and actions during the sessions." Of the 10 strugglers, 9 reached the solution via GenAI; "it appears that most of these ten who struggled thought they understood more than they actually did."
not_reported
not_reported
Interpretive synthesis of interview vs observed behavior (e.g., P7 claimed to use ChatGPT "like a personal tutor" while observed doing the opposite; P8/P17 helpfulness claims contradicted by session data). No inferential test; direction coded unclear per convention although the paper frames it as a harm.
F7 · Post-session interview Likert questions on programming/AI experience and perceived helpfulness (Fig. 3, e.g. "Using AI in this class is helping me learn"), plus self-reported prior experience
Observed cohort: CS1 students solving a programming problem in VSCode with GitHub Copilot (inline code suggestions) and browser ChatGPT available; no comparison arm (prior n=31 pre-GenAI cohort serves as historical reference only). vs AI availability at assessment not reported — cannot carry a learning claimImmediate post-testTransfer not assessedSelf-reportPrespecification unclear
Unclear
Distributions shown graphically only (Fig. 3); reported correlations: used AI in class x finds AI helpful in class r=0.661 (p=0.0011); programming experience before class x experience with AI r=0.622 (p=0.0026), both self-reported.
not_reported
not_reported
Researcher-written interview items, self-reported; the paper explicitly cautions that self-reported GenAI usage data "may not be as reliable as most currently think" given observed unawareness of Copilot use.
F8 · Self-efficacy (MSLQ self-efficacy subscale), collected one week after the lab session at the beginning of class
Observed cohort: CS1 students solving a programming problem in VSCode with GitHub Copilot (inline code suggestions) and browser ChatGPT available; no comparison arm (prior n=31 pre-GenAI cohort serves as historical reference only). vs AI availability at assessment not reported — cannot carry a learning claimDelayed post-test · 1 weekTransfer not assessedSelf-reportPrespecification unclear
Unclear
Self-efficacy score x course grade r=0.553 (p=0.0094); metacognitive issues count x negative self-efficacy count r=0.552 (p=0.009); lower self-efficacy students "were more likely to have metacognitive difficulties."
not_reported
0.0094
Validated instrument (MSLQ self-efficacy subscale) administered apart from the lab to avoid lab stress contamination; used as a correlate of difficulties/grades, not as an AI outcome, so no benefit/harm direction is attributable to AI despite the inferential tests.
Limitations
Authors' stated limitations: small sample (21 vs 31 in the original study), limiting generalizability; only two GenAI tools (ChatGPT and GitHub Copilot), not exhaustive of available tools; single site. Additional observations: no concurrent control condition (comparison is to a prior-cohort study); completion of the task with AI available is an assisted measure that the authors themselves argue overstates unassisted competence; self-reported AI/programming experience; text reports p=0.268 for the grade x old-difficulties correlation while Table 1 reports p=0.02682 for the same pair (apparent typo).

Who was studied

Level
university
Ages
not reported (age was collected in the post-session interview but no distribution is reported)
Country
USA
Prior knowledge
Novices enrolled in a CS1 course (C++); loops had been introduced two weeks before the study. The course used GenAI (Copilot, ChatGPT) openly from the first day of class, modeled by the professor.
Selection
Opt-in for extra credit: 21 of 27 students enrolled in the CS1 course at a small research university participated; non-participants were offered an alternative extra-credit task. 7 participants identified as women, 14 as men; 3 African-American, 2 Hispanic, 1 racially/ethnically Jewish, 15 white/Caucasian (2 of these from Europe). Institution described as growing toward Hispanic-Serving Institution status.

Methodological notes

Extracted from: https://arxiv.org/pdf/2405.17739 Access: arXiv preprint 2405.17739, version v1 (submitted 28 May 2024), cs.AI; author manuscript of the ICER 2024 paper ("Accepted to ICER 2024"; ACM copyright block present, DOI shown as placeholder XXXXXXX in this preprint version). Full 23+ page PDF read directly (fetched via arXiv, saved locally by the fetch tool). Abstract page https://arxiv.org/abs/2405.17739 also consulted for version metadata. Sample-size note: 21 lab sessions, one per participant (21 of 27 enrolled CS1 students opted in). Original study being replicated had n=31. Analyses combine observation notes, think-aloud verbalizations, eye tracking, interviews, Copilot suggestion accept rates, weekly quiz grades, and MSLQ self-efficacy scores.

02Funding

Funding and conflicts

Funding
"This research is funded by the Google Award for Inclusion Research Program and was also partially supported by the Research Council of Finland (Academy Research Fellow grant number 356114)."
Vendor funded
Partial
Vendor
None identified
Notes
Google (industry) award funding acknowledged, though Google does not make the studied tools (GitHub Copilot, ChatGPT). No conflict-of-interest statement in the preprint. The course professor was one of the researchers, and participation was incentivized with extra credit in that professor's course.

Funding is shown on every study and never used to score it.

03Claims

Claims this study bears on

04Source

The source, as retrieved

Abstract

No abstract retrieved.

Where this record came from

SourceRetrievedIdentifier
seedSep 18, 2026link
first ingestion