AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement But May Increase Adopters' Exam Performances

Design
Randomized trial · randomized at student level · 5831 participants
Subject and task
Programming
Introductory Python programming in Stanford's free 6-week online course Code in Place (spring 2023); treatment was access to a course-specific GPT-4 chat interface
Exposure: 6-week course (April 24 - June 5, 2023); GPT-4 access offered from the start of week 4 (May 15), 9 days before the optional diagnostic exam (May 24-26); Figure 1 labels the experiment period as May 15-24, and the paper does not state that access was revoked afterward
Population match
Secondary students: Different · University students: Partial
Attrition
No further post-randomization dropout is reported beyond the course's normal voluntary disengagement, which is itself the outcome: 54.2% of randomized students skipped the optional exam, differentially by arm (55.9% experiment vs 51.5% control). Homework completion declined in both arms across weeks (roughly 79% week 2 to ~45-50% week 6). An after-course survey found around 2% of students said they used ChatGPT outside the class.
Study tier
Study tier 2Large student-level randomized encouragement trial (N=5,831): the OFFER of GPT-4 was randomized, so the engagement ITT effects (exam participation, homework completion) are well identified, precisely estimated, and survive the authors' Bonferroni correction across 15 tests. The exam-performance conclusions are much weaker: the ITT effect on scores is null, and the headline adopter benefit is a LATE for the 14.2% self-selected compliers, resting on unverifiable exclusion and monotonicity assumptions plus ML imputation for the >50% missing exam outcomes, and the authors state it is not statistically significant after multiple-hypothesis adjustment. No preregistration is mentioned, and subgroup results (low-HDI, age, experience) are explicitly exploratory.
Assessment
Version 2 · AI: two passes agreed · Sep 18, 2026

Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.

In a free 6-week global online Python course with 5,831 active students, researchers randomly offered 60% of students a course-specific GPT-4 chat tool partway through the course. Only 14.2% of offered students used it, and simply offering and advertising it significantly REDUCED engagement: exam participation fell 4.3 percentage points and week-6 homework completion fell 4.6 points, though students from low-HDI countries showed the opposite (exploratory) pattern. Offering the tool did not change average exam scores overall; a causal (instrumental variable) estimate suggests the self-selected users may have scored about 6.8 points higher than they would have without GPT-4, but the authors note this is not statistically significant after adjusting for multiple comparisons.

01Findings

What the study measured

Arms

ArmWhat it got
A · Experiment: access to in-class GPT-4 chat interface (advertised by email and homepage sidebar button)
n = 3581
Intervention · teacher independent
Custom in-class ChatGPT-like chat interface built into the Code in Place course site · GPT-4 (OpenAI) · Prompted pedagogy · chat interface · unrestricted · used May 15-24, 2023 (start of week 4 through the optional exam window May 24-26; revocation after the experiment period not stated)
B · Control: no access to the in-class GPT-4 interface (no email; normal course experience)
n = 2250
Control · teacher not reported
Business-as-usual course: video lectures, weekly volunteer-taught sections, weekly homework, course materials; no in-class GPT-4 interface and no advertisement email. Use of public ChatGPT outside the class was possible but unverifiable (~2% self-reported).

Findings

A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.

Finding and conditionsResult
F1 · Diagnostic (optional midterm) exam participation rate: 44.1% vs 48.5% (ITT)
Experiment: access to in-class GPT-4 chat interface (advertised by email and homepage sidebar button) vs Control: no access to the in-class GPT-4 interface (no email; normal course experience)AI availability at assessment not reported — cannot carry a learning claimDuring interventionTransfer not assessedProcess / behaviorPrespecification unclear
Harm
Delta = -4.3 pp (SE=1.3); figure caption also reports SE=1.34
95% CI [-6.9, -1.7]
P=.020 Bonferroni-corrected (x15); unadjusted P=.001; Figure 3(b) caption reports P=0.006
Randomized ITT contrast on an objective platform-logged behavior (taking the voluntary 4-hour exam, May 24-26, right at the end of the 9-day experiment period). The exam earns an extra certificate distinction but is not required. Whether the in-class GPT-4 remained available during the exam window is not explicitly stated; the authors examined transcripts and found no students asking exam questions through the interface, and could not verify outside ChatGPT use.
F2 · Week 6 homework completion rate (all problems solved): 45.2% vs 49.8% (ITT)
Experiment: access to in-class GPT-4 chat interface (advertised by email and homepage sidebar button) vs Control: no access to the in-class GPT-4 interface (no email; normal course experience)AI availability at assessment not reported — cannot carry a learning claimDuring interventionTransfer not assessedProcess / behaviorPrespecification unclear
Harm
Delta = -4.6 pp (SE=1.3); figure caption reports SE=1.34
95% CI [-7.2, -1.9]
P=.01 Bonferroni-corrected; unadjusted P<.001
Randomized ITT contrast on logged homework completion. Arms had indistinguishable completion rates before the week-4 experiment start and diverged afterward. Week 6 falls after the labeled May 15-24 experiment period; whether treatment-arm access to GPT-4 continued into week 6 (i.e., whether homework could be AI-assisted) is not stated, hence ai_available_at_assessment not_reported.
F3 · Week 6 section attendance: 58.8% vs 61.6% (ITT)
Experiment: access to in-class GPT-4 chat interface (advertised by email and homepage sidebar button) vs Control: no access to the in-class GPT-4 interface (no email; normal course experience)AI availability at assessment not reported — cannot carry a learning claimDuring interventionTransfer not assessedProcess / behaviorPrespecification unclear
Null
Delta ~ -2.8 pp (SE=1.31, from Figure 3(b) caption)
CI [-5.41, -0.26] (Figure 3(b) caption; correction basis of this CI not fully specified)
P>.05 after Bonferroni correction (labeled NS in Figure 3(b))
Randomized ITT contrast on logged attendance at weekly human-taught sections. The reported interval excludes zero but the authors label the contrast non-significant after Bonferroni correction; coded null per the paper's own adjusted inference. Text describes it as part of a general decrease in engagement.
F4 · Diagnostic exam score, ITT / Advertisement Effect (E1), 0-100 scale
Experiment: access to in-class GPT-4 chat interface (advertised by email and homepage sidebar button) vs Control: no access to the in-class GPT-4 interface (no email; normal course experience)AI partly available at assessment — cannot carry a learning claimDuring interventionSame contentCourse examPrespecification unclear
Null
Impute-for-missingness (primary handling of >50% missing scores): Delta = +0.67 pp, ES=0.09; ignore-missingness: Delta = +0.91 pp, ES=0.05 (Table 2). Observed means: experiment takers 87.0 (SD 18.7, N=1,579) vs control takers 86.1 (SD 20.0, N=1,089).
90% CI [-0.37, 1.72] (impute); 90% CI [-0.00, 1.89] (ignore)
Not reported for E1 in Table 2; both 90% CIs include or touch 0
Randomized ITT contrast, but the exam was optional and taken by only 45.8% of students with differential participation by arm (F1), so observed-score comparisons condition on a post-treatment behavior; authors address this via ignore-missingness (MCAR assumption) and cross-fitted ML imputation (missing-at-random assumption). Exam: 5 multi-part coding problems of increasing difficulty, 4 hours, course content; many perfect scores (ceiling; median 95.9). AI availability during the exam not explicitly stated (see F1 note).
F5 · Diagnostic exam score for adopters: local average treatment effect of using GPT-4 (E2), 0-100 scale
Experiment: access to in-class GPT-4 chat interface (advertised by email and homepage sidebar button) vs Control: no access to the in-class GPT-4 interface (no email; normal course experience)AI partly available at assessment — cannot carry a learning claimDuring interventionSame contentCourse examPrespecification unclear
Null
LATE = +6.86 pp with imputation for missingness, ES=0.40 (Bansak effect size, "near medium"); +4.49 pp ignoring missingness, ES=0.23. Paper headline: adopters "could see a 6.8 average percentage point increase".
BCa 90% CI [0.30, 14.13] (impute); BCa 90% CI [-0.34, 8.98] (ignore)
No exact p reported; authors state the difference "is not statistically significant after adjusting for multiple hypothesis testing" (the unadjusted impute-based 90% BCa CI excludes 0)
NON-RANDOMIZED SUBGROUP EFFECT: adopters are the 14.2% (N=510) of the experiment arm who SELF-SELECTED into using GPT-4; adoption itself was not randomized. The estimate is an instrumental-variable LATE (Imbens-Angrist compliers) using the randomized offer as the instrument, so it is causal only under unverifiable exclusion and monotonicity assumptions and applies only to compliers, who were observably more engaged (e.g., higher prior section attendance, OR 1.24) and skewed toward lower-HDI, older, male students. Coded null per the paper's own adjusted inference despite the positive point estimate emphasized in the title/abstract. Naive descriptive comparisons (adopters 88.4 vs control takers 86.1) are reported separately and are likewise confounded by self-selection.
F6 · Exam participation among students from low-HDI countries (subgroup, N=180 active; 111 experiment / 69 control): 42.3% vs 27.5%
Experiment: access to in-class GPT-4 chat interface (advertised by email and homepage sidebar button) vs Control: no access to the in-class GPT-4 interface (no email; normal course experience)AI availability at assessment not reported — cannot carry a learning claimDuring interventionTransfer not assessedProcess / behaviorPrespecification unclear
Null
Delta = +14.8 pp (SE=7.2), Cohen's d=0.31 (small-to-medium)
95% CI [0.6, 29.0]
Unadjusted P=.045; P=.680 after multiple-hypothesis adjustment
Exploratory subgroup analysis (randomization preserved within subgroup, but small n and 1 of many tested moderators). The abstract and insight list present this as an increase for low-HDI students ("The trend was reversed"), yet the authors state it is not significant after their Bonferroni-style adjustment; coded null per that adjusted inference. A similar directional trend was seen for week-6 homework in this subgroup (no statistics reported).
F7 · Exam participation among students with prior coding experience 6-10 on a 0-18 scale ('less experienced' subgroup, N=2,037)
Experiment: access to in-class GPT-4 chat interface (advertised by email and homepage sidebar button) vs Control: no access to the in-class GPT-4 interface (no email; normal course experience)AI availability at assessment not reported — cannot carry a learning claimDuring interventionTransfer not assessedProcess / behaviorPrespecification unclear
Harm
Delta = -7.5 pp (SE=2.3)
95% CI [-12.0, -3.1]
P=.014 adjusted; unadjusted P<.001
Exploratory subgroup analysis supporting a "middle experience gap": the engagement decrease concentrated among students with medium prior coding experience (and ages 23-40, for which no test statistics are given). One of several demographic moderator analyses.
Limitations
Authors' stated limitations: a particular course, cohort, and moment in time (mid-2023 AI zeitgeist); at-will free course with no grade/credit stakes, so drop-out is easy; the course centers on human teachers and may attract students who prefer human instruction; mechanisms may be specific to programming; UX choices (email + sidebar advertisement, mandatory pros/cons handout, non-integrated interface) may have shaped outcomes. Additional limitations: optional exam creates outcome missingness correlated with treatment (their engagement effect contaminates the score analysis, addressed only by MCAR or missing-at-random imputation assumptions); adopter effect is self-selected compliers only (adopters were observably more engaged, e.g. higher prior section attendance); low adoption (14.2%) limits power; guardrail robustness against solution-seeking was not thoroughly tested; use of public ChatGPT outside the class could not be verified; figure vs text p-values differ slightly for exam participation (caption P=0.006 vs text corrected P=.020).

Who was studied

Level
other
Ages
Adults, mean 31.4 (SD ~10.4), median 29, range up to 84; 18-22 college-age subgroup N=1,241 of 5,831
Country
146 countries (global online course; run by Stanford University, USA)
Prior knowledge
Coding beginners in an intro course; self-reported prior coding experience mean 5.1 (SD 4.4) on a 0-18 scale; 3,180 of 5,831 reported no experience (0-5)
Selection
Applicants approved to enroll in the free online class (8,762); the experiment pool was the 5,831 students still active after week 1; exam-score analyses further condition on the 45.8% who chose to take the optional exam

Methodological notes

Extracted from: https://arxiv.org/pdf/2407.09975 Access: arXiv preprint v2 (arXiv:2407.09975v2 [cs.CY], 15 Jul 2025), fetched as PDF and read in full (main text pp. 1-20; appendix not extracted). This is the arXiv version, not the ACM Learning@Scale 2025 version of record (DOI 10.1145/3698205.3733960). arXiv title reads "...Reduced Engagement but Increased Adopters' Exam Performances"; the ACM title uses "May Increase". Sample-size note: 5,831 active-after-week-1 students randomized (3,581 experiment / 2,250 control, a 60/40 split) out of 8,762 approved enrollees. Engagement outcomes analyzed on the full randomized pool. Exam scores observed only for the 2,668 exam takers (experiment N=1,579, 44.1%; control N=1,089, 48.5%); only 510 experiment students (14.2%) actually used GPT-4 ("adopters"). Score analyses reported both ignoring missingness and with ML imputation for non-takers.

01bRisk of bias

How much each result can be relied on

Risk of bias is judged per result, not per paper: the same study can report one outcome at low risk and another at high. Two reviewers assess each result independently against a written guide, and every domain judgment below carries its reasoning and the sentences it rests on.

Risk of bias is not the same as whether the numbers work. A trial can be well conducted and still print a standard deviation that is arithmetically impossible, or a table whose sign contradicts its own text. Risk of bias has no question for that, so where it was found it is recorded separately below, under Source integrity.

These words are not quality scores. On the RoB 2 scale, High risk of bias is the worst rating a result can receive and Low risk of bias is the best — the opposite of how the same words read on some other scales. Levels are always written out in full here for that reason.

ENGAGE

Some concerns about risk of bias · RoB 2 · Two reviewers agreed on every domain

How the overall was reached: Randomization, adherence (for an assignment-effect estimand), missing data, and outcome measurement are all low risk for these fully observed administrative outcomes; the only concern is the absence of a preregistered analysis plan with multiple outcomes tested (partially mitigated by Bonferroni correction). Overall equals the worst domain (D5: Some concerns).

DomainJudgment and reasoning
D1 · Randomisation process
Low risk of bias
Students active after week 1 (N=5,831) were randomly assigned 60/40 to experiment and control; the unequal ratio is explained by an a priori variance consideration, not by any outcome-related process. Baseline demographic covariates (age, gender, application score, English fluency, coding score, prior experience, section attendance, country HDI) are reported by arm in Table 1 and are closely balanced, consistent with intact randomization. The exact randomization mechanism (e.g., software seed) and allocation concealment are not described, but with n=5,831, administrative assignment, and the observed balance there is no indication of a problem.
We randomly assigned 60% of the students (n=3,581) to the experiment condition and 40% of the students (n=2,250) to the control condition.
We give slightly more students access to GPT-4 because we expect the measurable educational outcome to have a high variance for the experiment group.
D2 · Deviations from intended interventions
Low risk of bias
The effect of interest is assignment (the "Advertisement Effect", which the authors explicitly equate with intent-to-treat). The intervention is the offer/advertisement of the GPT-4 interface, so low take-up (14.2%) is a feature of the encouragement design, not a deviation. Participants were necessarily unblinded (the email IS the intervention), but no deviations beyond the intended offer are apparent for engagement outcomes; contamination of the control arm with the in-class tool was impossible (no access), and only ~2% of students reported using ChatGPT outside the class. Engagement analyses compare all experiment vs all control students, i.e., analysis followed assignment.
In econometrics, medicine, and statistics, the Advertisement Effect is commonly known as the "intent-to-treat effect", and the Effect for Adopters is known as the "treatment effect on the compliers".
Only 14.2% (N=510) students in the experiment group used GPT-4 prior to the midterm exam.
D3 · Missing outcome data
Low risk of bias
Exam participation (a binary took/did-not-take indicator) and week-6 homework completion are platform-logged administrative outcomes that are defined and observed for every randomized student who remains in the analysis pool; there is no missing-outcome mechanism for these measures. Group Ns in the participation comparison match the randomized Ns (3,581 vs 2,250 per Table 1).
In Code-in-Place, student activity is tracked in three ways: their submission of weekly homework, their attendance for weekly sections, and whether they watched video lectures or interacted with the course materials.
D4 · Measurement of the outcome
Low risk of bias
Outcomes are objective administrative measures recorded identically for both arms by the course platform (exam-taking status; homework marked complete when all problems are solved). Measurement is automated and could not plausibly differ by assignment, and outcome ascertainers/algorithms are not influenced by knowledge of arm. Note one interpretive caveat the authors themselves raise: disengagement from the platform could partly reflect students learning elsewhere (e.g., on public ChatGPT), which is a construct-validity issue rather than differential measurement of the recorded outcome.
note a homework is considered "complete" when a student has solved all the problems it contains, which range from 1 to 6 per week.
D5 · Selection of the reported result
Some concerns about risk of bias
No preregistration or pre-analysis plan is mentioned anywhere in the paper. Multiple engagement outcomes were tested (exam participation, week-6 section attendance, week-6 homework completion, plus numerous demographic subgroups), and the authors apply a Bonferroni correction across 15 tests, which mitigates but does not eliminate the risk that the reported outcome set and analyses were selected after results were known. Absent a prespecified plan, RoB 2 assigns at least some concerns.
We used Bonferroni corrections to account for multiple hypothesis testing. To apply Bonferroni correction, we multiplied the P-value by 15, representing the number of hypothesis tests we conducted in this study.

EXAMSCORE

High risk of bias · RoB 2 · Two reviewers agreed on every domain

How the overall was reached: D3 is high risk: more than half of randomized students have no exam score, missingness was itself affected by assignment (the study's own headline engagement result), and the MAR-based ML imputation cannot be verified to remove treatment-induced differential self-selection. With additional some-concerns in D2, D4, and D5 (unproctored exam with the tool not stated to be disabled; no preregistration), the overall judgment follows the worst domain: High.

DomainJudgment and reasoning
D1 · Randomisation process
Low risk of bias
Same randomization as ENGAGE: 60/40 random assignment of 5,831 active students with closely balanced baseline covariates reported by arm in Table 1. No indication of allocation problems.
We randomly assigned 60% of the students (n=3,581) to the experiment condition and 40% of the students (n=2,250) to the control condition.
D2 · Deviations from intended interventions
Some concerns about risk of bias
The exam was an optional 4-hour assessment taken remotely over a 3-day window in a free online course; the paper never states that it was proctored or that the in-class GPT-4 tool was disabled during the exam. Treatment-arm students therefore retained access to an assistance tool during outcome collection that control students did not have — a potential deviation from the intended (learning-support) intervention arising from the trial context. The authors checked in-class transcripts and found no exam questions asked, but they explicitly could not verify use of the public ChatGPT interface outside the class (in either arm). Analysis itself was appropriate for assignment (all randomized students compared by arm).
After 9 days (at the end of week 5), all students were invited to take an optional 4-hour midterm exam as part of a regular course offering.
After examining the transcripts, we did not find students asking diagnostic exam questions through the GPT-4 interface. However, we are not able to verify whether students have used OpenAI's public ChatGPT interface outside the class.
D3 · Missing outcome data
High risk of bias
Only 45.8% of students took the exam, so exam scores are missing for more than half of the randomized sample. Worse, missingness itself was affected by the intervention (participation 44.1% in the experiment arm vs 48.5% in control — the ENGAGE result), so the missingness mechanism differs between arms and is very likely related to the true score. The authors acknowledge MCAR is implausible and present a "missing at random given covariates" ML imputation with cross-fitting alongside a complete-case ("Ignore Missingness") analysis; MAR conditional on observed covariates is itself untestable and unlikely to fully capture treatment-induced differential self-selection into exam-taking. With >50% missingness that is differential by arm and outcome-related, risk of bias is high despite the sensitivity analyses.
Since this exam is not mandatory for a student to earn their course certificate, 45.8% of students took the exam.
MCAR is a strong assumption and unlikely to be valid in our setting, as we expect there to be differences among students who do or do not take the exam.
D4 · Measurement of the outcome
Some concerns about risk of bias
The exam is scored on coding problems and scoring is presumably applied identically to both arms, so assessor bias is unlikely. The concern is that outcome measurement occurred under unproctored, take-home conditions in which the measurement circumstances differed systematically by arm: experiment students had the in-class GPT-4 tool available during the exam window (never stated to be disabled), so observed scores could reflect assisted performance rather than learning. The transcript audit found no exam questions asked through the in-class interface, which limits but does not eliminate this pathway (outside-tool use was unverifiable). Participants were also aware of assignment while self-administering the assessment.
After examining the transcripts, we did not find students asking diagnostic exam questions through the GPT-4 interface. However, we are not able to verify whether students have used OpenAI's public ChatGPT interface outside the class.
The exam consists of 5 multi-part coding problems of increasing difficulty.
D5 · Selection of the reported result
Some concerns about risk of bias
No preregistration or pre-analysis plan is mentioned. The ITT exam score effect (E1) is reported under two missing-data strategies ("Ignore Missingness" +0.91pp and "Impute for Missingness" +0.67pp, both with 90% CIs crossing or nearly crossing zero), and many hypothesis tests were run with post hoc Bonferroni adjustment (x15). Reporting both missingness strategies is transparent, but without a prespecified primary outcome/analysis the possibility of result-driven selection among estimators and outcomes cannot be excluded.
We used Bonferroni corrections to account for multiple hypothesis testing. To apply Bonferroni correction, we multiplied the P-value by 15, representing the number of hypothesis tests we conducted in this study.
02Funding

Funding and conflicts

Funding
Stanford HAI Hoffman-Yee grant (in part); no other funding disclosed in the arXiv version
Vendor funded
Partial
Vendor
None identified
Notes
No conflict-of-interest statement in the arXiv version. All authors are Stanford-affiliated and several are the course's creators/instructors (evaluating their own course context). The study used OpenAI's GPT-4; no OpenAI funding, involvement, or free-credit arrangement is disclosed either way.

Funding is shown on every study and never used to score it.

03Claims

Claims this study bears on

04Source

The source, as retrieved

Abstract

No abstract retrieved.

Where this record came from

SourceRetrievedIdentifier
seedSep 18, 2026link
first ingestion