Generative AI without guardrails can harm learning: Evidence from high school mathematics
- Design
- Cluster-randomized trial · randomized at class level · 1000 participants
- Subject and task
- Math
solving teacher-designed math practice problems (e.g., combinatorics) reviewing previously covered curriculum, followed by a closed-book exam on conceptually similar problems
Exposure: four 90-minute sessions during the Fall 2023 semester (intervention active only in the practice portion of each session) - Population match
- Secondary students: Direct · University students: Different
- Attrition
- No differential attrition: authors checked differential student absenteeism and 'find no differential attrition in student attendance across arms or sessions' (details in SI Appendix B.6). Five class sessions were noncompliant (e.g., laptops did not arrive); handled by intention-to-treat, with a noncomplier-omitting robustness check reported as consistent.
- Study tier
- Study tier 1Preregistered randomized controlled trial with randomization at the classroom level (classrooms assigned to arms via an integer program matching observable characteristics; students randomly assigned to classrooms), analyzed by intention-to-treat with standard errors clustered at the classroom level (the unit of randomization). Large sample (nearly 1,000 students, ~50 classes, 2,848 student-session observations), independent blinded-workload graders with teacher-designed rubrics, covariate balance reported, and no differential attrition across arms. Comparison arm is a credible business-as-usual condition (textbooks/notes, no devices). Main caveats are a single school in one country, short-term (same-session) outcomes only, and constrained rather than simple randomization of classes to arms.
- Assessment
- Version 2 · AI: two passes agreed · Sep 18, 2026
Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.
In a preregistered field experiment at a Turkish high school, about a thousand students in roughly fifty classes did in-class math practice with either a ChatGPT-like GPT-4 chatbot (GPT Base), a guardrailed GPT-4 tutor built with teacher-provided solutions and hints (GPT Tutor), or no AI. With AI available, practice scores rose sharply (48% for GPT Base, 127% for GPT Tutor), but on the closed-book exam that immediately followed, GPT Base students scored 17% worse than students who never had AI, while GPT Tutor students were statistically indistinguishable from control. Message logs suggest students used the unguarded chatbot as a crutch, mostly asking it for answers, and students did not perceive that the AI had hurt their learning.
What the study measured
Arms
| Arm | What it got |
|---|---|
| A · Control (business-as-usual practice with textbooks and course notes, no devices) | Control · teacher supervised Worked the same teacher-designed practice problems with access to course notes and the course textbook only; no laptops or generative AI. Teacher lecture (part 1) and unassisted exam (part 3) identical across arms; teachers did not interact with students during practice or exam. |
| B · GPT Base (standard ChatGPT-like GPT-4 chat interface) | Intervention · teacher supervised GPT Base (custom tool built by the authors) · GPT-4 (OpenAI) · Vanilla chat · chat interface · in class · used Fall semester of the 2023-2024 academic year (Fall 2023) |
| C · GPT Tutor (GPT-4 chat with teacher-designed guardrails: hints without answers, correct solutions and common mistakes in prompt) | Intervention · teacher supervised GPT Tutor (custom tool built by the authors) · GPT-4 (OpenAI) · Guardrailed tutor · tutor scaffold interface · in class · used Fall semester of the 2023-2024 academic year (Fall 2023) |
Findings
A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.
| Finding and conditions | Result |
|---|---|
| F1 · Normalized grade on assisted practice problems (0-1), graded by independent graders using teacher-designed rubric GPT Base (standard ChatGPT-like GPT-4 chat interface) vs Control (business-as-usual practice with textbooks and course notes, no devices)With AI at assessment — assisted performance, not learning evidenceDuring interventionSame contentResearcher-developed testPrespecified | Benefit +0.137 (out of 1) vs. control mean 0.284; a 48% improvement SE 0.031 (HC1 robust, clustered by class) P < 0.01 Teacher-designed practice problems counted toward course grades; graded by hired independent graders balanced across arms per grade-session to reduce grader bias; rubric-based. |
| F2 · Normalized grade on assisted practice problems (0-1) GPT Tutor (GPT-4 chat with teacher-designed guardrails: hints without answers, correct solutions and common mistakes in prompt) vs Control (business-as-usual practice with textbooks and course notes, no devices)With AI at assessment — assisted performance, not learning evidenceDuring interventionSame contentResearcher-developed testPrespecified | Benefit +0.361 (out of 1) vs. control mean 0.284; a 127% improvement SE 0.032 (HC1 robust, clustered by class) P < 0.01 Same practice-problem measure as F1. |
| F3 · Normalized grade on unassisted closed-book, closed-laptop exam (0-1); designated primary preregistered outcome GPT Base (standard ChatGPT-like GPT-4 chat interface) vs Control (business-as-usual practice with textbooks and course notes, no devices)Measured without AIImmediate post-testNear transferResearcher-developed testPrespecified | Harm -0.054 (out of 1) vs. control mean 0.321; a 17% reduction SE 0.022 (HC1 robust, clustered by class) P < 0.05 Exam taken in the same session immediately after practice; each exam problem is conceptually very similar to a paired practice problem (near transfer, very short retention interval). Counted toward final grades; independent rubric-based grading. Preregistration specified pairwise t tests; the regression is described as a variation, with t-test results reported as qualitatively similar in SI A.6. |
| F4 · Normalized grade on unassisted closed-book, closed-laptop exam (0-1); designated primary preregistered outcome GPT Tutor (GPT-4 chat with teacher-designed guardrails: hints without answers, correct solutions and common mistakes in prompt) vs Control (business-as-usual practice with textbooks and course notes, no devices)Measured without AIImmediate post-testNear transferResearcher-developed testPrespecified | Null -0.004 (out of 1); statistically indistinguishable from control SE 0.013 (HC1 robust, clustered by class) not significant (not < 0.05) Same exam measure as F3. Authors note no positive effect was observed despite the large practice-phase advantage. |
| F5 · Preregistered heterogeneous treatment effects by student ability, resources, and effort vs Control (business-as-usual practice with textbooks and course notes, no devices)AI availability at assessment not reported — cannot carry a learning claimImmediate post-testNear transferResearcher-developed testPrespecified | Null limited to no statistically significant heterogeneity, particularly for unassisted exam performance (details in SI Appendix B.4) not_reported not_reported Moderators preregistered; numeric estimates in SI Appendix only, not in fetched main text. |
| F6 · Students' self-reported perception of their exam performance and learning (end-of-session survey) GPT Base (standard ChatGPT-like GPT-4 chat interface) vs Control (business-as-usual practice with textbooks and course notes, no devices)AI availability at assessment not reported — cannot carry a learning claimImmediate post-testTransfer not assessedSelf-reportPrespecification unclear | Null students did not perceive worse performance or learning despite actual 17% exam decrement not_reported not_reported Perception measure; details in SI Appendix B.3. Shows a mismatch between perceived and actual learning. |
| F7 · Students' self-reported perception of their exam performance and learning (end-of-session survey) GPT Tutor (GPT-4 chat with teacher-designed guardrails: hints without answers, correct solutions and common mistakes in prompt) vs Control (business-as-usual practice with textbooks and course notes, no devices)AI availability at assessment not reported — cannot carry a learning claimImmediate post-testTransfer not assessedSelf-reportPrespecification unclear | Mixed students perceived they performed significantly better, although actual exam performance did not differ from control not_reported described as significant; statistic not in fetched main text Perception measure; details in SI Appendix B.3. Coded mixed because the significant self-reported benefit is an overestimate relative to the null objective result. |
- Limitations
- Source-stated: single subject (math) at a single high school in Turkey; deployment in Fall 2023 when generative AI was new and models have since improved; short-term outcomes only (partner-school constraint), so long-term learning not measured; mechanism analyses could be complemented by more controlled experiments; GPT Tutor remains passive rather than proactively engaging students. Observed: exam followed practice within the same 90-minute session, so 'learning' is very short-horizon retention; honors classrooms excluded from main sample (robustness check including them reported as similar); class-to-arm assignment used constrained optimization rather than pure randomization; note published correction at PNAS 122(34):e2518204122 (content of correction not described in fetched text).
Who was studied
- Level
- secondary
- Ages
- 9th, 10th, and 11th grade (approx. 14-17; exact ages not reported)
- Country
- Turkey (large high school; TED Ankara Koleji named as research partner in acknowledgments)
- Prior knowledge
- material previously covered in the course; prior-year normalized GPA used as control (mean 0.82, SD 0.11)
- Selection
- about fifty 9th-11th grade classes at one school; students randomly assigned to classrooms; honors-designated classrooms excluded from main sample
Methodological notes
Extracted from: https://pmc.ncbi.nlm.nih.gov/articles/PMC12232635/ Access: full text (main article body; SI Appendix PDF not retrieved, so appendix-only details such as arm-level ns and exact survey statistics were unavailable) Sample-size note: Stated as 'nearly 1,000 students' across about fifty classes (no exact randomized N in main text; 1000 is the stated approximation). Primary student-level regressions use 2,848 student-session observations; problem-level regressions use 13,484 (practice) and 11,392 (exam) observations. Arm-level ns not reported in main text.
How much each result can be relied on
Risk of bias is judged per result, not per paper: the same study can report one outcome at low risk and another at high. Two reviewers assess each result independently against a written guide, and every domain judgment below carries its reasoning and the sentences it rests on.
Risk of bias is not the same as whether the numbers work. A trial can be well conducted and still print a standard deviation that is arithmetically impossible, or a table whose sign contradicts its own text. Risk of bias has no question for that, so where it was found it is recorded separately below, under Source integrity.
These words are not quality scores. On the RoB 2 scale, High risk of bias is the worst rating a result can receive and Low risk of bias is the best — the opposite of how the same words read on some other scales. Levels are always written out in full here for that reason.
Assessed, not settled (1)
Recorded rather than omitted. Leaving these out would make the appraisal look cleaner than it is, and a disagreement between two careful readers is itself worth knowing.
EXAM
Assessed, not settled · RoB 2 · Two reviewers reached different judgments and nothing has settled it
This result has no overall judgment, because the two independent reviews did not reach one. Both readings are shown. The domains they did agree on are below.
| Where they differ | The two readings |
|---|---|
| D1 · Randomisation process | One reviewer: Some concerns about risk of bias The other: Low risk of bias |
| Domain | Judgment and reasoning |
|---|---|
| D1 · Randomisation process Some concerns about risk of bias | Cluster design: intact classrooms were the unit of assignment to the three arms. The cluster-specific concerns are largely mitigated: students were placed into classrooms by the school's own random assignment process before and independently of the trial, so identification/recruitment of individuals could not be influenced by knowledge of the cluster's arm, and the exam was administered to whole classes within the same sessions. However, the main text says only that the authors 'assigned each classroom' to an arm — it does not describe how the allocation sequence for classrooms was generated (random component, stratification, concealment), and the covariate balance results are relegated to SI Appendix A.4, which could not be checked from the PMC main text. With ~50 clusters across three arms, chance imbalance is a live possibility that cannot be verified here, so the domain cannot be rated Low from the main text alone. Analysis matches the randomization unit: SEs are clustered at the classroom level.At this school, students are randomly assigned to classrooms (with the exception of honors-designated classrooms, which we exclude from our main sample). We assigned each classroom to one of three treatment arms—control, GPT Base, and GPT Tutor. |
| D2 · Deviations from intended interventions Low risk of bias | Effect of assignment. Participants and teachers were necessarily unblinded, but the identified deviations were delivery failures (five class sessions could not use the assigned treatment for logistical reasons such as laptops not arriving), not trial-context deviations likely to bias the effect. The primary analysis is explicitly intention-to-treat, preserving randomization, and a robustness analysis omitting noncompliant sessions (SI Appendix B.1) reportedly yields the same conclusions. The exam portion itself was identical across arms.Five class sessions did not use the assigned treatment due to external circumstances (e.g., laptops did not arrive on time). Our primary specifications use an intention-to-treat analysis—i.e., to preserve randomization, we consider all students in a treatment arm as treated, regardless of whether they actually received that treatment. |
| D3 · Missing outcome data Low risk of bias | The main missingness mechanism in this design is student absence from a session (and hence the exam). The authors directly examined differential absenteeism and report no differential attrition across arms or sessions (detail in SI Appendix B.6, not independently checkable from the main text, but the null finding is stated in the article itself). No other indication that outcome data availability depended on arm.Last, we check whether differential student absenteeism may impact our results, and find no differential attrition in student attendance across arms or sessions; see SI Appendix, Appendix B.6. |
| D4 · Measurement of the outcome Low risk of bias | The outcome is an objectively scored closed-book paper exam, identical across arms, graded by hired independent graders (not the classroom teachers) against a teacher-designed rubric, with grader workloads balanced across arms within each grade-session pair. All students submitted answers on paper, so scripts would not reveal arm membership; no explicit statement of grader blinding appears in the main text, but nothing in a paper exam script identifies the arm, and the grader-assignment scheme further limits any grader-level bias. Measurement method and timing were the same across arms.We hired independent graders to evaluate student performance to reduce potential teacher bias (e.g., self-fulfilling prophecy), and ensured that each grader was assigned a similar number of papers across all three arms in each grade–session pair to reduce potential grader bias. The first and third parts are identical across all treatment arms. |
| D5 · Selection of the reported result Low risk of bias | The trial was preregistered (AsPredicted) with the unassisted exam comparison across arms designated as the primary analysis, and that is exactly the result assessed here and reported in Table 1. The additional mechanism/heterogeneity analyses are presented as secondary and do not displace the prespecified primary comparison. The preregistration document itself was not fetched, so exact correspondence of the analysis specification could not be line-checked, but the reported result matches the stated prespecified primary outcome.We preregistered this RCT with a designated primary analysis of comparing students' unassisted exam outcomes across arms. |
Funding and conflicts
- Funding
- Wharton AI & Analytics Initiative, the Fishman-Davidson Center, and Wharton Global Initiatives
- Vendor funded
- No
- Vendor
- None identified
- Notes
- 'The authors declare no competing interest.' The tutors were custom-built by the authors on OpenAI's GPT-4; no author affiliation with OpenAI is stated. Anonymized data and code deposited at https://github.com/obastani/GenAICanHarmLearning. Seminar feedback acknowledged from participants at Google and Microsoft (feedback venues, not funders).
Funding is shown on every study and never used to score it.
Claims this study bears on
- Unrestricted use of a general-purpose LLM chatbot during practice improves secondary students' subsequent performance on assessments taken without AI.Contradicts · rests on finding F3
Preregistered primary outcome: vanilla GPT-4 chat practice reduced unassisted exam performance 17% vs business-as-usual (cluster RCT, HIGH). Direct test; direction opposite to the claim. RoB for this result is assessed-not-settled on D1 (reporting detail).
- Guardrailed AI tutors (hints, no direct answers, pedagogical prompting) improve secondary students' performance on assessments taken without AI, compared with business-as-usual instruction.Contradicts · rests on finding F4
Guardrailed tutor indistinguishable from business-as-usual on the unassisted same-session exam (-0.004) despite +127% assisted practice: a fairly precise null in a large preregistered trial. One site, one subject, immediate only.
- Guardrails prevent the unassisted-performance harm observed when secondary students practice with answer-giving chatbots.Supports · rests on finding F3
Establishes, within the same trial, the vanilla-chat harm the guardrails prevent.
- Guardrails prevent the unassisted-performance harm observed when secondary students practice with answer-giving chatbots.Supports · rests on finding F4
Under identical conditions the guardrailed arm showed no decrement: the harm eliminated (not reversed) by teacher-designed guardrails. Single study, one product and guardrail design.
- Secondary students with low prior knowledge gain less, or are harmed more, by AI assistance during practice than high-prior-knowledge peers.Consistent, but doesn't test the claim · rests on finding F5
Preregistered heterogeneity analyses found limited-to-no ability moderation — but AI availability at assessment is not_reported for this finding, so per the gate it cannot support or contradict; recorded as context.
The source, as retrieved
Abstract
No abstract retrieved.
Where this record came from
| Source | Retrieved | Identifier |
|---|---|---|
| seed | Sep 18, 2026 | link first ingestion |