[{"id":1,"doi":"10.1073/pnas.2422633122","pmid":null,"arxiv_id":null,"openalex_id":null,"nct_ids":"[]","study_group_id":null,"publication_status":"peer_reviewed","title":"Generative AI without guardrails can harm learning: Evidence from high school mathematics","authors":"[\"Hamsa Bastani\", \"Osbert Bastani\", \"Alp Sungu\", \"Haosen Ge\", \"Özge Kabakcı\", \"Rei Mariman\"]","journal":"Proceedings of the National Academy of Sciences","publication_date":null,"year":2025,"publication_type":"journal_article","peer_reviewed":"yes","abstract":null,"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC12232635/","source_name":"seed","source_tier":1,"coi_statement":null,"assessment_version":2,"assessed_by":"ai:two-pass","assessed_at":"2026-09-18T19:22:56+00:00","study_design":"cluster_rct","drugs":null,"drug_details":null,"domains":null,"outcome_type":null,"primary_outcome":null,"endpoints":null,"effect_estimate":null,"confidence_interval":null,"p_value":null,"sample_size":1000,"follow_up":null,"direction":null,"population":"{\"level\": \"secondary\", \"ages\": \"9th, 10th, and 11th grade (approx. 14-17; exact ages not reported)\", \"country\": \"Turkey (large high school; TED Ankara Koleji named as research partner in acknowledgments)\", \"prior_knowledge\": \"material previously covered in the course; prior-year normalized GPA used as control (mean 0.82, SD 0.11)\", \"selection\": \"about fifty 9th-11th grade classes at one school; students randomly assigned to classrooms; honors-designated classrooms excluded from main sample\"}","applicability":null,"applicability_rationale":null,"randomization_level":"class","subject":"math","learning_task":"solving teacher-designed math practice problems (e.g., combinatorics) reviewing previously covered curriculum, followed by a closed-book exam on conceptually similar problems","exposure_duration":"four 90-minute sessions during the Fall 2023 semester (intervention active only in the practice portion of each session)","population_match_secondary":"direct","population_match_university":"different","attrition":"No differential attrition: authors checked differential student absenteeism and 'find no differential attrition in student attendance across arms or sessions' (details in SI Appendix B.6). Five class sessions were noncompliant (e.g., laptops did not arrive); handled by intention-to-treat, with a noncomplier-omitting robustness check reported as consistent.","adjusted_for":null,"evidence_level":"HIGH","evidence_components":null,"evidence_rationale":"Preregistered randomized controlled trial with randomization at the classroom level (classrooms assigned to arms via an integer program matching observable characteristics; students randomly assigned to classrooms), analyzed by intention-to-treat with standard errors clustered at the classroom level (the unit of randomization). Large sample (nearly 1,000 students, ~50 classes, 2,848 student-session observations), independent blinded-workload graders with teacher-designed rubrics, covariate balance reported, and no differential attrition across arms. Comparison arm is a credible business-as-usual condition (textbooks/notes, no devices). Main caveats are a single school in one country, short-term (same-session) outcomes only, and constrained rather than simple randomization of classes to arms.","funding_source":"Wharton AI & Analytics Initiative, the Fishman-Davidson Center, and Wharton Global Initiatives","industry_funded":"no","manufacturer":null,"author_conflicts":null,"sponsor_role":null,"independent_replication_exists":null,"conflict_notes":"'The authors declare no competing interest.' The tutors were custom-built by the authors on OpenAI's GPT-4; no author affiliation with OpenAI is stated. Anonymized data and code deposited at https://github.com/obastani/GenAICanHarmLearning. Seminar feedback acknowledged from participants at Google and Microsoft (feedback venues, not funders).","adverse_events":null,"limitations":"Source-stated: single subject (math) at a single high school in Turkey; deployment in Fall 2023 when generative AI was new and models have since improved; short-term outcomes only (partner-school constraint), so long-term learning not measured; mechanism analyses could be complemented by more controlled experiments; GPT Tutor remains passive rather than proactively engaging students. Observed: exam followed practice within the same 90-minute session, so 'learning' is very short-horizon retention; honors classrooms excluded from main sample (robustness check including them reported as similar); class-to-arm assignment used constrained optimization rather than pure randomization; note published correction at PNAS 122(34):e2518204122 (content of correction not described in fetched text).","plain_summary":"In a preregistered field experiment at a Turkish high school, about a thousand students in roughly fifty classes did in-class math practice with either a ChatGPT-like GPT-4 chatbot (GPT Base), a guardrailed GPT-4 tutor built with teacher-provided solutions and hints (GPT Tutor), or no AI. With AI available, practice scores rose sharply (48% for GPT Base, 127% for GPT Tutor), but on the closed-book exam that immediately followed, GPT Base students scored 17% worse than students who never had AI, while GPT Tutor students were statistically indistinguishable from control. Message logs suggest students used the unguarded chatbot as a crutch, mostly asking it for answers, and students did not perceive that the AI had hurt their learning.","methodological_notes":"Extracted from: https://pmc.ncbi.nlm.nih.gov/articles/PMC12232635/\nAccess: full text (main article body; SI Appendix PDF not retrieved, so appendix-only details such as arm-level ns and exact survey statistics were unavailable)\nSample-size note: Stated as 'nearly 1,000 students' across about fifty classes (no exact randomized N in main text; 1000 is the stated approximation). Primary student-level regressions use 2,848 student-session observations; problem-level regressions use 13,484 (practice) and 11,392 (exam) observations. Arm-level ns not reported in main text."},{"id":2,"doi":"10.1038/s41598-025-97652-6","pmid":null,"arxiv_id":null,"openalex_id":null,"nct_ids":"[]","study_group_id":null,"publication_status":"peer_reviewed","title":"AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting","authors":"[\"Gregory Kestin\", \"Kelly Miller\", \"Anna Klales\", \"Timothy Milbourne\", \"Gregorio Ponti\"]","journal":"Scientific Reports","publication_date":null,"year":2025,"publication_type":"journal_article","peer_reviewed":"yes","abstract":null,"url":"https://www.nature.com/articles/s41598-025-97652-6","source_name":"seed","source_tier":1,"coi_statement":null,"assessment_version":2,"assessed_by":"ai:two-pass","assessed_at":"2026-09-18T19:22:56+00:00","study_design":"crossover","drugs":null,"drug_details":null,"domains":null,"outcome_type":null,"primary_outcome":null,"endpoints":null,"effect_estimate":null,"confidence_interval":null,"p_value":null,"sample_size":194,"follow_up":null,"direction":null,"population":"{\"level\": \"university\", \"ages\": \"not_reported\", \"country\": \"USA\", \"prior_knowledge\": \"students in an introductory physics course for the life sciences; FCI pretest scores comparable to students at other universities; over 90% reported they had not studied the two lesson topics in depth before the course\", \"selection\": \"consenting students enrolled in Harvard's Physical Sciences 2 course (Fall 2023) who participated in both conditions and completed all pre- and post-tests (194 of 233 enrolled)\"}","applicability":null,"applicability_rationale":null,"randomization_level":"small_group","subject":"physics","learning_task":"introductory college physics lessons (surface tension; fluid flow) taught via activity worksheets with problem-solving","exposure_duration":"two lessons in consecutive weeks (each student experienced one lesson per condition); in-class lesson 60 minutes of learning time, AI-tutored lesson self-paced at home (median 49 minutes)","population_match_secondary":"different","population_match_university":"direct","attrition":"233 enrolled, 194 (83%) eligible and included; exclusions were for lack of consent, non-participation in either condition, or incomplete pre-/post-tests; no differential-attrition analysis reported","adjusted_for":null,"evidence_level":"MODERATE","evidence_components":null,"evidence_rationale":"Randomized crossover trial in an authentic course with within-student comparison, an active-learning (not passive) comparator, controls for prior knowledge, topic, test version, and time on task, and a large, highly significant effect (p < 10^-8; adjusted d = 0.63, quantile-regression estimate 0.73-1.3 SD). However, the AI arm bundles elements the comparison lacks: self-paced at-home study, professionally produced pre-recorded instructor videos, and pre-written step-by-step solutions delivered through a bespoke scaffolded platform, so the contrast is condition-bundle vs. classroom rather than AI per se. Outcomes are immediate researcher-developed post-tests over only two lessons at a single elite institution, with ceiling effects acknowledged, randomization in 2-3 student clusters, and no delayed retention measure; no preregistration is mentioned.","funding_source":"not_reported (no funding statement in the retrieved full text)","industry_funded":"unclear","manufacturer":null,"author_conflicts":null,"sponsor_role":null,"independent_replication_exists":null,"conflict_notes":"The paper states 'The authors declare no competing interests.' The AI tutor platform was conceived, designed, and engineered by author G.K. (in-house academic tool; no AI-vendor ties stated). Authors also note they overlap with the authors of the prior literature validating the in-class active-learning approach used as the comparator.","adverse_events":null,"limitations":"Only two lessons; post-tests immediate, with no delayed/retention measure; ceiling effect on post-test acknowledged by the authors; researcher-developed tests; randomization clustered in 2-3 student peer groups; AI condition differs from control on medium, location (home vs. class), pacing, videos, and pre-written solutions simultaneously; self-report outcomes unblinded; authors state they 'do not presume that structured AI tutoring will always outperform in-class active learning in all contexts, for example, those requiring complex synthesis of multiple concepts and higher-order critical thinking'; whether the AI tutor was accessible during post-tests is never stated; single course at a single institution; per-condition post-test Ns (142 vs 174) unexplained.","plain_summary":"Harvard researchers had 194 students in a large introductory physics course each experience one lesson taught in class with well-implemented active learning and one lesson taught at home by 'PS2 Pal,' a custom GPT-4 tutor engineered with pedagogical prompts, scaffolded question sequences, and pre-written step-by-step solutions. Students scored substantially higher on immediate post-tests after the AI-tutored lessons (median 4.5 vs 3.5; adjusted effect size 0.63, ceiling-corrected 0.73-1.3 SD) while spending less time (median 49 vs 60 minutes), and reported feeling more engaged and motivated, with enjoyment and growth mindset comparable. The AI condition bundled self-pacing, videos, and vetted solutions with the AI itself, and outcomes were immediate tests over just two lessons, so the result shows what a carefully engineered AI-tutoring package can do rather than what generic chatbot use does.","methodological_notes":"Extracted from: https://www.nature.com/articles/s41598-025-97652-6\nAccess: full text (HTML of open-access article, including Abstract, Introduction, Results, Discussion, Methods, Notes, Acknowledgements, author information, and competing-interests declaration; supplementary tables S1-S3 and Supplementary Material 1 not retrieved; no funding statement present in retrieved text)\nSample-size note: 194 analyzed students (of 233 enrolled) in a crossover design; reported post-test observation counts are N = 142 (AI condition) and N = 174 (in-class condition), pre-test baseline N = 316 pooled; the paper does not explain why per-condition post-test Ns are below 194 each"},{"id":3,"doi":"10.1111/bjet.13544","pmid":null,"arxiv_id":"2412.09315","openalex_id":null,"nct_ids":"[]","study_group_id":null,"publication_status":"peer_reviewed","title":"Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance","authors":"[\"Yizhou Fan\", \"Luzhen Tang\", \"Huixiao Le\", \"Kejie Shen\", \"Shufang Tan\", \"Yueying Zhao\", \"Yuan Shen\", \"Xinyu Li\", \"Dragan Gašević\"]","journal":"British Journal of Educational Technology","publication_date":null,"year":2025,"publication_type":"journal_article","peer_reviewed":"yes","abstract":null,"url":"https://bera-journals.onlinelibrary.wiley.com/doi/10.1111/bjet.13544","source_name":"seed","source_tier":1,"coi_statement":null,"assessment_version":2,"assessed_by":"ai:two-pass","assessed_at":"2026-09-18T19:22:56+00:00","study_design":"rct","drugs":null,"drug_details":null,"domains":null,"outcome_type":null,"primary_outcome":null,"endpoints":null,"effect_estimate":null,"confidence_interval":null,"p_value":null,"sample_size":117,"follow_up":null,"direction":null,"population":"{\"level\": \"university\", \"ages\": \"mean 22.61 (SD=3.39); 70% female; 55% undergraduates\", \"country\": \"Not explicitly stated in the arXiv version (participants' first language is redacted as \\\"[disclosed]\\\"); lead authors are at Peking University, Beijing, China, and all participants were English-as-second-language speakers.\", \"prior_knowledge\": \"Not directly reported beyond a pre-test knowledge score (no significant group differences, F=1.294, p=0.281); participants came from \\\"a diverse array of disciplines\\\".\", \"selection\": \"117 university students recruited July-September 2023; recruitment method not reported.\"}","applicability":null,"applicability_rationale":null,"randomization_level":"student","subject":"writing_language","learning_task":"English (L2) reading-and-writing task: read materials on three topics (AI, differentiated teaching, scaffolding teaching) and write, then revise, an essay envisioning the future of education in 2035, scored against a 25-point rubric.","exposure_duration":"Single lab session: 2-hour reading and writing task (stage 1, no differentiated support) followed by a 1-hour revising task (stage 2, in which the group-specific support was available); post-test completed within one day.","population_match_secondary":"different","population_match_university":"direct","attrition":"not_reported as such; analysis Ns vary by outcome (117 for essay scores, 114 for IMI, 97-107 for knowledge/transfer tests) with no explanation of missing data or exclusions in the arXiv text.","adjusted_for":null,"evidence_level":"MODERATE","evidence_components":null,"evidence_rationale":"Randomised experiment with an unusually informative set of comparison conditions (ChatGPT vs. human expert vs. writing-analytics checklist vs. no support), objective trace data, and blinded-rubric essay scoring with high inter-rater reliability on a subsample. However, it is a single ~3-hour lab session with small per-arm samples (24-35), researcher-developed knowledge tests, unexplained variation in analysis Ns and internally inconsistent group sizes across tables, and the headline \"metacognitive laziness\" claim rests partly on descriptive process-mining contrasts without inferential statistics. Null findings on knowledge gain and transfer are imprecise given the sample size.","funding_source":"National Natural Science Foundation of China (Grant 62407001); Society for Learning Analytics Research (ECR Research Grant, 2023).","industry_funded":"no","manufacturer":null,"author_conflicts":null,"sponsor_role":null,"independent_replication_exists":null,"conflict_notes":"not_reported (no conflict-of-interest statement in the arXiv text)","adverse_events":null,"limitations":"Source-stated: possible lack of power from task duration and sample size; 70% female sample limiting representativeness; single reading/writing task; no long-term follow-up; no targeted, mature measure of metacognitive laziness. Observed: group Ns inconsistent between methods text, essay table, IMI table, and knowledge-test tables with no attrition accounting; the appendix \"Essay Score after Revision\" row for CL duplicates its before-revision values (14.367, SD 3.882), an apparent table error; the essay improvement (the only significant performance benefit) was produced while the tool was in use, and authors themselves observed learners copy-pasting ChatGPT-generated sentences to cater to the rubric; knowledge/transfer post-test was completed \"within one day\" after the session, with supervision conditions unstated; process-mining comparisons use a 10% transition-probability threshold, not significance tests; ChatGPT was restricted to task topics and advice-giving, so results may not generalise to unrestricted chatbot use.","plain_summary":"117 university students wrote an English essay in a 3-hour lab session and were randomly given one of four kinds of help for the 1-hour revision stage: a task-restricted ChatGPT-4, a human writing expert, an AI-powered checklist/writing-analytics toolkit, or nothing. The ChatGPT group improved their essays significantly more than every other group - including the human-expert group - but on tests of what they actually learned (knowledge of the topic and transfer to a new domain, taken after the session without the tool) they did no better than anyone else, and motivation did not differ. Behaviour traces showed the ChatGPT group's revising revolved around the chatbot, with relatively fewer self-checking (metacognitive) moves than the human-expert and checklist groups, which the authors call potential \"metacognitive laziness\". In short: the AI made the assisted product better without making the learner better.","methodological_notes":"Extracted from: https://arxiv.org/pdf/2412.09315\nAccess: Full text extracted from the arXiv version (arXiv:2412.09315, PDF converted to text locally), not the Wiley version of record (DOI 10.1111/bjet.13544). The arXiv copy is partially anonymised: participants' first language is \"[disclosed]\" and some self-citations appear as \"Author (2023)\" / \"Authors, 2022, 2023\".\nSample-size note: Methods text: CN (control) 30, AI (ChatGPT) 35, HE (human expert) 25, CL (checklist) 27. Inconsistent with appendix essay-score table (CN 27, AI 35, HE 25, CL 30, total 117) and with the IMI motivation table (CN 27, AI 35, HE 24, CL 28, total 114). Knowledge-test analysis Ns are smaller still (pretest 107, posttest 105, gain 97, transfer 107)."},{"id":4,"doi":"10.1145/3544548.3580919","pmid":null,"arxiv_id":"2302.07427","openalex_id":null,"nct_ids":"[]","study_group_id":null,"publication_status":"peer_reviewed","title":"Studying the effect of AI Code Generators on Supporting Novice Learners in Introductory Programming","authors":"[\"Majeed Kazemitabaar\", \"Justin Chow\", \"Carl Ka To Ma\", \"Barbara J. Ericson\", \"David Weintrop\", \"Tovi Grossman\"]","journal":"CHI Conference on Human Factors in Computing Systems (CHI 2023)","publication_date":null,"year":2023,"publication_type":"journal_article","peer_reviewed":"yes","abstract":null,"url":"https://arxiv.org/abs/2302.07427","source_name":"seed","source_tier":1,"coi_statement":null,"assessment_version":2,"assessed_by":"ai:two-pass","assessed_at":"2026-09-18T19:22:56+00:00","study_design":"rct","drugs":null,"drug_details":null,"domains":null,"outcome_type":null,"primary_outcome":null,"endpoints":null,"effect_estimate":null,"confidence_interval":null,"p_value":null,"sample_size":69,"follow_up":null,"direction":null,"population":"{\"level\": \"secondary\", \"ages\": \"10-17 (M=12.5, SD=1.8)\", \"country\": \"Canada (recruited through coding camps described as located in two major North American cities; acknowledgments name camps in Ottawa and the Greater Toronto Area)\", \"prior_knowledge\": \"No prior text-based programming experience (self-reported); 64 of 69 had used a block-based environment like Scratch or Code.org, and 27 had taken a programming-related class; balanced on a 25-item Scratch pre-test\", \"selection\": \"Volunteers from more than 200 coding-camp sign-ups; 90 reporting no prior text-based programming experience were contacted and started; compensated with a $50 gift card; ethics-board approved\"}","applicability":null,"applicability_rationale":null,"randomization_level":"student","subject":"programming","learning_task":"Introductory Python programming in a self-paced web environment (Coding Steps): 45 two-part tasks (code-authoring followed by code-modification) plus 40 multiple-choice questions across basics, data-types, conditionals, loops, and arrays","exposure_duration":"Ten 90-minute sessions over three consecutive weeks (1 introduction session, 7 training sessions during which the Codex group had AI access on authoring tasks, 2 evaluation sessions); first session was 2 hours","population_match_secondary":"direct","population_match_university":"different","attrition":"21 of 90 starters (23%) did not complete: 11 participated only in the first session, 4 in less than half of the sessions, and 6 missed one of the last two evaluation sessions. Paper states no common dropout factors were identified (disability, native language, computer/internet access) and that completers in the two groups had similar Scratch pre-test means (Codex M=62.7%, Baseline M=60%, t(67)=0.54, p=.67). Attrition counts per assigned condition are not reported.","adjusted_for":null,"evidence_level":"MODERATE","evidence_components":null,"evidence_rationale":"True experiment with student-level random assignment within pre-test-matched pairs and a well-controlled active learning environment identical across arms except for the code generator. However, the sample is small (n=69, 33/36 per arm), outcome measures are researcher-developed and unvalidated, post-tests stayed within trained content (no far transfer), the key retention contrasts are nonsignificant and likely underpowered, and 23% attrition from randomization to analysis is not broken down by condition.","funding_source":"not_reported (no funding statement; acknowledgments thank after-school coding and STEM camps — Coder Sports, CodeZilla, Hive5 Innovative Center, Junior Innovators — for help recruiting)","industry_funded":"unclear","manufacturer":null,"author_conflicts":null,"sponsor_role":null,"independent_replication_exists":null,"conflict_notes":"No conflict-of-interest statement observed. Authors are university researchers (University of Toronto, University of Michigan, University of Maryland). The paper notes \"Coding Steps was approved by the OpenAI App Review team prior to running the study\" — API access approval, not stated as funding or employment ties.","adverse_events":null,"limitations":"Source-stated: correctness scored with the authors' own simple rubric; several between-group differences showed effect-size trends but did not reach significance, possibly due to sample size; post-tests \"did not leave the boundaries of what learners were trained on\" (no far-transfer assessment); most participants were non-native English speakers; study covered only code generation, not other AI-assistant capabilities; qualitative usage analysis left to future work. Observed: 23% attrition between assignment and analysis without per-condition breakdown; no preregistration mentioned; prior-competency moderation analysis is a post-hoc quartile-style split (n=16-18 per cell); retention interval only one week; heavy AI usage patterns (49% of AI-used tasks submitted with AI-generated code unmodified) complicate interpretation of training-phase \"performance.\"","plain_summary":"69 novices aged 10-17 with no prior text-based coding experience learned Python over ten sessions; half could use an OpenAI Codex-based code generator on the code-writing (authoring) tasks during training. With the AI available, that group completed more tasks, wrote more correct code, made fewer errors, and worked faster, and did just as well as the control group on manual code-modification tasks done without the AI. On post-tests taken without any AI (one day and one week after training), the two groups performed similarly overall — the Codex group scored somewhat higher at one-week retention, but not statistically significantly. Learners with higher Scratch pre-test scores who had trained with the AI did significantly better on the retention test than comparable controls.","methodological_notes":"Extracted from: https://arxiv.org/pdf/2302.07427\nAccess: Full text (arXiv version, arXiv:2302.07427v2, 21 Feb 2023), extracted from the arXiv PDF rather than the ACM version of record. PDF downloaded and read page-by-page; quotes transcribed from the rendered pages.\nSample-size note: 69 completers analyzed (21 female/48 male): Codex group n=33, Baseline group n=36. 90 participants started and were pair-matched into groups after the Scratch pre-test."},{"id":5,"doi":"10.7759/cureus.85767","pmid":null,"arxiv_id":null,"openalex_id":null,"nct_ids":"[]","study_group_id":null,"publication_status":"peer_reviewed","title":"ChatGPT as a Learning Tool for Medical Students: Results From a Randomized Controlled Trial","authors":"[\"Kazi A. Kalam\", \"Fatima D. Masoud\", \"Ahmed Muntaser\", \"Rohan Ranga\", \"Xue Geng\", \"Munish Goyal\"]","journal":"Cureus","publication_date":null,"year":2025,"publication_type":"journal_article","peer_reviewed":"yes","abstract":null,"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC12248138/","source_name":"seed","source_tier":1,"coi_statement":null,"assessment_version":2,"assessed_by":"ai:two-pass","assessed_at":"2026-09-18T19:22:56+00:00","study_design":"rct","drugs":null,"drug_details":null,"domains":null,"outcome_type":null,"primary_outcome":null,"endpoints":null,"effect_estimate":null,"confidence_interval":null,"p_value":null,"sample_size":33,"follow_up":null,"direction":null,"population":"{\"level\": \"university\", \"ages\": \"not_reported\", \"country\": \"USA (Washington, DC)\", \"prior_knowledge\": \"first-year MD students in good academic standing; quiz covered material students were expected to have mastered at that point in the semester\", \"selection\": \"volunteers recruited via email and campus bulletin board flyers in April 2025; inclusion: first-year MD enrollment, good academic standing, written informed consent; no exclusion criteria\"}","applicability":null,"applicability_rationale":null,"randomization_level":"student","subject":"other","learning_task":"answering a 10-item multiple-choice quiz on first-year medical curriculum content while consulting an assigned study resource","exposure_duration":"single 15-minute proctored quiz session in Week 1 (resource available only during the quiz; no separate study period)","population_match_secondary":"different","population_match_university":"partial","attrition":"not_reported","adjusted_for":null,"evidence_level":"LOW","evidence_components":null,"evidence_rationale":"True individually randomized trial with two comparison arms, but very small (N=33, 10-12 per arm), single-institution, and unblinded for both participants and researchers. The significant Week 1 result is an open-book contrast (each arm had its assigned resource available during the quiz), so it measures resource-assisted performance rather than learning; the learning-relevant closed-book retention contrast was null and, by the authors' own post hoc calculation, would have needed 72-132 participants for 80% power. No trial registration is mentioned, and precision at this sample size is too low to support any retention conclusion.","funding_source":"None stated; authors declared no financial support was received from any organization for the submitted work","industry_funded":"no","manufacturer":null,"author_conflicts":null,"sponsor_role":null,"independent_replication_exists":null,"conflict_notes":"Authors declared no financial relationships within the previous three years with organizations that might have an interest in the work, and no other relationships or activities that could appear to have influenced it","adverse_events":null,"limitations":"Source-stated: retention comparisons underpowered (N=33 vs 72-132 needed for 80% power); single institution and first-year students only; no baseline academic indicators collected; no blinding of participants or researchers (performance, expectancy, and novelty biases possible); survey outcomes self-reported and subject to response/social desirability bias; no long-term follow-up. Observed: the primary outcome was taken with resources in hand, so the Week 1 \"benefit\" conflates access to an answer-finding tool with learning; the same 10-item quiz was reused in Week 2 (retest exposure); no trial registration reported; ceiling effects likely in Groups A and B on the 10-point Week 1 quiz (means 9.60 and 9.08).","plain_summary":"Thirty-three first-year medical students at Georgetown were randomized to take a 10-question quiz while using ChatGPT-4.0, non-AI internet resources, or their course's own materials. With resources in hand, the ChatGPT and internet groups scored much higher than the institutional-materials group, and ChatGPT was no better than ordinary internet resources. One week later, on the same quiz with no resources allowed, all three groups' scores dropped and no longer differed significantly, so the study found no retention advantage for ChatGPT - though it was far too small to detect modest retention differences.","methodological_notes":"Extracted from: https://pmc.ncbi.nlm.nih.gov/articles/PMC12248138/\nAccess: full text (PMC open-access HTML, including methods, results, discussion, and disclosure statements; appendix survey tables referenced but percentages not extracted)\nSample-size note: Group A (ChatGPT-4.0) n=10; Group B (external non-AI online resources) n=12; Group C (institutional resources) n=11"},{"id":6,"doi":null,"pmid":null,"arxiv_id":null,"openalex_id":null,"nct_ids":"[]","study_group_id":null,"publication_status":"working_paper","title":"From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria","authors":"[\"Martín E. De Simone\", \"Federico Tiberti\", \"Maria Barron Rodriguez\", \"Federico Manolio\", \"Wuraola Mosuro\", \"Eliot Jolomi Dikoru\"]","journal":"World Bank Policy Research Working Paper 11125","publication_date":null,"year":2025,"publication_type":"report","peer_reviewed":"no","abstract":null,"url":"https://documents1.worldbank.org/curated/en/099548105192529324/pdf/IDU-c09f40d8-9ff8-42dc-b315-591157499be7.pdf","source_name":"seed","source_tier":1,"coi_statement":null,"assessment_version":2,"assessed_by":"ai:two-pass","assessed_at":"2026-09-18T19:37:38+00:00","study_design":"rct","drugs":null,"drug_details":null,"domains":null,"outcome_type":null,"primary_outcome":null,"endpoints":null,"effect_estimate":null,"confidence_interval":null,"p_value":null,"sample_size":1328,"follow_up":null,"direction":null,"population":"{\"level\": \"secondary\", \"ages\": \"first-year senior secondary (SS1), typically 15 years old\", \"country\": \"Nigeria (Benin City, Edo State)\", \"prior_knowledge\": \"regular first-year English curriculum students; baseline measured by first- and second-term curricular exam scores; for poorer students often first experience using computers\", \"selection\": \"9 public schools selected for computer-lab availability; all SS1 students invited, 52% expressed interest within a 10-day window; randomization by lottery only among interested volunteers; guardians signed consent\"}","applicability":null,"applicability_rationale":null,"randomization_level":"student","subject":"writing_language","learning_task":"Curriculum-aligned English language practice via dialogue with an LLM chatbot acting as a virtual tutor (students in pairs, teacher-provided starter prompts)","exposure_duration":"6 weeks (June-July 2024; up to twelve 90-minute after-school sessions, max two per week; mean attendance ~72% of sessions)","population_match_secondary":"direct","population_match_university":"different","attrition":"Substantial and differential: 422/657 (~64%) of treatment and 337/671 (~50%) of control completed the endline; the paper states the treatment-control difference in attrition is significant and addresses it with Lee bounds (effects remain positive and significant; e.g. total weighted score bounds 0.255-0.327) and inverse-probability weighting (estimates largely unchanged)","adjusted_for":null,"evidence_level":"PRELIMINARY","evidence_components":null,"evidence_rationale":"Individually randomized RCT (lottery among volunteer students within 9 schools) with a business-as-usual control, school fixed effects, baseline-score control, and robust standard errors; because randomization is at the student level there is no cluster-inference problem, but the design invites spillovers and documented control-group contamination (some control students attended sessions). The sample is self-selected (52% of eligible students volunteered) and endline attrition is large and differential (64% vs 50%), though Lee bounds and IPW checks hold. Effects are precisely estimated (SEs ~0.07). Rated PRELIMINARY because this is a non-peer-reviewed Policy Research Working Paper series entry authored by the implementing World Bank team, notwithstanding the verified reproducibility package.","funding_source":"Mastercard Foundation (financial support acknowledged); study conducted and authored by World Bank staff (Education Global Department) in collaboration with Edo State education authorities","industry_funded":"no","manufacturer":null,"author_conflicts":null,"sponsor_role":null,"independent_replication_exists":null,"conflict_notes":"The tutoring tool was Microsoft Copilot powered by GPT-4, but the paper reports no funding, involvement, or role of Microsoft or OpenAI; the tool was used as a free off-the-shelf product. Funder is the Mastercard Foundation (philanthropy). Authors are World Bank employees evaluating a pilot their own institution helped implement (evaluator-implementer overlap, though not commercial). A verified reproducibility package is available at reproducibility.worldbank.org.","adverse_events":null,"limitations":"Source-stated: student-level randomization may cause spillovers to controls; some control students inadvertently gained access to sessions (teachers unwilling to enforce the distinction early on), and early implementation problems (account creation, internet disruptions, power outages) attenuate estimates; differential attrition at endline (addressed with Lee bounds and IPW); volunteer sample with no demographic data on non-interested students, limiting representativeness checks; the female heterogeneity result may be driven by one girls-only school with low baseline performance; urban schools selected for having computer labs, limiting external validity (rural settings untested); short 6-week duration with no long-term follow-up; year-long extrapolations (1.2-2.2 SD) rest on strong linearity and constant-effect assumptions. Observed: the endline instrument is researcher-commissioned and partly measures content (AI knowledge, digital skills) plausibly taught only to the treatment group; the paper is internally inconsistent about the third-term exam, describing its content in the text as broader than but including intervention-period material while Table 2's note calls its content \"unrelated to the intervention's material\"; treatment group also received extra instructional time and adult supervision, so the AI component is not isolated (authors acknowledge and argue against pure time effects); no pre-registration or pre-analysis plan is mentioned.","plain_summary":"In Benin City, Nigeria, volunteer first-year senior secondary students were randomly chosen to attend up to twelve 90-minute after-school sessions over six weeks in which pairs of students chatted with Microsoft Copilot (GPT-4) as an English tutor, guided by trained teachers using curated prompts. On a pencil-and-paper endline test, selected students scored about 0.31 SD higher overall and 0.24 SD higher in English, and also did about 0.21 SD better on their regular third-term school English exam. Gains were larger for girls, for students with higher baseline scores, and for higher-SES students, and grew roughly linearly with days attended. The authors rate the program among the most cost-effective learning interventions, though the study is a non-peer-reviewed working paper with volunteer participants and substantial differential attrition.","methodological_notes":"Extracted from: https://documents1.worldbank.org/curated/en/099548105192529324/pdf/IDU-c09f40d8-9ff8-42dc-b315-591157499be7.pdf\nAccess: full text (official World Bank PDF downloaded and text-extracted locally; version note on title page says originally published May 2025, this version updated December 2025)\nSample-size note: 1,328 randomized (657 treatment, 671 control); 759 completed the final assessment (422 treatment, 337 control); regression Ns 636-654 depending on outcome (654 for endline assessment scores, 636 for third-term exam)"},{"id":7,"doi":"10.1145/3698205.3733960","pmid":null,"arxiv_id":"2407.09975","openalex_id":null,"nct_ids":"[]","study_group_id":null,"publication_status":"peer_reviewed","title":"The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement But May Increase Adopters' Exam Performances","authors":"[\"Allen Nie\", \"Yash Chandak\", \"Miroslav Suzara\", \"Ali Malik\", \"Juliette Woodrow\", \"Matt Peng\", \"Mehran Sahami\", \"Emma Brunskill\", \"Chris Piech\"]","journal":"Proceedings of the Twelfth ACM Conference on Learning @ Scale","publication_date":null,"year":2025,"publication_type":"journal_article","peer_reviewed":"yes","abstract":null,"url":"https://arxiv.org/abs/2407.09975","source_name":"seed","source_tier":1,"coi_statement":null,"assessment_version":2,"assessed_by":"ai:two-pass","assessed_at":"2026-09-18T20:14:05+00:00","study_design":"rct","drugs":null,"drug_details":null,"domains":null,"outcome_type":null,"primary_outcome":null,"endpoints":null,"effect_estimate":null,"confidence_interval":null,"p_value":null,"sample_size":5831,"follow_up":null,"direction":null,"population":"{\"level\": \"other\", \"ages\": \"Adults, mean 31.4 (SD ~10.4), median 29, range up to 84; 18-22 college-age subgroup N=1,241 of 5,831\", \"country\": \"146 countries (global online course; run by Stanford University, USA)\", \"prior_knowledge\": \"Coding beginners in an intro course; self-reported prior coding experience mean 5.1 (SD 4.4) on a 0-18 scale; 3,180 of 5,831 reported no experience (0-5)\", \"selection\": \"Applicants approved to enroll in the free online class (8,762); the experiment pool was the 5,831 students still active after week 1; exam-score analyses further condition on the 45.8% who chose to take the optional exam\"}","applicability":null,"applicability_rationale":null,"randomization_level":"student","subject":"programming","learning_task":"Introductory Python programming in Stanford's free 6-week online course Code in Place (spring 2023); treatment was access to a course-specific GPT-4 chat interface","exposure_duration":"6-week course (April 24 - June 5, 2023); GPT-4 access offered from the start of week 4 (May 15), 9 days before the optional diagnostic exam (May 24-26); Figure 1 labels the experiment period as May 15-24, and the paper does not state that access was revoked afterward","population_match_secondary":"different","population_match_university":"partial","attrition":"No further post-randomization dropout is reported beyond the course's normal voluntary disengagement, which is itself the outcome: 54.2% of randomized students skipped the optional exam, differentially by arm (55.9% experiment vs 51.5% control). Homework completion declined in both arms across weeks (roughly 79% week 2 to ~45-50% week 6). An after-course survey found around 2% of students said they used ChatGPT outside the class.","adjusted_for":null,"evidence_level":"MODERATE","evidence_components":null,"evidence_rationale":"Large student-level randomized encouragement trial (N=5,831): the OFFER of GPT-4 was randomized, so the engagement ITT effects (exam participation, homework completion) are well identified, precisely estimated, and survive the authors' Bonferroni correction across 15 tests. The exam-performance conclusions are much weaker: the ITT effect on scores is null, and the headline adopter benefit is a LATE for the 14.2% self-selected compliers, resting on unverifiable exclusion and monotonicity assumptions plus ML imputation for the >50% missing exam outcomes, and the authors state it is not statistically significant after multiple-hypothesis adjustment. No preregistration is mentioned, and subgroup results (low-HDI, age, experience) are explicitly exploratory.","funding_source":"Stanford HAI Hoffman-Yee grant (in part); no other funding disclosed in the arXiv version","industry_funded":"partial","manufacturer":null,"author_conflicts":null,"sponsor_role":null,"independent_replication_exists":null,"conflict_notes":"No conflict-of-interest statement in the arXiv version. All authors are Stanford-affiliated and several are the course's creators/instructors (evaluating their own course context). The study used OpenAI's GPT-4; no OpenAI funding, involvement, or free-credit arrangement is disclosed either way.","adverse_events":null,"limitations":"Authors' stated limitations: a particular course, cohort, and moment in time (mid-2023 AI zeitgeist); at-will free course with no grade/credit stakes, so drop-out is easy; the course centers on human teachers and may attract students who prefer human instruction; mechanisms may be specific to programming; UX choices (email + sidebar advertisement, mandatory pros/cons handout, non-integrated interface) may have shaped outcomes. Additional limitations: optional exam creates outcome missingness correlated with treatment (their engagement effect contaminates the score analysis, addressed only by MCAR or missing-at-random imputation assumptions); adopter effect is self-selected compliers only (adopters were observably more engaged, e.g. higher prior section attendance); low adoption (14.2%) limits power; guardrail robustness against solution-seeking was not thoroughly tested; use of public ChatGPT outside the class could not be verified; figure vs text p-values differ slightly for exam participation (caption P=0.006 vs text corrected P=.020).","plain_summary":"In a free 6-week global online Python course with 5,831 active students, researchers randomly offered 60% of students a course-specific GPT-4 chat tool partway through the course. Only 14.2% of offered students used it, and simply offering and advertising it significantly REDUCED engagement: exam participation fell 4.3 percentage points and week-6 homework completion fell 4.6 points, though students from low-HDI countries showed the opposite (exploratory) pattern. Offering the tool did not change average exam scores overall; a causal (instrumental variable) estimate suggests the self-selected users may have scored about 6.8 points higher than they would have without GPT-4, but the authors note this is not statistically significant after adjusting for multiple comparisons.","methodological_notes":"Extracted from: https://arxiv.org/pdf/2407.09975\nAccess: arXiv preprint v2 (arXiv:2407.09975v2 [cs.CY], 15 Jul 2025), fetched as PDF and read in full (main text pp. 1-20; appendix not extracted). This is the arXiv version, not the ACM Learning@Scale 2025 version of record (DOI 10.1145/3698205.3733960). arXiv title reads \"...Reduced Engagement but Increased Adopters' Exam Performances\"; the ACM title uses \"May Increase\".\nSample-size note: 5,831 active-after-week-1 students randomized (3,581 experiment / 2,250 control, a 60/40 split) out of 8,762 approved enrollees. Engagement outcomes analyzed on the full randomized pool. Exam scores observed only for the 2,668 exam takers (experiment N=1,579, 44.1%; control N=1,089, 48.5%); only 510 experiment students (14.2%) actually used GPT-4 (\"adopters\"). Score analyses reported both ignoring missingness and with ML imputation for non-takers."},{"id":8,"doi":"10.1186/s41239-023-00425-2","pmid":null,"arxiv_id":null,"openalex_id":null,"nct_ids":"[]","study_group_id":null,"publication_status":"peer_reviewed","title":"AI-generated feedback on writing: insights into efficacy and ENL student preference","authors":"[\"Juan Escalante\", \"Austin Pack\", \"Alex Barrett\"]","journal":"International Journal of Educational Technology in Higher Education","publication_date":null,"year":2023,"publication_type":"journal_article","peer_reviewed":"yes","abstract":null,"url":"https://link.springer.com/article/10.1186/s41239-023-00425-2","source_name":"seed","source_tier":1,"coi_statement":null,"assessment_version":2,"assessed_by":"ai:two-pass","assessed_at":"2026-09-18T19:37:38+00:00","study_design":"quasi","drugs":null,"drug_details":null,"domains":null,"outcome_type":null,"primary_outcome":null,"endpoints":null,"effect_estimate":null,"confidence_interval":null,"p_value":null,"sample_size":48,"follow_up":null,"direction":null,"population":"{\"level\": \"university\", \"ages\": \"Study 1: 20-30; Study 2: 19-36\", \"country\": \"not named; small liberal arts university in the Asia-Pacific region (authors affiliated with Brigham Young University-Hawaii)\", \"prior_knowledge\": \"At least CEFR B1 English proficiency per the institution's English Language Admission Test; enrolled in an academic reading and writing language course\", \"selection\": \"Non-probability self-selection recruitment of 91 participants across both studies; voluntary consent, no academic credit offered\"}","applicability":null,"applicability_rationale":null,"randomization_level":"none","subject":"writing_language","learning_task":"Academic English paragraph writing (300-word source-integrated paragraphs) by university English-as-a-new-language (ENL) learners; weekly draft-feedback-revise cycles\n","exposure_duration":"6 weeks","population_match_secondary":"different","population_match_university":"direct","attrition":"Study 1: \"The scores of one student who only completed the posttest were not included in the analysis.\" No other Study 1 attrition reported. Study 2: weekly survey completion varied from 32 to 41 of 43 (average 37.7); fewer students completed the survey in the final week.\n","adjusted_for":null,"evidence_level":"LOW","evidence_components":null,"evidence_rationale":"Quasi-experimental two-group pretest-posttest design; the paper says participants \"were divided into two groups\" with no randomization described and self-selected enrollment, so baseline equivalence is not guaranteed (pretest difference favored EG, d=0.520, p=0.078). The comparison arm is genuinely active (weekly 30-min one-on-one sessions with a trained, CRLA-certified human tutor), which strengthens interpretability of the contrast, but modality differs (face-to-face dialogue vs emailed written feedback). The headline result is a nonsignificant group-by-time interaction at N=48 (23 vs 25); at this sample size a null difference is weak evidence of equivalence between AI and human tutor feedback, and the interaction was near-threshold (p=0.085) with the descriptive gain favoring the human-tutor group (6.86 vs 5.848).\n","funding_source":"None: no specific grant from any public, commercial, or not-for-profit funding agency","industry_funded":"no","manufacturer":null,"author_conflicts":null,"sponsor_role":null,"independent_replication_exists":null,"conflict_notes":"Authors declared no competing interests. Authors are university faculty (BYU-Hawaii, Florida State); no AI-industry ties disclosed.\n","adverse_events":null,"limitations":"No dedicated limitations section. Author-acknowledged constraints: students could not ask follow-up questions of the AI (access was deliberately limited; a TA generated and emailed the feedback), and GPT-4 is \"not optimized for AWE purposes\" (e.g., no text annotation). Additional reviewer-noted limitations: non-random assignment and self-selection; small N; posttest CG scores non-normal (Shapiro-Wilk p<0.001); rubric was researcher/institution-developed; feedback modality confounded with source (spoken interactive tutoring vs written emailed AI feedback); single site and single shortened semester; per-protocol handling of the one incomplete case.\n","plain_summary":"University ENL students in an academic writing course got weekly feedback on 300-word paragraph drafts for six weeks, either from ChatGPT (GPT-4, via a teaching-assistant-run engineered prompt, emailed to students) or from a trained human tutor in weekly 30-minute one-on-one sessions. Both groups improved substantially from a diagnostic pretest to a final writing exam scored by two blinded-to-nothing-stated human raters on a 40-point rubric, and neither the group-by-time interaction nor the between-group contrasts were statistically significant, so AI feedback was neither better nor detectably worse than human tutor feedback at this small sample size. A separate group of 43 students who received both kinds of feedback split roughly evenly on which they preferred, rating both highly.\n","methodological_notes":"Extracted from: https://link.springer.com/article/10.1186/s41239-023-00425-2\nAccess: Open-access full text (CC BY) read via the Springer Nature Link HTML page in a browser session (direct WebFetch/curl blocked by Springer's cookie/bot challenge). Full article body, funding, and competing-interests sections captured; Tables 1 and 3 read from their dedicated table pages (per-group Ns, means, SDs, t-tests). Table 2 (RM-ANOVA) statistics taken from the results text. Appendices A-C (survey items, example prompt/feedback) not retrieved.\n\nSample-size note: Study 1 (learning outcomes): N=48 analyzed (EG/AI n=23, CG/human tutor n=25, per Tables 1 and 3). Study 2 (preference, different participants): N=43, with complete weekly questionnaire responses ranging 32-41 (mean 37.7) across the six weeks. 91 recruited across both studies.\n"},{"id":9,"doi":"10.1016/j.compedu.2023.104967","pmid":null,"arxiv_id":null,"openalex_id":null,"nct_ids":"[]","study_group_id":null,"publication_status":"peer_reviewed","title":"Impact of AI assistance on student agency","authors":"[\"Ali Darvishi\", \"Hassan Khosravi\", \"Shazia Sadiq\", \"Dragan Gašević\", \"George Siemens\"]","journal":"Computers & Education","publication_date":null,"year":2024,"publication_type":"journal_article","peer_reviewed":"yes","abstract":null,"url":"https://www.sciencedirect.com/science/article/pii/S0360131523002440","source_name":"seed","source_tier":1,"coi_statement":null,"assessment_version":2,"assessed_by":"ai:two-pass","assessed_at":"2026-09-18T20:14:05+00:00","study_design":"rct","drugs":null,"drug_details":null,"domains":null,"outcome_type":null,"primary_outcome":null,"endpoints":null,"effect_estimate":null,"confidence_interval":null,"p_value":null,"sample_size":1625,"follow_up":null,"direction":null,"population":"{\"level\": \"university (undergraduate)\", \"ages\": \"not_reported\", \"country\": \"Australia (The University of Queensland; RiPPLE courses)\", \"prior_knowledge\": \"not_reported; all groups statistically equivalent on all six peer-review measures during the shared 4-week AI-prompt phase\", \"selection\": \"Undergraduates in 10 courses using RiPPLE in second semester 2020, spanning Humanities & Social Sciences, Health & Behavioural Sciences, Medicine, Business, and Engineering; only students who consented to data use were included\"}","applicability":null,"applicability_rationale":null,"randomization_level":"student","subject":"other","learning_task":"Peer review of student-created learning resources in RiPPLE (a learnersourcing platform): students rated resources on a 4-item rubric (alignment, correctness, difficulty, critical thinking) and wrote justification/feedback comments to the author. AI quality-control functions (rule-based suggestion detection, SBERT relatedness score, GLEU similarity score) flagged low-quality comments and prompted revision.\n","exposure_duration":"8 weeks total (first 8 weeks of semester 2, 2020). Phase 1 (weeks 1-4): all participants received AI prompts during peer review. Phase 2 (weeks 5-8): random assignment to four conditions (AI continued / NR no prompts / SR self-monitoring checklist / SAI checklist plus prompts).\n","population_match_secondary":"different","population_match_university":"direct","attrition":"not_reported as a flow statement; group counts are identical in the phase-1 and phase-2 summary tables (396/409/402/418, total 1625), indicating the same consenting cohort was analysed in both phases, but per-student dropout or reduced review activity is not reported (phase-2 review volume is lower than phase 1: 11,243 vs 16,007 reviews).\n","adjusted_for":null,"evidence_level":"MODERATE","evidence_components":null,"evidence_rationale":"Randomised controlled field experiment with individual-level assignment of 1625 students to four conditions after a shared four-week AI-scaffold phase, with demonstrated baseline equivalence on all six measures during that phase and large samples per arm giving good precision (one-way ANOVA plus Tukey HSD pairwise tests with CIs and Cohen's d). Rated MODERATE rather than HIGH because outcomes are automated process/quality proxies for peer-review quality (flag rate, SBERT relatedness, GLEU similarity, length, time, likes) partly generated by the same algorithms that powered the intervention prompts, not independent learning assessments; no preregistration is reported; and the design lacks a never-assisted arm, so \"reliance\" is inferred from the withdrawal contrast against a still-assisted control.\n","funding_source":"Australian Government, Australian Research Council Industrial Transformation Training Centre for Information Resilience (CIRES), project IC200100022 (partial support)","industry_funded":"no","manufacturer":null,"author_conflicts":null,"sponsor_role":null,"independent_replication_exists":null,"conflict_notes":"No conflict-of-interest statement observed in the article. The AI-assistance functions were built by the research team into RiPPLE, an adaptive platform associated with author Khosravi's group at UQ, so the evaluated tool is researcher-developed. Data availability: \"The authors do not have permission to share data.\"\n","adverse_events":null,"limitations":"Authors note: experiment limited to the first eight weeks of semester (no longer-term follow-up); the NR comparison group had received AI assistance for the first four weeks, so there is no never-assisted comparison; reliance on quantitative behavioural measures without qualitative interview/open-ended data; the null SAI result may reflect a design issue (AI prompts and checklist not tailored to the hybrid condition); biases of AI-generated feedback not thoroughly investigated; course content, instructional methods, motivation and instructor preferences vary across the 10 courses; generalisability beyond peer feedback tasks untested.\n","plain_summary":"In 10 Australian university courses, 1625 students wrote peer reviews with AI quality prompts for four weeks, then were randomly split into four groups: prompts kept, prompts removed, prompts replaced by a self-monitoring checklist, or checklist plus prompts. When the AI prompts were removed, review quality dropped (more flagged reviews, more generic/repetitive and less relevant, shorter comments), suggesting students had relied on the AI rather than learning from it. The self-monitoring checklist partly, but not fully, compensated for the removed AI, and adding the checklist on top of AI prompts was no better than AI prompts alone.\n","methodological_notes":"Extracted from: https://researchmgt.monash.edu/ws/portalfiles/portal/572678646/564026501_oa.pdf\nAccess: full text (Monash University institutional repository, open-access publisher PDF, CC BY)\nSample-size note: 1625 students across 10 courses; phase 2 group sizes AI n=396, NR n=409, SR n=402, SAI n=418; 16,007 peer reviews on 4,501 resources in phase 1 and 11,243 peer reviews on 3,573 resources in phase 2.\n"},{"id":10,"doi":"10.1016/j.chb.2024.108386","pmid":null,"arxiv_id":null,"openalex_id":null,"nct_ids":"[]","study_group_id":null,"publication_status":"peer_reviewed","title":"Cognitive ease at a cost: LLMs reduce mental effort but compromise depth in student scientific inquiry","authors":"[\"Matthias Stadler\", \"Maria Bannert\", \"Michael Sailer\"]","journal":"Computers in Human Behavior","publication_date":null,"year":2024,"publication_type":"journal_article","peer_reviewed":"yes","abstract":null,"url":"https://opus.bibliothek.uni-augsburg.de/opus4/files/114673/114673.pdf","source_name":"seed","source_tier":1,"coi_statement":null,"assessment_version":2,"assessed_by":"ai:two-pass","assessed_at":"2026-09-18T20:14:05+00:00","study_design":"lab","drugs":null,"drug_details":null,"domains":null,"outcome_type":null,"primary_outcome":null,"endpoints":null,"effect_estimate":null,"confidence_interval":null,"p_value":null,"sample_size":91,"follow_up":null,"direction":null,"population":"{\"level\": \"university\", \"ages\": \"M = 22.3 years (SD = 4.11)\", \"country\": \"Germany\", \"prior_knowledge\": \"Some prior nanotechnology knowledge (M = 2.96, SD = 1.52, of 8 on an adapted Public Knowledge of Nanotechnology test); medicine, pharmacy, and biology students excluded a priori because of potential topic knowledge; groups did not differ on prior knowledge (t(89) = -0.28, p = 0.777, d = -0.06).\\n\", \"selection\": \"Convenience sample of students from various academic programs at \\\"a prestigious German university\\\" (April-May 2023); participation credits offered as compensation for some programs (e.g., psychology majors and minors); 67 female, 24 male in the analyzed sample.\\n\"}","applicability":null,"applicability_rationale":null,"randomization_level":"student","subject":"other","learning_task":"20-minute individual web research on an unsettled socio-scientific issue (safety of zinc-oxide/titanium-dioxide nanoparticles in sunscreen) to advise a fictitious friend, followed by writing a recommendation with justifications from memory (no notes, webpages, or ChatGPT conversations available while writing).\n","exposure_duration":"Single lab session; exactly 20 minutes of tool-assisted research, plus demographic questions, cognitive-load questionnaire, recommendation writing, and a prior-knowledge test afterward.\n","population_match_secondary":"different","population_match_university":"direct","attrition":"One of 92 participants excluded for not following instructions (used both an LLM and a search engine); no other attrition reported — all measures were completed within the single session.\n","adjusted_for":null,"evidence_level":"MODERATE","evidence_components":null,"evidence_rationale":"Individually randomized single-session lab experiment (~20 min exposure) with an active comparison (ChatGPT-3.5 vs Google search) and a prior-knowledge covariate and randomization check. Outcomes are proximal: self-reported cognitive load about the task and a researcher-coded justification written immediately afterward (notably WITHOUT tool access while writing); there is no delayed retention, transfer, or validated learning test. Sample is small (n = 91), non-representative, and effects, while moderate-to-large and consistent, come from one site and one task, so estimates are imprecise and generalization is limited.\n","funding_source":"not_reported — the paper contains no funding or acknowledgements statement (only CRediT authorship, declaration of competing interest, and an OSF data-availability note; materials at https://osf.io/jpxyt).\n","industry_funded":"unclear","manufacturer":null,"author_conflicts":null,"sponsor_role":null,"independent_replication_exists":null,"conflict_notes":"Authors declare no known competing financial interests or personal relationships; no funding statement is given, so funding source cannot be determined from the paper.\n","adverse_events":null,"limitations":"Authors note: no think-aloud protocol or search-log analysis, so cognitive/ metacognitive processes and strategies are not observed; prompting skill and prior LLM experience unexamined as moderators; fixed 20-minute window may have been too long for the LLM group; small, non-representative, university-only sample with presumably high digital literacy; artificial constraint to a single tool (real learning mixes sources, and search engines now embed LLMs); knowledge test given after the task, so scores could be inflated by the search itself.\n","plain_summary":"91 German university students were randomly assigned to research the safety of nanoparticles in sunscreen for 20 minutes using either ChatGPT-3.5 or Google search, then wrote a recommendation with justifications from memory. ChatGPT users reported substantially lower intrinsic, extraneous, and germane cognitive load, but their written justifications contained significantly fewer relevant arguments than those of the Google group, and an exploratory analysis suggested the quality gap was fully mediated by reduced germane (schema-building) load. The diversity of the actual recommendations (for/neutral/against) did not differ between groups.\n","methodological_notes":"Extracted from: https://opus.bibliothek.uni-augsburg.de/opus4/files/114673/114673.pdf\nAccess: Direct fetch attempts (curl, WebFetch) from this session timed out / were refused, but the completed PDF download from this exact URL was present in the session scratchpad (stadler2024.pdf, 1,315,728 bytes). Verified the PDF header, title, author list, and DOI (10.1016/j.chb.2024.108386) against the assignment, then produced an independent text extraction with pdftotext (stadler2024-passb.txt) and worked only from that extraction. No repository files, other passes, or press coverage were read. Full paper text available (CC BY 4.0 Augsburg OPUS copy of the CHB open-access article).\n\nSample-size note: 92 recruited; 1 excluded for using both tools; analyzed n = 91 (web search n = 47, LLM n = 44).\n"},{"id":11,"doi":"10.1145/3632620.3671116","pmid":null,"arxiv_id":"2405.17739","openalex_id":null,"nct_ids":"[]","study_group_id":null,"publication_status":"peer_reviewed","title":"The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers","authors":"[\"James Prather\", \"Brent N. Reeves\", \"Juho Leinonen\", \"Stephen MacNeil\", \"Arisoa S. Randrianasolo\", \"Brett A. Becker\", \"Bailey Kimmel\", \"Jared Wright\", \"Ben Briggs\"]","journal":"ACM Conference on International Computing Education Research (ICER 2024)","publication_date":null,"year":2024,"publication_type":"journal_article","peer_reviewed":"yes","abstract":null,"url":"https://arxiv.org/abs/2405.17739","source_name":"seed","source_tier":1,"coi_statement":null,"assessment_version":2,"assessed_by":"ai:two-pass","assessed_at":"2026-09-18T20:14:05+00:00","study_design":"qualitative","drugs":null,"drug_details":null,"domains":null,"outcome_type":null,"primary_outcome":null,"endpoints":null,"effect_estimate":null,"confidence_interval":null,"p_value":null,"sample_size":21,"follow_up":null,"direction":null,"population":"{\"level\": \"university\", \"ages\": \"not reported (age was collected in the post-session interview but no distribution is reported)\", \"country\": \"USA\", \"prior_knowledge\": \"Novices enrolled in a CS1 course (C++); loops had been introduced two weeks before the study. The course used GenAI (Copilot, ChatGPT) openly from the first day of class, modeled by the professor.\", \"selection\": \"Opt-in for extra credit: 21 of 27 students enrolled in the CS1 course at a small research university participated; non-participants were offered an alternative extra-credit task. 7 participants identified as women, 14 as men; 3 African-American, 2 Hispanic, 1 racially/ethnically Jewish, 15 white/Caucasian (2 of these from Europe). Institution described as growing toward Hispanic-Serving Institution status.\"}","applicability":null,"applicability_rationale":null,"randomization_level":"none","subject":"programming","learning_task":"Solving one CS1 C++ programming problem (\"More Positive or Negative\": count whether more positive or negative integers were entered, loop with sentinel 0) with GitHub Copilot and ChatGPT available, submitted to the Athene automated assessment tool; replication of Prather et al.'s pre-GenAI think-aloud study.","exposure_duration":"Single lab session: warm-up \"Hello World\" task, then up to 35 minutes on the target problem (observed solve times 5 to 35 minutes, mean 17.1, SD 8.1), followed by a post-session interview.","population_match_secondary":"different","population_match_university":"direct","attrition":"No dropout reported among the 21 participants; all completed the session (1 of 21 did not produce a working program within the 35-minute limit). 6 of 27 enrolled students chose not to participate, including 1 student who identified as African-American and 2 who identified as Hispanic.","adjusted_for":null,"evidence_level":"LOW","evidence_components":null,"evidence_rationale":"Small single-site observational lab study (n=21) with no comparison group and no randomization; the pre-GenAI baseline is a separate prior cohort (n=31), so causal claims about GenAI effects cannot be supported. However, as mechanism evidence it is comparatively strong for its genre: triangulated think-aloud, screen replay, and eye-tracking data; a replication protocol reusing the prior study's task, time limit, and codebook; dual coding with Cohen's Kappa 0.74; and correlational statistics with reported p-values. Findings should be read as hypotheses about how GenAI interacts with novice metacognition, not as causal effect estimates on learning outcomes (the authors state they did not measure learning outcomes).","funding_source":"\"This research is funded by the Google Award for Inclusion Research Program and was also partially supported by the Research Council of Finland (Academy Research Fellow grant number 356114).\"","industry_funded":"partial","manufacturer":null,"author_conflicts":null,"sponsor_role":null,"independent_replication_exists":null,"conflict_notes":"Google (industry) award funding acknowledged, though Google does not make the studied tools (GitHub Copilot, ChatGPT). No conflict-of-interest statement in the preprint. The course professor was one of the researchers, and participation was incentivized with extra credit in that professor's course.","adverse_events":null,"limitations":"Authors' stated limitations: small sample (21 vs 31 in the original study), limiting generalizability; only two GenAI tools (ChatGPT and GitHub Copilot), not exhaustive of available tools; single site. Additional observations: no concurrent control condition (comparison is to a prior-cohort study); completion of the task with AI available is an assisted measure that the authors themselves argue overstates unassisted competence; self-reported AI/programming experience; text reports p=0.268 for the grade x old-difficulties correlation while Table 1 reports p=0.02682 for the same pair (apparent typo).","plain_summary":"Researchers watched 21 first-semester programming students solve a standard beginner problem while using GitHub Copilot and ChatGPT, with think-aloud, eye tracking, and interviews, replicating an earlier study run before generative AI existed. Almost everyone (20 of 21) finished the problem, versus 20 of 31 in the earlier no-AI study, but about half struggled: all five previously known metacognitive difficulties reappeared and three new AI-driven ones emerged (constant interruption by suggestions, being misled down wrong paths, and being conceptually behind while feeling confident). Students with lower grades and lower self-efficacy showed more of these difficulties and accepted more AI suggestions, while stronger students used the AI to speed up code they already intended to write. The authors conclude that AI let struggling students finish while leaving them with an illusion of competence, potentially widening the gap between well-prepared and under-prepared students.","methodological_notes":"Extracted from: https://arxiv.org/pdf/2405.17739\nAccess: arXiv preprint 2405.17739, version v1 (submitted 28 May 2024), cs.AI; author manuscript of the ICER 2024 paper (\"Accepted to ICER 2024\"; ACM copyright block present, DOI shown as placeholder XXXXXXX in this preprint version). Full 23+ page PDF read directly (fetched via arXiv, saved locally by the fetch tool). Abstract page https://arxiv.org/abs/2405.17739 also consulted for version metadata.\nSample-size note: 21 lab sessions, one per participant (21 of 27 enrolled CS1 students opted in). Original study being replicated had n=31. Analyses combine observation notes, think-aloud verbalizations, eye tracking, interviews, Copilot suggestion accept rates, weekly quiz grades, and MSLQ self-efficacy scores."},{"id":12,"doi":"10.1057/s41599-026-07019-z","pmid":null,"arxiv_id":null,"openalex_id":null,"nct_ids":"[]","study_group_id":null,"publication_status":"peer_reviewed","title":"ChatGPT's impact on student learning outcomes: a meta-analysis of 35 experimental studies","authors":"[\"Xinning Wu\", \"Pei Zhu\", \"Jinliang Zhang\", \"Mengwei Yin\", \"Yingxi Wang\"]","journal":"Humanities and Social Sciences Communications","publication_date":null,"year":2026,"publication_type":"journal_article","peer_reviewed":"yes","abstract":null,"url":"https://www.nature.com/articles/s41599-026-07019-z","source_name":"seed","source_tier":1,"coi_statement":null,"assessment_version":2,"assessed_by":"ai:two-pass","assessed_at":"2026-09-18T19:37:38+00:00","study_design":"meta_analysis","drugs":null,"drug_details":null,"domains":null,"outcome_type":null,"primary_outcome":null,"endpoints":null,"effect_estimate":null,"confidence_interval":null,"p_value":null,"sample_size":4193,"follow_up":null,"direction":null,"population":"{\"level\": \"mixed\", \"ages\": \"Not reported as ages; grade bands: primary education (k=2), secondary education (k=7), higher education (k=26).\", \"country\": \"Multi-country; not itemized. 28 English-language studies (80%) and 7 Chinese-language studies (20%); Chinese literature deliberately added to cover ChatGPT-restricted contexts.\", \"prior_knowledge\": \"Not reported at the pooled level; discussion notes inexperienced students struggle to prompt ChatGPT effectively.\", \"selection\": \"Included studies had to (1) use ChatGPT as a direct or supported learning tool with focus on student learning outcomes; (2) use experimental or quasi-experimental designs; (3) have experimental and control groups or pre-/post-test assessments; (4) sample mainstream primary, secondary, or college students (special education students, adult learners excluded); (5) report sufficient data to compute effect sizes. Search window November 2022 - June 2024; 2038 deduplicated records screened to 35 included articles.\"}","applicability":null,"applicability_rationale":null,"randomization_level":null,"subject":"other","learning_task":"Pools student learning outcomes from 35 experimental and quasi-experimental ChatGPT studies across many subjects (physics, chemistry, English, mathematics, computer science, teaching skills, literature, science, interdisciplinary/STEM, and others). Outcomes span a cognitive dimension (learning achievement, critical thinking, problem-solving, creative thinking, social skills) and a non-cognitive dimension (learning interest, self-efficacy, learning engagement, learning motivation), the latter largely self-report constructs.","exposure_duration":"Included interventions coded into three bands: less than one month (n=7), one to three months (n=18), and more than three months (n=7). Individual studies cited in the review ranged from ~10 days to 15-16 weeks; the paper notes definitions of \"long-term\" varied from 3 to 12 months across studies.","population_match_secondary":"partial","population_match_university":"partial","attrition":"not_reported","adjusted_for":null,"evidence_level":"LOW","evidence_components":null,"evidence_rationale":"Multi-database search (Web of Science, Wiley, SpringerLink, ProQuest Education, Elsevier ScienceDirect, CNKI) with PRISMA screening, dual independent coding (kappa = 0.851), a 7-point quality checklist (mean 6.11/7), sensitivity analyses, and a three-method publication-bias assessment are genuine strengths. However, heterogeneity is very high (I2 = 91.4%, Q = 409.067, p < 0.01) and largely unexplained; 134 effect sizes from 35 studies are analyzed in CMA 3.0 without meta-regression or three-level modeling (the authors themselves flag this), so effect-size dependency is not handled. Critically for our question, the meta-analysis pools post-test outcomes without ever distinguishing whether ChatGPT was available to students during outcome assessment, and it mixes objective achievement measures with self-report non-cognitive scales in the overall g = 0.670. Quasi-experiments are pooled with experiments, and the search window (Nov 2022 - Jun 2024) captures only early, mostly short studies.","funding_source":"General Project of the National Social Science Fund (Education), Grant No. BIA250124, and the Scientific Research Fund of Hunan Provincial Education Department, Grant No. 25A0349 (Chinese government/academic funding).","industry_funded":"no","manufacturer":null,"author_conflicts":null,"sponsor_role":null,"independent_replication_exists":null,"conflict_notes":"The authors declare no competing interests. No AI-industry involvement disclosed.","adverse_events":null,"limitations":"Source-stated: only published peer-reviewed literature (possible publication bias toward positive findings); moderator coverage incomplete (intervention setting, ChatGPT's role, teacher's role not examined); primary school subgroup has only k = 2 studies with limited statistical power; limited search timeframe; focus on basic cognitive/non-cognitive outcomes with little attention to AI literacy, computational thinking, or ethics; authors recommend meta-regression or three-level meta-analysis for future rigor. Observed: the paper never reports whether pooled post-test outcomes were measured with or without ChatGPT access, so assisted and unassisted performance are presumably pooled together; the overall estimate mixes objective test scores with self-report questionnaire outcomes (motivation, engagement, self-efficacy, interest); Begg's test is borderline (p = 0.079); one sensitivity exclusion (\"Study B\") changed a subgroup effect substantially (to g = 1.329 with Qbet = 7.884, p = 0.005), suggesting fragility in the educational-level subgroup conclusion; quasi-experimental designs are pooled with randomized ones without subgroup separation by design.","plain_summary":"This meta-analysis pooled 35 experimental and quasi-experimental studies (4193 students, 134 effect sizes) published between November 2022 and June 2024 and found a moderate overall benefit of ChatGPT on student learning outcomes (g = 0.670, 95% CI 0.495-0.844), with larger effects on cognitive outcomes (g = 0.872) than non-cognitive ones (g = 0.539). Effects were stronger for interventions longer than three months, in traditional (teacher-centred) instruction, and in some subjects (physics, chemistry, English), while education level and knowledge type did not significantly moderate results. Heterogeneity was very high (I2 = 91.4%) and the paper does not distinguish whether outcomes were measured while students still had ChatGPT available, so assisted and unassisted performance appear to be pooled together.","methodological_notes":"Extracted from: https://www.nature.com/articles/s41599-026-07019-z\nAccess: Open-access full text retrieved from Nature (Humanities and Social Sciences Communications 13:684, published 26 March 2026). Full HTML article text was accessible, including Methods, Results (overall effect, sensitivity analysis, publication bias, heterogeneity, moderator analyses), Discussion, Limitations, and funding/competing-interest statements. Tables 1-4 and Supplementary Tables S1-S5 render as links only; all statistics below are taken from the running text, which reports the key values.\nSample-size note: k = 35 studies, 134 effect sizes, 4193 pooled participants."},{"id":13,"doi":"10.1057/s41599-025-04787-y","pmid":null,"arxiv_id":null,"openalex_id":null,"nct_ids":"[]","study_group_id":null,"publication_status":"retracted","title":"The effect of ChatGPT on students' learning performance, learning perception, and higher-order thinking: insights from a meta-analysis","authors":"[\"Jin Wang\", \"Wenxiang Fan\"]","journal":"Humanities and Social Sciences Communications","publication_date":null,"year":2025,"publication_type":"journal_article","peer_reviewed":"yes","abstract":null,"url":"https://www.nature.com/articles/s41599-026-07310-z","source_name":"seed","source_tier":1,"coi_statement":null,"assessment_version":2,"assessed_by":"manual","assessed_at":"2026-09-18T19:37:38+00:00","study_design":"meta_analysis","drugs":null,"drug_details":null,"domains":null,"outcome_type":null,"primary_outcome":null,"endpoints":null,"effect_estimate":null,"confidence_interval":null,"p_value":null,"sample_size":null,"follow_up":null,"direction":null,"population":null,"applicability":null,"applicability_rationale":null,"randomization_level":null,"subject":"other","learning_task":null,"exposure_duration":null,"population_match_secondary":"not_applicable","population_match_university":"not_applicable","attrition":null,"adjusted_for":null,"evidence_level":"NOT_RATED","evidence_components":null,"evidence_rationale":"RETRACTED. The Retraction Note (22 Apr 2026, DOI 10.1057/s41599-026-07310-z) cites concerns relating to discrepancies in the meta-analysis, initially raised by Magnus Ingebrigtsen and Marko Lukic, which undermine the Editor's confidence in the validity of the analysis and its conclusions. A retracted analysis's numbers may not be used and no findings are recorded.","funding_source":null,"industry_funded":null,"manufacturer":null,"author_conflicts":null,"sponsor_role":null,"independent_replication_exists":null,"conflict_notes":null,"adverse_events":null,"limitations":"Retracted; the authors did not respond to correspondence regarding the retraction. Many secondary sources still cite the original pooled estimate.","plain_summary":"This meta-analysis of ChatGPT's effect on learning performance was, for a time, the field's most-cited pooled estimate. It was retracted in April 2026 over discrepancies in the analysis. It is kept in this collection so that readers who encounter it cited elsewhere can see its status; its results carry no evidential weight here.","methodological_notes":"Retracted-record handling per EVIDENCE-MODEL.md Rule 14."},{"id":14,"doi":"10.1007/978-3-031-64315-6_34","pmid":null,"arxiv_id":"2402.09809","openalex_id":null,"nct_ids":"[]","study_group_id":null,"publication_status":"peer_reviewed","title":"Effective and Scalable Math Support: Experimental Evidence on the Impact of an AI Math Tutor in Ghana","authors":"[\"Owen Henkel\", \"Hannah Horne-Robinson\", \"Nessie Kozhakhmetova\", \"Amanda Lee\"]","journal":"Artificial Intelligence in Education (AIED 2024)","publication_date":null,"year":2024,"publication_type":"journal_article","peer_reviewed":"yes","abstract":null,"url":"https://link.springer.com/chapter/10.1007/978-3-031-64315-6_34","source_name":"seed","source_tier":1,"coi_statement":null,"assessment_version":2,"assessed_by":"ai:two-pass","assessed_at":"2026-09-18T19:37:38+00:00","study_design":"cluster_rct","drugs":null,"drug_details":null,"domains":null,"outcome_type":null,"primary_outcome":null,"endpoints":null,"effect_estimate":null,"confidence_interval":null,"p_value":null,"sample_size":477,"follow_up":null,"direction":null,"population":"{\"level\": \"mixed\", \"ages\": \"Mean age 12.10 (SD 1.95) among students completing both tests; grades 3-8 per Participation section (grades 3-9 stated in the study overview and arXiv landing-page abstract)\", \"country\": \"Ghana\", \"prior_knowledge\": \"Baseline mean approximately 20.2/35 on the study assessment in both groups; assessment covered grade 3-5 skills, and some higher-grade students scored at ceiling at baseline\", \"selection\": \"Students from 11 Rising Academies network schools selected for similarity in geography, demographics, curricula, and teaching methodologies; participation defined by taking the baseline assessment\"}","applicability":null,"applicability_rationale":null,"randomization_level":"school","subject":"math","learning_task":"Independent practice of numeracy and algebra micro-lessons (Global Proficiency Framework curriculum) via a WhatsApp chat tutor during supervised study hall","exposure_duration":"Two 30-minute sessions per week during study hall, approximately 8 months (early February to late August 2023)","population_match_secondary":"partial","population_match_university":"different","attrition":"160 of 637 baseline students (25.1%) did not complete the endline, attributed primarily to inconsistent school attendance. Dropouts had lower baseline scores (M=18.53, SD=7.70) than completers (M=22.26, SD=7.57); no significant differences by age or gender. Authors note the growth-score analysis uses completers only.","adjusted_for":null,"evidence_level":"LOW","evidence_components":null,"evidence_rationale":"Randomization was at the school level with only 11 clusters (5 treatment, 6 control), but all analyses (independent-samples t-tests on student-level growth scores) ignore clustering, so the reported p < 0.001 is likely overstated. Attrition was 25% and differential on baseline ability, the outcome was a researcher-developed 35-item test with acknowledged ceiling effects, and no CI or preregistration is reported. The unresolved discrepancies between the arXiv landing-page abstract (~1,000 students, grades 3-9, d=0.37) and the published/v2 text (~500, grades 3-8 analyzed, d=0.36) further reduce confidence in the exact reported quantities, though the direction of effect is consistent across versions. The authors themselves label the study a preliminary evaluation of year 1.\n","funding_source":"Not stated; the paper contains no funding or acknowledgements statement.","industry_funded":"partial","manufacturer":null,"author_conflicts":null,"sponsor_role":null,"independent_replication_exists":null,"conflict_notes":"Rori is a product of Rising Academies, and the study ran inside Rising Academies' own school network. Co-author Hannah Horne-Robinson is affiliated with Rising Academies (per the author list: Henkel - University of Oxford; Horne-Robinson - Rising Academies; Kozhakhmetova and Lee - J-PAL North America). No conflict-of-interest statement is provided in either version.\n","adverse_events":null,"limitations":"Source-stated: possible Hawthorne effects at treatment schools; possible unobserved baseline differences in prior ability or socio-economic background; same assessment used for all grades produced ceiling effects (some higher-grade students scored perfectly at baseline, so their growth was unobservable and gains may be understated - or Rori may mainly help on easier topics); results cover only year 1; attrition warrants further examination; dose-response not analyzed. Observed: only 11 clusters randomized at school level with no clustering adjustment in the t-test analysis; completer-only analysis with 25% differential attrition; no CI, no preregistration, no funding/COI statement; the AI model powering Rori is not specified. Version discrepancies: arXiv landing-page abstract vs v2 PDF/Springer differ on N (~1,000 vs ~500), grade range (3-9 vs 3-8 in the analyzed sample; v2 body itself states both), effect size (0.37 vs 0.36), and session framing; v2 text vs Table 1 differ on treatment n (236 vs 237).\n","plain_summary":"Eleven schools in the Rising Academies network in Ghana were randomly split so that students in five schools spent two 30-minute study-hall sessions per week chatting with Rori, an AI math tutor on WhatsApp, for about 8 months, while six schools continued normal schooling. Among the 477 students who took both the baseline and endline of a 35-question math test, the Rori group gained about 3 points more (Cohen's d = 0.36, p < 0.001). The study is promising but preliminary: schools, not students, were randomized yet the analysis treats students as independent, a quarter of students dropped out, and the test had ceiling effects. Reported figures also differ between the arXiv landing-page abstract and the published version.\n","methodological_notes":"Extracted from: https://arxiv.org/pdf/2402.09809 (v2 full text, PDF); https://arxiv.org/abs/2402.09809 (arXiv landing-page abstract and version history); https://link.springer.com/chapter/10.1007/978-3-031-64315-6_34 (Springer AIED 2024 chapter abstract)\nAccess: Values extracted from the arXiv full text (v2 PDF, dated 5 May 2024; 10 pages, read page-by-page from the downloaded PDF). The Springer chapter abstract (AIED 2024, pp. 373-381, published 2 July 2024) was fetched for comparison and matches the v2 PDF abstract on all key figures. The arXiv LANDING-PAGE abstract still carries older (v1-era) figures that conflict with both the v2 PDF and the Springer version; version history shows v1 (15 Feb 2024) and v2 (5 May 2024). The v1 PDF itself was not fetched.\n\nSample-size note: VERSION DISCREPANCIES. (1) N: the arXiv landing-page abstract says \"approximately 1,000 students in grades 3-9 across 11 schools\", but the v2 full-text PDF and the Springer abstract both say \"approximately 500 students\"; the full text reports 637 students at baseline (336 treatment, 301 control) and 477 analyzed (text: 241 control, 236 treatment). The ~1,000 figure appears to be a v1-era metadata holdover and appears nowhere in the v2 full text. (2) Grade range: arXiv landing-page abstract says grades 3-9; the v2 full text is internally inconsistent, saying \"students in grades 3-9\" in the study overview but \"637 students in grades 3-8\" in the Participation and attrition sections; the Springer abstract states no grade range. (3) Effect size: arXiv landing-page abstract says 0.37; the v2 full text and Springer abstract both say 0.36. (4) Session description: arXiv landing-page abstract says \"two 30-minute sessions per week over 8 months\"; the v2/Springer abstracts say \"one hour a week\" with a phone during study hall (Springer adds \"low-cost smartphone\" and \"monitored\"); the v2 body says two 30-minute weekly sessions - total dose is consistent (1 hr/week) but the framing differs. (5) Internal inconsistency in v2: text says treatment n=236 but Table 1 says Treatment (N=237). The recorded sample_size of 477 is the analyzed completer sample from the v2 full text.\n"}]