AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting

Design
Crossover trial · randomized at small group level · 194 participants
Subject and task
Physics
introductory college physics lessons (surface tension; fluid flow) taught via activity worksheets with problem-solving
Exposure: two lessons in consecutive weeks (each student experienced one lesson per condition); in-class lesson 60 minutes of learning time, AI-tutored lesson self-paced at home (median 49 minutes)
Population match
Secondary students: Different · University students: Direct
Attrition
233 enrolled, 194 (83%) eligible and included; exclusions were for lack of consent, non-participation in either condition, or incomplete pre-/post-tests; no differential-attrition analysis reported
Study tier
Study tier 2Randomized crossover trial in an authentic course with within-student comparison, an active-learning (not passive) comparator, controls for prior knowledge, topic, test version, and time on task, and a large, highly significant effect (p < 10^-8; adjusted d = 0.63, quantile-regression estimate 0.73-1.3 SD). However, the AI arm bundles elements the comparison lacks: self-paced at-home study, professionally produced pre-recorded instructor videos, and pre-written step-by-step solutions delivered through a bespoke scaffolded platform, so the contrast is condition-bundle vs. classroom rather than AI per se. Outcomes are immediate researcher-developed post-tests over only two lessons at a single elite institution, with ceiling effects acknowledged, randomization in 2-3 student clusters, and no delayed retention measure; no preregistration is mentioned.
Assessment
Version 2 · AI: two passes agreed · Sep 18, 2026

Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.

Harvard researchers had 194 students in a large introductory physics course each experience one lesson taught in class with well-implemented active learning and one lesson taught at home by 'PS2 Pal,' a custom GPT-4 tutor engineered with pedagogical prompts, scaffolded question sequences, and pre-written step-by-step solutions. Students scored substantially higher on immediate post-tests after the AI-tutored lessons (median 4.5 vs 3.5; adjusted effect size 0.63, ceiling-corrected 0.73-1.3 SD) while spending less time (median 49 vs 60 minutes), and reported feeling more engaged and motivated, with enjoyment and growth mindset comparable. The AI condition bundled self-pacing, videos, and vetted solutions with the AI itself, and outcomes were immediate tests over just two lessons, so the result shows what a carefully engineered AI-tutoring package can do rather than what generic chatbot use does.

01Findings

What the study measured

Arms

ArmWhat it got
A · AI-tutored lesson at home (PS2 Pal)
n = 194
Intervention · teacher independent
PS2 Pal (custom AI tutor platform built by the authors) · GPT-4 (OpenAI) · Guardrailed tutor · tutor scaffold interface · homework · used Fall 2023 semester (study run in weeks 9-10 of the course)
B · in-class active learning lesson
n = 194
Active comparison · teacher led
75-minute in-class session (60 minutes learning after tests) using research-based active learning: instructor introduces each activity, students work through the identical worksheet in self-selected peer-instruction groups of 2-3 with support from course staff, and the instructor gives targeted feedback addressing questions and difficulties

Findings

A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.

Finding and conditionsResult
F1 · post-test content-mastery score (pre/post quizzes written by a separate team member from learning goals; both topics/weeks combined)
AI-tutored lesson at home (PS2 Pal) vs in-class active learning lessonAI availability at assessment not reported — cannot carry a learning claimImmediate post-testSame contentResearcher-developed testPrespecification unclear
Benefit
median post-score 4.5 (AI, N = 142) vs 3.5 (in-class, N = 174) against pooled pre-test baseline median 2.75 (N = 316); median learning gains in the AI group over double those in-class
not_reported
p < 10^-8 (Mann-Whitney rank-sum, z = -5.6)
tests constructed by a team member independent of lesson/AI design, based on learning goals not lesson content; authors report a ceiling effect that makes gains an underestimate; per-week trends matched the combined result (footnote)
F2 · post-test score, regression-adjusted (controls: pre-test, midterm, FCI, ChatGPT experience, topic, test version, time on task; clustered at student level)
AI-tutored lesson at home (PS2 Pal) vs in-class active learning lessonAI availability at assessment not reported — cannot carry a learning claimImmediate post-testSame contentResearcher-developed testPrespecification unclear
Benefit
linear-regression effect size 0.63 SD (authors call this an underestimate due to ceiling effect); quantile regression estimate 0.73 to 1.3 SD; robust to clustering at peer-group level (p < 0.001)
not_reported (range 0.73-1.3 SD given for quantile-regression estimate)
p < 10^-8
regression table is in supplement (Table S1), not retrieved; ceiling effect acknowledged by authors
F3 · self-reported engagement (5-point Likert agreement)
AI-tutored lesson at home (PS2 Pal) vs in-class active learning lessonAI availability at assessment not reported — cannot carry a learning claimImmediate post-testTransfer not assessedSelf-reportPrespecification unclear
Benefit
mean 4.1 (SD 0.98) AI vs 3.6 (SD 0.92) in-class
not_reported
p < 0.0001 (dependent t-test, t(311) = -4.5)
single Likert item, unblinded self-report collected after each lesson
F4 · self-reported motivation (5-point Likert agreement)
AI-tutored lesson at home (PS2 Pal) vs in-class active learning lessonAI availability at assessment not reported — cannot carry a learning claimImmediate post-testTransfer not assessedSelf-reportPrespecification unclear
Benefit
mean 3.4 (SD 1.0) AI vs 3.1 (SD 0.86) in-class
not_reported
p < 0.001 (dependent t-test, t(311) = -3.4)
single Likert item, unblinded self-report
F5 · self-reported enjoyment (5-point Likert agreement)
AI-tutored lesson at home (PS2 Pal) vs in-class active learning lessonAI availability at assessment not reported — cannot carry a learning claimImmediate post-testTransfer not assessedSelf-reportPrespecification unclear
Null
not statistically significantly different between groups (means not reported in text; shown in Fig. 3)
not_reported
not_reported (nonsignificant)
single Likert item, unblinded self-report
F6 · self-reported growth mindset (5-point Likert agreement)
AI-tutored lesson at home (PS2 Pal) vs in-class active learning lessonAI availability at assessment not reported — cannot carry a learning claimImmediate post-testTransfer not assessedSelf-reportPrespecification unclear
Null
not statistically significantly different between groups (means not reported in text; shown in Fig. 3)
not_reported
not_reported (nonsignificant)
single Likert item, unblinded self-report
F7 · time on task (platform-tracked for AI group; assumed 60 minutes for in-class group)
AI-tutored lesson at home (PS2 Pal) vs in-class active learning lessonAI availability at assessment not reported — cannot carry a learning claimDuring interventionTransfer not assessedProcess / behaviorPrespecification unclear
Benefit
median 49 minutes (AI) vs 60 minutes assumed (in-class); 70% of AI-group students spent under 60 minutes; no correlation between time on task and post-test score
not_reported
not_reported (no statistical test of the time difference reported)
in-class time is an assumption (75-minute period minus 15 minutes of testing), not a measurement; AI time is logged platform interaction time
Limitations
Only two lessons; post-tests immediate, with no delayed/retention measure; ceiling effect on post-test acknowledged by the authors; researcher-developed tests; randomization clustered in 2-3 student peer groups; AI condition differs from control on medium, location (home vs. class), pacing, videos, and pre-written solutions simultaneously; self-report outcomes unblinded; authors state they 'do not presume that structured AI tutoring will always outperform in-class active learning in all contexts, for example, those requiring complex synthesis of multiple concepts and higher-order critical thinking'; whether the AI tutor was accessible during post-tests is never stated; single course at a single institution; per-condition post-test Ns (142 vs 174) unexplained.

Who was studied

Level
university
Ages
not_reported
Country
USA
Prior knowledge
students in an introductory physics course for the life sciences; FCI pretest scores comparable to students at other universities; over 90% reported they had not studied the two lesson topics in depth before the course
Selection
consenting students enrolled in Harvard's Physical Sciences 2 course (Fall 2023) who participated in both conditions and completed all pre- and post-tests (194 of 233 enrolled)

Methodological notes

Extracted from: https://www.nature.com/articles/s41598-025-97652-6 Access: full text (HTML of open-access article, including Abstract, Introduction, Results, Discussion, Methods, Notes, Acknowledgements, author information, and competing-interests declaration; supplementary tables S1-S3 and Supplementary Material 1 not retrieved; no funding statement present in retrieved text) Sample-size note: 194 analyzed students (of 233 enrolled) in a crossover design; reported post-test observation counts are N = 142 (AI condition) and N = 174 (in-class condition), pre-test baseline N = 316 pooled; the paper does not explain why per-condition post-test Ns are below 194 each

01bRisk of bias

How much each result can be relied on

Risk of bias is judged per result, not per paper: the same study can report one outcome at low risk and another at high. Two reviewers assess each result independently against a written guide, and every domain judgment below carries its reasoning and the sentences it rests on.

Risk of bias is not the same as whether the numbers work. A trial can be well conducted and still print a standard deviation that is arithmetically impossible, or a table whose sign contradicts its own text. Risk of bias has no question for that, so where it was found it is recorded separately below, under Source integrity.

These words are not quality scores. On the RoB 2 scale, High risk of bias is the worst rating a result can receive and Low risk of bias is the best — the opposite of how the same words read on some other scales. Levels are always written out in full here for that reason.

POSTTEST

High risk of bias · RoB 2 · Two reviewers agreed on every domain

How the overall was reached: Worst-domain rule: D3 is High because roughly 18% of expected post-test observations are absent from the primary analysis (142 + 174 = 316 of an expected 388 from 194 crossover participants), the shortfall is differential across conditions and entirely unexplained, and completion was voluntary in an unsupervised at-home arm, making outcome-dependent missingness plausible with no sensitivity analysis to rule out bias. D1, D2, D4, and D5 each carry Some concerns (unreported sequence generation/concealment with small-cluster randomization; unblinded, unproctored at-home delivery and non-ITT eligibility filtering; arm-dependent, possibly unproctored test administration with a ceiling effect; no preregistered analysis plan), which compound the overall risk.

DomainJudgment and reasoning
D1 · Randomisation process
Some concerns about risk of bias
Random assignment is stated and randomization was at the level of standing 2-3-student peer-instruction groups (small clusters), but the paper gives no information on how the allocation sequence was generated or whether it was concealed. Reported baseline comparability is reassuring: demographics and prior physics measures (FCI, midterm) were comparable across groups (Tables S2A/B), and a sensitivity regression clustered at the peer-group level yielded similar results, mitigating the small-cluster concern. Carryover risk is structurally low: the two lesson topics (surface tension week 1, fluid flow week 2) were distinct and described as independent of each other, and the adjusted model controlled for topic and test version; however, no formal test of period or sequence effects is reported, and the combined-weeks presentation relies on a footnote asserting the same trend in each week. Absent sequence-generation and concealment detail, Some concerns is warranted.
Students were randomly assigned to two groups, respecting the constraint that students who regularly worked together in class during peer instruction were placed in the same group in order to maximize the effectiveness of their in-class learning.
The demographics of the two groups were comparable (see table S2A), as were previous measures of their physics background knowledge (see Table S2B).
D2 · Deviations from intended interventions
Some concerns about risk of bias
Participants and instructors were necessarily aware of assigned condition (no blinding possible). The experimental condition was delivered at home, unproctored, and self-paced, so condition-specific deviations from the intended intervention (use of outside resources, collaboration with peers, splitting the lesson across sittings) were possible and uncheckable, whereas the in-class condition was supervised; time on task in fact varied widely in the AI arm. The trial context (authentic course, equal participation credit in both conditions, honest-effort instruction) reduces but does not eliminate this risk. In addition, the analysis is not intention-to-treat with respect to randomization: eligibility for analysis was conditioned post hoc on participating in both conditions and completing all tests, so non-adherent randomized students were excluded rather than analyzed as assigned. Period/order effects from the fixed condition-order-by-group structure were addressed only via topic and test-version covariates in the adjusted model, not in the unadjusted Mann-Whitney.
During, the first week, group 1 engaged with an AI-supported lesson at home while group 2 participated in an active learning lesson in class. The conditions were reversed the following week.
Eligibility was based on students’ consent, participation in both in-class and AI-tutored instruction, and completion of all pre-tests and post-tests.
D3 · Missing outcome data
High risk of bias
Of 233 enrolled students, only 194 were deemed eligible for analysis (about 83%), with exclusions driven by consent, participation, and test completion. More importantly, the primary Mann-Whitney comparison rests on 142 AI-condition and 174 in-class post-test observations (sum 316), whereas 194 students each experiencing both conditions should yield 388 observations; roughly 18% of expected post-tests are missing, the shortfall is markedly asymmetric across conditions (142 vs 174), and the paper offers no explanation, flow diagram, or sensitivity analysis for this. Completion was voluntary and effort-based, so missingness could plausibly depend on how much a student learned (e.g., students who struggled with the unsupervised at-home lesson skipping its post-test), and the larger shortfall in the AI arm is in the direction that could inflate that arm's scores. With unexplained, differential, plausibly outcome-related missingness and no analysis correcting for it, risk is High.
Of the 233 enrolled students, 194 were eligible for inclusion in the study.
Students in the AI group exhibited a higher median (M) post-score (M = 4.5, N = 142) compared to those in the in-class active learning group (M = 3.5, N = 174).
D4 · Measurement of the outcome
Some concerns about risk of bias
The outcome is a researcher-developed, non-validated test, though sensible safeguards are described: items were written by a team member separate from those designing the AI tutor or teaching, and were based on learning goals rather than lesson content. The salient problem is that measurement conditions differed systematically by arm: the in-class group took the post-test inside the proctored 75-minute class period, while the AI group's lesson (and therefore, by implication, its post-test) occurred at home; the paper never states where the AI-condition post-test was taken, whether it was proctored, or whether the AI tutor or other resources were accessible during it. Because the test was self-administered by unblinded participants under arm-dependent conditions, score inflation in the unsupervised arm cannot be excluded. A ceiling effect in post-test scores is acknowledged, which compresses the scale (biasing the estimate toward the null rather than away from it) but further limits measurement quality. No information supports an affirmative finding that measurement did differ in effect, so Some concerns rather than High.
To prevent the specific test questions from influencing the teaching or AI tutor design, the tests were constructed by a separate team member from those involved in designing the AI or teaching the lessons.
During a 75-minute period, the in-class students spent 15 minutes taking the pre- and post-tests; we assume 60 minutes spent on learning.
D5 · Selection of the reported result
Some concerns about risk of bias
No preregistration, trial registration, or pre-specified statistical analysis plan is mentioned anywhere in the article, so there is no way to verify that the reported analyses (unadjusted Mann-Whitney, linear regression with a particular covariate set, quantile regression for effect size) were selected in advance rather than after seeing the data. Multiple analytic routes and multiple candidate outcome summaries (median, mean, gain scores) are discussed, with the authors explicitly preferring the measures that handle the ceiling effect; all reported analyses point in the same direction, which limits the practical concern, but the absence of a pre-specified plan means Some concerns under RoB 2.
We conducted a two-sample rank-sum (Mann–Whitney) test to compare the distribution of post-scores of the two groups.
Note that measures that are less sensitive to ceiling effect, such as the median, will be more reliable than measures that are more sensitive to ceiling effect, such as straight gain or mean.
02Funding

Funding and conflicts

Funding
not_reported (no funding statement in the retrieved full text)
Vendor funded
Unclear
Vendor
None identified
Notes
The paper states 'The authors declare no competing interests.' The AI tutor platform was conceived, designed, and engineered by author G.K. (in-house academic tool; no AI-vendor ties stated). Authors also note they overlap with the authors of the prior literature validating the in-class active-learning approach used as the comparator.

Funding is shown on every study and never used to score it.

04Source

The source, as retrieved

Abstract

No abstract retrieved.

Where this record came from

SourceRetrievedIdentifier
seedSep 18, 2026link
first ingestion