AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

ChatGPT as a Learning Tool for Medical Students: Results From a Randomized Controlled Trial

Design
Randomized trial · randomized at student level · 33 participants
Subject and task
Other
answering a 10-item multiple-choice quiz on first-year medical curriculum content while consulting an assigned study resource
Exposure: single 15-minute proctored quiz session in Week 1 (resource available only during the quiz; no separate study period)
Population match
Secondary students: Different · University students: Partial
Attrition
not_reported
Study tier
Study tier 3True individually randomized trial with two comparison arms, but very small (N=33, 10-12 per arm), single-institution, and unblinded for both participants and researchers. The significant Week 1 result is an open-book contrast (each arm had its assigned resource available during the quiz), so it measures resource-assisted performance rather than learning; the learning-relevant closed-book retention contrast was null and, by the authors' own post hoc calculation, would have needed 72-132 participants for 80% power. No trial registration is mentioned, and precision at this sample size is too low to support any retention conclusion.
Assessment
Version 2 · AI: two passes agreed · Sep 18, 2026

Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.

Thirty-three first-year medical students at Georgetown were randomized to take a 10-question quiz while using ChatGPT-4.0, non-AI internet resources, or their course's own materials. With resources in hand, the ChatGPT and internet groups scored much higher than the institutional-materials group, and ChatGPT was no better than ordinary internet resources. One week later, on the same quiz with no resources allowed, all three groups' scores dropped and no longer differed significantly, so the study found no retention advantage for ChatGPT - though it was far too small to detect modest retention differences.

01Findings

What the study measured

Arms

ArmWhat it got
A · ChatGPT-4.0 during quiz
n = 10
Intervention · teacher independent
ChatGPT · GPT-4 · Vanilla chat · chat interface · in class · used April 2025 (two-week study period)
B · External non-AI online resources during quiz
n = 12
Active comparison · teacher independent
Publicly available online materials (e.g., Google, PubMed, third-party educational websites), excluding AI-assisted tools, available during the Week 1 quiz
C · Internal institutional resources during quiz
n = 11
Control · teacher independent
Internal institutional resources such as lecture materials, electronic textbooks, and course-provided slides, available during the Week 1 quiz

Findings

A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.

Finding and conditionsResult
F2 · Week 1 quiz score, Tukey HSD pairwise comparison
ChatGPT-4.0 during quiz vs Internal institutional resources during quizWith AI at assessment — assisted performance, not learning evidenceDuring interventionSame contentResearcher-developed testPrespecification unclear
Benefit
9.60 +/- 0.52 vs 6.64 +/- 1.57
not_reported
p < 0.001
Same open-resource quiz; benefit reflects ChatGPT access during the assessment
F3 · Week 1 quiz score, Tukey HSD pairwise comparison
ChatGPT-4.0 during quiz vs External non-AI online resources during quizWith AI at assessment — assisted performance, not learning evidenceDuring interventionSame contentResearcher-developed testPrespecification unclear
Null
9.60 +/- 0.52 vs 9.08 +/- 0.79
not_reported
p = 0.50
ChatGPT conferred no significant advantage over ordinary internet resources even with both available during the quiz; near-ceiling scores in both arms limit sensitivity
F4 · Week 1 quiz score, Tukey HSD pairwise comparison
External non-AI online resources during quiz vs Internal institutional resources during quizMeasured without AIDuring interventionSame contentResearcher-developed testPrespecification unclear
Benefit
9.08 +/- 0.79 vs 6.64 +/- 1.57
not_reported
p < 0.001
Non-AI contrast; shows the Week 1 gap is about open-web lookup vs course materials, not AI per se
F5 · Week 2 retention quiz score (same 10 MCQs, no resources allowed), one-way ANOVA
ChatGPT-4.0 during quiz vs Measured without AIDelayed post-test · 1 week after the Week 1 quizSame contentResearcher-developed testPrespecified
Null
A 6.20 +/- 1.93 vs B 5.58 +/- 2.07 vs C 4.36 +/- 2.01
not_reported
p = 0.118
Closed-book readministration of the identical quiz; the learning-relevant outcome, but authors' post hoc power calculation says 72-132 participants were needed for 80% power, so the null is uninformative about modest effects
F6 · Week 1 quiz completion time (minutes)
ChatGPT-4.0 during quiz vs AI partly available at assessment — cannot carry a learning claimDuring interventionTransfer not assessedProcess / behaviorPrespecification unclear
Null
A 8.92 min vs B 8.83 min vs C 10.04 min
not_reported
p = 0.505
Objective timing from electronic quiz delivery; no efficiency advantage for ChatGPT
F7 · Week 2 quiz completion time (minutes)
ChatGPT-4.0 during quiz vs Measured without AIDelayed post-test · 1 week after the Week 1 quizTransfer not assessedProcess / behaviorPrespecification unclear
Null
A 2.34 min (SD 1.07) vs B 2.38 min (SD 1.62) vs C 2.03 min
not_reported
p = 0.771
Objective timing; all groups much faster on retest
F8 · Post-quiz Likert survey perceptions (usefulness, satisfaction, perceived efficiency, confidence)
ChatGPT-4.0 during quiz vs AI availability at assessment not reported — cannot carry a learning claimImmediate post-testTransfer not assessedSelf-reportPrespecification unclear
Unclear
Group A perceived ChatGPT as enhancing quiz efficiency and reported high satisfaction, but mixed confidence in AI as a primary resource; percentages relegated to appendices
not_reported
not_reported
Unvalidated researcher-designed Likert survey; self-report subject to response and social desirability bias (author-acknowledged); perceived efficiency contradicts the null objective timing result
Limitations
Source-stated: retention comparisons underpowered (N=33 vs 72-132 needed for 80% power); single institution and first-year students only; no baseline academic indicators collected; no blinding of participants or researchers (performance, expectancy, and novelty biases possible); survey outcomes self-reported and subject to response/social desirability bias; no long-term follow-up. Observed: the primary outcome was taken with resources in hand, so the Week 1 "benefit" conflates access to an answer-finding tool with learning; the same 10-item quiz was reused in Week 2 (retest exposure); no trial registration reported; ceiling effects likely in Groups A and B on the 10-point Week 1 quiz (means 9.60 and 9.08).

Who was studied

Level
university
Ages
not_reported
Country
USA (Washington, DC)
Prior knowledge
first-year MD students in good academic standing; quiz covered material students were expected to have mastered at that point in the semester
Selection
volunteers recruited via email and campus bulletin board flyers in April 2025; inclusion: first-year MD enrollment, good academic standing, written informed consent; no exclusion criteria

Methodological notes

Extracted from: https://pmc.ncbi.nlm.nih.gov/articles/PMC12248138/ Access: full text (PMC open-access HTML, including methods, results, discussion, and disclosure statements; appendix survey tables referenced but percentages not extracted) Sample-size note: Group A (ChatGPT-4.0) n=10; Group B (external non-AI online resources) n=12; Group C (institutional resources) n=11

01bRisk of bias

How much each result can be relied on

Risk of bias is judged per result, not per paper: the same study can report one outcome at low risk and another at high. Two reviewers assess each result independently against a written guide, and every domain judgment below carries its reasoning and the sentences it rests on.

Risk of bias is not the same as whether the numbers work. A trial can be well conducted and still print a standard deviation that is arithmetically impossible, or a table whose sign contradicts its own text. Risk of bias has no question for that, so where it was found it is recorded separately below, under Source integrity.

These words are not quality scores. On the RoB 2 scale, High risk of bias is the worst rating a result can receive and Low risk of bias is the best — the opposite of how the same words read on some other scales. Levels are always written out in full here for that reason.

Assessed, not settled (1)

Recorded rather than omitted. Leaving these out would make the appraisal look cleaner than it is, and a disagreement between two careful readers is itself worth knowing.

RETENTION

Assessed, not settled · RoB 2 · Two reviewers reached different judgments and nothing has settled it

This result has no overall judgment, because the two independent reviews did not reach one. Both readings are shown. The domains they did agree on are below.

Where they differThe two readings
D4 · Measurement of the outcomeOne reviewer: Some concerns about risk of bias
The other: Low risk of bias
DomainJudgment and reasoning
D1 · Randomisation process
Some concerns about risk of bias
A computer-generated random sequence is reported (pseudo-random number generator), and the unequal group sizes (10/12/11) are consistent with simple randomization of N=33. However, nothing is reported about allocation concealment (who generated and applied the sequence, or whether assignment could be foreseen), and no baseline characteristics table is provided against which comparability could be checked - the authors state baseline academic indicators were not collected. With adequate sequence generation but no information on concealment and no baseline data to assess imbalance, Some concerns.
Randomization was performed using a pseudo-random number generator.
Baseline academic indicators (e.g., college GPA) were not collected, so we could not assess whether ChatGPT's impact varied by prior academic achievement.
D2 · Deviations from intended interventions
Some concerns about risk of bias
Participants and researchers were both unblinded, and the authors themselves acknowledge the resulting risk of performance and expectancy biases. For the retention outcome specifically, the two-week interval between quizzes was uncontrolled: nothing is reported about whether participants were told not to study, whether they knew the identical quiz would be readministered, or whether groups differed in between-week study behavior (e.g., continued ChatGPT use, or compensatory studying by the lower-scoring institutional-resources group). No deviations are actually reported, and the analysis appears to include all randomized participants in their assigned groups (identical Ns at both weeks), so there is no evidence of inappropriate analysis - but the absence of information on trial-context deviations between weeks precludes Low risk.
Participants and researchers were not blinded to group assignments.
Furthermore, the absence of blinding presents opportunities for performance and expectancy biases.
D3 · Missing outcome data
Low risk of bias
The article contains no explicit attrition or flow statement, but the group Ns reported for the Week-2 retention analysis (Group A N=10, Group B N=12, Group C N=11) are identical to the Ns at enrollment and at Week 1, indicating outcome data were available for all 33 randomized participants at both timepoints. With apparently complete outcome data, Low risk.
Participants (N = 33) were assigned to one of three groups: Group A (ChatGPT-4.0; N = 10, 30.3%), Group B (external resources; N = 12, 36.4%), and Group C (institutional resources; N = 11, 33.3%).
Delayed mean repeat quiz scores revealed a nonsignificant trend toward improved retention: Group A (N = 10, mean = 6.20 ± 1.93), Group B (N = 12, mean = 5.58 ± 2.07), and Group C (N = 11, mean = 4.36 ± 2.01) (p = 0.118)
D4 · Measurement of the outcome
Some concerns about risk of bias
The measure itself is relatively objective (10 faculty-derived MCQs delivered electronically via Qualtrics, proctored and closed-book, identical items and timing for all groups), so scoring bias is unlikely. However, the retention outcome reused the identical 10 items from Week 1, so scores incorporate retest/practice effects, and the degree of Week-1 item exposure differed systematically by group (Week-1 scores ranged from 6.64 to 9.60 across groups); the article does not state whether answer feedback was withheld. Participants - who effectively generate the outcome - were also aware of their assignment, and the authors note expectancy bias as a possibility. The measurement method was the same for all groups, but assessment of "retention" via an identical readministered quiz by unblinded participants could plausibly be influenced differentially, so Some concerns rather than Low.
In Week 1, each group completed an identical 15-minute proctored quiz using only their assigned resources.
In Week 2, the same quiz was readministered without any resource access to assess short-term knowledge retention.
D5 · Selection of the reported result
Some concerns about risk of bias
No trial registration or pre-specified protocol/statistical analysis plan is mentioned anywhere in the article, so reported results cannot be compared against pre-specified intentions. Mitigating this, the text designates the retention score as the stated secondary outcome, the analysis is a single conventional one (one-way ANOVA across the three groups), and the nonsignificant result is reported in full with means, SDs, and p-value - there is no indication of selection among multiple measurements or analyses. Absent registration, Some concerns is the ceiling.
The primary outcome was the initial quiz score. The secondary outcome was the retention score, evaluated through the second quiz without resources.
A one‑way analysis of variance (ANOVA) compared continuous outcomes across the three study groups, and Tukey's honestly significant difference (HSD) test provided pairwise comparisons for initial quiz scores (Week 1).
02Funding

Funding and conflicts

Funding
None stated; authors declared no financial support was received from any organization for the submitted work
Vendor funded
No
Vendor
None identified
Notes
Authors declared no financial relationships within the previous three years with organizations that might have an interest in the work, and no other relationships or activities that could appear to have influenced it

Funding is shown on every study and never used to score it.

03Claims

Claims this study bears on

04Source

The source, as retrieved

Abstract

No abstract retrieved.

Where this record came from

SourceRetrievedIdentifier
seedSep 18, 2026link
first ingestion