AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

Cognitive ease at a cost: LLMs reduce mental effort but compromise depth in student scientific inquiry

Design
Laboratory experiment · randomized at student level · 91 participants
Subject and task
Other
20-minute individual web research on an unsettled socio-scientific issue (safety of zinc-oxide/titanium-dioxide nanoparticles in sunscreen) to advise a fictitious friend, followed by writing a recommendation with justifications from memory (no notes, webpages, or ChatGPT conversations available while writing).
Exposure: Single lab session; exactly 20 minutes of tool-assisted research, plus demographic questions, cognitive-load questionnaire, recommendation writing, and a prior-knowledge test afterward.
Population match
Secondary students: Different · University students: Direct
Attrition
One of 92 participants excluded for not following instructions (used both an LLM and a search engine); no other attrition reported — all measures were completed within the single session.
Study tier
Study tier 2Individually randomized single-session lab experiment (~20 min exposure) with an active comparison (ChatGPT-3.5 vs Google search) and a prior-knowledge covariate and randomization check. Outcomes are proximal: self-reported cognitive load about the task and a researcher-coded justification written immediately afterward (notably WITHOUT tool access while writing); there is no delayed retention, transfer, or validated learning test. Sample is small (n = 91), non-representative, and effects, while moderate-to-large and consistent, come from one site and one task, so estimates are imprecise and generalization is limited.
Assessment
Version 2 · AI: two passes agreed · Sep 18, 2026

Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.

91 German university students were randomly assigned to research the safety of nanoparticles in sunscreen for 20 minutes using either ChatGPT-3.5 or Google search, then wrote a recommendation with justifications from memory. ChatGPT users reported substantially lower intrinsic, extraneous, and germane cognitive load, but their written justifications contained significantly fewer relevant arguments than those of the Google group, and an exploratory analysis suggested the quality gap was fully mediated by reduced germane (schema-building) load. The diversity of the actual recommendations (for/neutral/against) did not differ between groups.

01Findings

What the study measured

Arms

ArmWhat it got
A · LLM condition: ChatGPT (GPT-3.5) as sole information-gathering tool
n = 44
Intervention · teacher independent
ChatGPT · GPT · Vanilla chat · chat interface · in class · used April-May 2023
B · Web search condition: Google search engine as sole information-gathering tool
n = 47
Active comparison · teacher independent
Traditional web search: Firefox homepage set to google.com; students searched and read webpages themselves for the same 20-minute task in the same lab setting.

Findings

A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.

Finding and conditionsResult
F1 · Extraneous cognitive load (ECL subscale, Klepsch et al. 2017; 1-7 agreement)
LLM condition: ChatGPT (GPT-3.5) as sole information-gathering tool vs Web search condition: Google search engine as sole information-gathering toolAI availability at assessment not reported — cannot carry a learning claimDuring interventionTransfer not assessedSelf-reportPrespecified
Benefit
Web search M = 3.96 (SD 1.55) vs LLM M = 3.00 (SD 1.67); F = 8.26, eta-sq = 0.09 (ANCOVA adjusting for prior knowledge)
not_reported
0.005
One of three subscales of a single 7-item validated instrument (Klepsch et al., 2017); omega = 0.70 in this sample. Questionnaire administered immediately after the 20-min search task and before writing, referring to load during the task. Coded benefit: lower extraneous load with the LLM is, per cognitive load theory as the paper presents it, a reduction of learning-irrelevant burden ("cognitive ease"); it did not, however, translate into better output quality (see F4).
F2 · Intrinsic cognitive load (ICL subscale, Klepsch et al. 2017; 1-7 agreement)
LLM condition: ChatGPT (GPT-3.5) as sole information-gathering tool vs Web search condition: Google search engine as sole information-gathering toolAI availability at assessment not reported — cannot carry a learning claimDuring interventionTransfer not assessedSelf-reportPrespecified
Unclear
Web search M = 4.43 (SD 1.60) vs LLM M = 3.16 (SD 1.46); F = 16.48, eta-sq = 0.15 (ANCOVA adjusting for prior knowledge)
not_reported
<0.001
Same instrument as F1; omega = 0.78. The group difference is statistically clear (LLM users perceived the task as less complex), but its learning relevance is ambiguous: the paper frames it both as a beneficial reduction in burden and as a sign the LLM task was "too simple" to require self-regulation, so direction is coded unclear despite the significant test.
F3 · Germane cognitive load (GCL subscale, Klepsch et al. 2017; 1-7 agreement)
LLM condition: ChatGPT (GPT-3.5) as sole information-gathering tool vs Web search condition: Google search engine as sole information-gathering toolAI availability at assessment not reported — cannot carry a learning claimDuring interventionTransfer not assessedSelf-reportPrespecified
Harm
Web search M = 4.79 (SD 1.18) vs LLM M = 3.14 (SD 1.53); F = 33.06, eta-sq = 0.27 (ANCOVA adjusting for prior knowledge)
not_reported
<0.001
Same instrument as F1; omega = 0.69; largest of the three load effects. Coded harm relative to learning-relevant quality: germane load indexes schema-construction effort, and the paper interprets the LLM group's lower GCL as reduced deep processing that (in the exploratory mediation, F6) fully accounts for their lower-quality justifications.
F4 · Quality of justifications: count of relevant expert-derived aspects (0-7 possible) in the written recommendation
LLM condition: ChatGPT (GPT-3.5) as sole information-gathering tool vs Web search condition: Google search engine as sole information-gathering toolMeasured without AIImmediate post-testSame contentGraded artifactPrespecified
Harm
Web search M = 1.87 (SD 1.10) vs LLM M = 1.20 (SD 0.77); F = 11.18, eta-sq = 0.11 (ANCOVA adjusting for prior knowledge)
not_reported
0.001
Researcher-developed coding scheme (7 relevant aspects) built from nanotechnology- expert consultation plus bottom-up analysis; two independent raters, kappa = 0.92, disagreements resolved by discussion; scheme not shown to participants. Importantly, the recommendation was written WITHOUT the tool: students had no notes, webpages, or ChatGPT conversations available while writing, so this measures unassisted output immediately after tool-assisted research, not tool-assisted production. LLM group produced fewer relevant arguments, hence harm. Observed scores were low overall (max 4 of 7); as an index score it was not expected to be internally consistent.
F5 · Homogeneity of recommendations (distribution of against/neutral/in-favor codings)
LLM condition: ChatGPT (GPT-3.5) as sole information-gathering tool vs Web search condition: Google search engine as sole information-gathering toolMeasured without AIImmediate post-testSame contentGraded artifactPrespecified
Null
LLM 9 against / 7 neutral / 28 in favor vs Web search 15 / 8 / 24; chi-sq(2) = 1.78
not_reported
0.411
Recommendation stance coded +1/0/-1 by two independent raters (kappa = 0.89); chi-squared test of distribution. H3 predicted less variation in the LLM group; no significant difference was found, so the hypothesized narrowing of perspectives did not appear. Coded "null" (nonsignificant), from the same tool-free written artifact as F4.
F6 · Exploratory mediation: condition -> germane cognitive load -> quality of justifications
LLM condition: ChatGPT (GPT-3.5) as sole information-gathering tool vs Web search condition: Google search engine as sole information-gathering toolMeasured without AIImmediate post-testSame contentGraded artifactNot prespecified
Harm
Significant indirect effect via GCL: beta = 0.15; direct effect nonsignificant: beta = 0.19 (p = 0.095); interpreted as full mediation
not_reported
0.020
Explicitly exploratory analysis (jAMM mediation model with prior knowledge as covariate), contingent on H2; not a primary contrast. Supports the mechanism that the LLM group's lower-quality reasoning ran through reduced germane load. Cross- sectional single-session mediation, so causal ordering of load and quality is assumed, not demonstrated.
Limitations
Authors note: no think-aloud protocol or search-log analysis, so cognitive/ metacognitive processes and strategies are not observed; prompting skill and prior LLM experience unexamined as moderators; fixed 20-minute window may have been too long for the LLM group; small, non-representative, university-only sample with presumably high digital literacy; artificial constraint to a single tool (real learning mixes sources, and search engines now embed LLMs); knowledge test given after the task, so scores could be inflated by the search itself.

Who was studied

Level
university
Ages
M = 22.3 years (SD = 4.11)
Country
Germany
Prior knowledge
Some prior nanotechnology knowledge (M = 2.96, SD = 1.52, of 8 on an adapted Public Knowledge of Nanotechnology test); medicine, pharmacy, and biology students excluded a priori because of potential topic knowledge; groups did not differ on prior knowledge (t(89) = -0.28, p = 0.777, d = -0.06).
Selection
Convenience sample of students from various academic programs at "a prestigious German university" (April-May 2023); participation credits offered as compensation for some programs (e.g., psychology majors and minors); 67 female, 24 male in the analyzed sample.

Methodological notes

Extracted from: https://opus.bibliothek.uni-augsburg.de/opus4/files/114673/114673.pdf Access: Direct fetch attempts (curl, WebFetch) from this session timed out / were refused, but the completed PDF download from this exact URL was present in the session scratchpad (stadler2024.pdf, 1,315,728 bytes). Verified the PDF header, title, author list, and DOI (10.1016/j.chb.2024.108386) against the assignment, then produced an independent text extraction with pdftotext (stadler2024-passb.txt) and worked only from that extraction. No repository files, other passes, or press coverage were read. Full paper text available (CC BY 4.0 Augsburg OPUS copy of the CHB open-access article). Sample-size note: 92 recruited; 1 excluded for using both tools; analyzed n = 91 (web search n = 47, LLM n = 44).

01bRisk of bias

How much each result can be relied on

Risk of bias is judged per result, not per paper: the same study can report one outcome at low risk and another at high. Two reviewers assess each result independently against a written guide, and every domain judgment below carries its reasoning and the sentences it rests on.

Risk of bias is not the same as whether the numbers work. A trial can be well conducted and still print a standard deviation that is arithmetically impossible, or a table whose sign contradicts its own text. Risk of bias has no question for that, so where it was found it is recorded separately below, under Source integrity.

These words are not quality scores. On the RoB 2 scale, High risk of bias is the worst rating a result can receive and Low risk of bias is the best — the opposite of how the same words read on some other scales. Levels are always written out in full here for that reason.

QUALITY

Some concerns about risk of bias · RoB 2 · Two reviewers agreed on every domain

How the overall was reached: Worst-domain rule: D1 (randomization procedure and concealment not described), D4 (rater blinding to condition not reported for a researcher-coded outcome requiring judgment), and D5 (no preregistration or analysis plan) are each Some concerns; D2 and D3 are Low. No domain reaches High: the exclusion was a single non-adherent participant, outcome data are essentially complete, coding reliability is high, and the reported analysis matches the stated hypotheses. Overall risk of bias for the quality-of-justifications ANCOVA result is Some concerns.

DomainJudgment and reasoning
D1 · Randomisation process
Some concerns about risk of bias
Random assignment to the two tool conditions is explicitly stated, and a randomization check found no group difference in prior knowledge (t(89) = -0.28, p = 0.777, d = -0.06), with no other baseline imbalance suggesting a problem. However, the paper gives no information on how the allocation sequence was generated or whether allocation was concealed until enrolment. With randomization stated, no information on concealment, and no imbalance signal, the RoB 2 algorithm yields Some concerns; concealment risk is plausibly minor in a single-session lab study, but it cannot be verified.
Students were randomly assigned to one of two groups using different tools of information search.
First, we found in the randomization check that the two groups did not differ significantly in their level of prior knowledge (t(89) = −0.28; p = 0.777, d = −0.06).
D2 · Deviations from intended interventions
Low risk of bias
Participants and researchers were necessarily aware of the assigned tool (conditions cannot be blinded), but the sessions were tightly standardized — computers preconfigured to Google or ChatGPT-3.5, identical 20-minute task in a supervised lab — and no deviations arising from the trial context are reported. Adherence was otherwise complete: the single deviation was one participant who used both an LLM and a search engine and was excluded, so the analysis is a modified intention-to-treat on 91 of 92 randomized participants. Excluding 1 of 92 (~1%), a participant who effectively received both conditions and fits neither arm, has no potential to substantially impact the estimated effect of assignment.
One participant did not follow the instructions, using both an LLM and a search engine for the research, and was therefore excluded from the study.
Thus, the final sample size was 91 students, with 67 female and 24 male participants.
D3 · Missing outcome data
Low risk of bias
The written recommendation was produced by every retained participant at the end of the single lab session, so outcome data on justification quality were available for all 91 analyzed participants (Table 3 recommendations also sum to 91: 44 LLM + 47 web search). Aside from the one pre-analysis exclusion handled under D2, no missing outcome data are reported and there is no plausible mechanism for missingness related to the true value in a supervised single-session design.
Thus, the final sample size was 91 students, with 67 female and 24 male participants.
Students were informed that they had exactly 20 min to research this issue after which they would be asked to provide a written recommendation with justifications without any notes (web-pages or conversations with ChatGPT).
D4 · Measurement of the outcome
Some concerns about risk of bias
The measure itself is defensible: an appropriate coding scheme (developed with nanotechnology experts plus bottom-up analysis), applied identically to both groups, with two independent raters and high inter-rater agreement (k = 0.92, disagreements resolved by discussion), and recommendations were written without notes or ChatGPT transcripts, reducing direct condition cues in the coded texts. However, the paper does not state that raters were blinded to condition. Counting "relevant aspects mentioned" involves rater judgment, and if coders knew (or could infer from stylistic cues) which condition produced a text, assessment could plausibly have been influenced, in either direction. No information on assessor blinding yields Some concerns.
Two raters independently coded all justifications, achieving an overall inter-rater agreement of k = 0.92. The disagreements were resolved through discussion.
We used the number of relevant aspects each participant mentioned in their recommendation as a dependent measure.
D5 · Selection of the reported result
Some concerns about risk of bias
Hypotheses (H2) and the ANCOVA adjusting for prior knowledge are clearly stated in the paper, a single quality outcome and a single analysis are reported for this result, and the mediation analysis is explicitly labeled exploratory. However, there is no preregistration or registered analysis plan; the OSF repository (osf.io/jpxyt) is described as hosting instructions, scales, items, and data, not a prospective protocol. Without a pre-specified plan predating unblinded data, selection from multiple possible outcome codings or analyses cannot be ruled out, though there is no positive indication of selective reporting.
To test the Hypotheses based on RQ1 and RQ2, we conducted analyses of covariance (ANCOVA) comparing the mean scores of the two groups controlling for prior knowledge.
Data is available on osf
02Funding

Funding and conflicts

Funding
not_reported — the paper contains no funding or acknowledgements statement (only CRediT authorship, declaration of competing interest, and an OSF data-availability note; materials at https://osf.io/jpxyt).
Vendor funded
Unclear
Vendor
None identified
Notes
Authors declare no known competing financial interests or personal relationships; no funding statement is given, so funding source cannot be determined from the paper.

Funding is shown on every study and never used to score it.

03Claims

Claims this study bears on

04Source

The source, as retrieved

Abstract

No abstract retrieved.

Where this record came from

SourceRetrievedIdentifier
seedSep 18, 2026link
first ingestion