Changes
Every change in the evidence, typed and dated: new results, status changes, corrections, retractions. Each entry says what the new evidence is and whether any claim status changed because of it — or why not.
Also available as an RSS feed. Claim pages carry their own full status history.
- Claim assessed: Preliminary SignalClaim status changed · Sep 18, 2026 · Notable
"Secondary students with low prior knowledge gain less, or are harmed more, by AI assistance during practice than high-prior-knowledge peers." received its first assessed status: PRELIMINARY_SIGNAL (unclear). Two independent weak analyses point the claim's way (retention gains confined to high-prior learners; a positive treatment-by-baseline interaction), while the strongest trial's preregistered heterogeneity analysis — gated for stance because its assessment conditions are not reported — found no detectable ability moderation. An early, internally inconsistent signal.
- Claim assessed: Insufficient EvidenceClaim status changed · Sep 18, 2026 · Notable
"Learning gains from AI-assisted practice persist on delayed (one week or longer) retention tests taken without AI." received its first assessed status: INSUFFICIENT_EVIDENCE (stable). One underpowered null (N=33, ~43% power). The claim is effectively untested at meaningful precision.
- Claim assessed: Insufficient EvidenceClaim status changed · Sep 18, 2026 · Notable
"Learning gains from AI-assisted practice persist on delayed (one week or longer) retention tests taken without AI." received its first assessed status: INSUFFICIENT_EVIDENCE (stable). One study measured delayed unassisted retention: three concordant nulls with favorable trends at N=69, plus an unprespecified positive high-prior subgroup. Neither persistence nor decay is demonstrated.
- Claim assessed: Insufficient EvidenceClaim status changed · Sep 18, 2026 · Notable
"AI feedback on writing produces unassisted writing growth noninferior to human feedback — no worse than a stated margin." received its first assessed status: INSUFFICIENT_EVIDENCE (stable). No noninferiority evidence exists: the one head-to-head comparison is small, non-randomized (Serious ROBINS-I), and directionally favors the human tutor without significance. A nonsignificant difference does not establish the claim; as written it is untested.
- Claim assessed: Proxy OnlyClaim status changed · Sep 18, 2026 · Notable
"University students with low prior knowledge gain less, or are harmed more, by AI assistance during practice than high-prior-knowledge peers." received its first assessed status: PROXY_ONLY (unclear). The only relevant evidence is qualitative and process-tracing (struggling CS1 students' difficulties compounding, with an illusion of competence): consistent with differential harm, but entirely proxy measures with no unassisted outcome and no comparison arm.
- Claim assessed: Preliminary SignalClaim status changed · Sep 18, 2026 · Notable
"Guardrails prevent the unassisted-performance harm observed when secondary students practice with answer-giving chatbots." received its first assessed status: PRELIMINARY_SIGNAL (stable). One strong three-arm trial shows the vanilla-chat harm eliminated (not reversed) by teacher-designed guardrails — the claim's exact contrast, within a single study, site, subject, and product. Pass B read this as PROBABLE on the strength of the internal contrast; the single-study rule holds it at PRELIMINARY_SIGNAL until replicated.
- Claim assessed: ContradictoryClaim status changed · Sep 18, 2026 · Notable
"Guardrailed AI tutors (hints, no direct answers, pedagogical prompting) improve secondary students' performance on assessments taken without AI, compared with business-as-usual instruction." received its first assessed status: CONTRADICTORY (unclear). A well-powered trial shows no unassisted gain when guardrailed tutoring replaces equivalent practice time; a preliminary supervised after-school program shows +0.24 SD when the program adds instruction. Candidate explanation: what the comparator holds constant (time and materials). No design yet isolates it.
- Claim assessed: Insufficient EvidenceClaim status changed · Sep 18, 2026 · Notable
"Unrestricted use of a general-purpose LLM chatbot during practice improves university students' subsequent performance on assessments taken without AI." received its first assessed status: INSUFFICIENT_EVIDENCE (unclear). Evidence leans against the claim — a significant post-use quality deficit vs web search, and an underpowered retention null — but nothing adequately powered tests it directly, and the largest randomized offer cannot (partial AI at its exam). Pass B read the same evidence as NOT_SUPPORTED; the conservative status stands until something powered exists.
- Claim assessed: Not SupportedClaim status changed · Sep 18, 2026 · Major
"Unrestricted use of a general-purpose LLM chatbot during practice improves secondary students' subsequent performance on assessments taken without AI." received its first assessed status: NOT_SUPPORTED (stable). The claim asserts improvement. The one direct, gate-passing test (preregistered cluster RCT, HIGH) found a 17% unassisted-exam deficit; nothing gate-passing supports the claim. A single setting and product generation keep this NOT_SUPPORTED rather than a harm claim of its own; replication would move it.