AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

AI feedback on writing produces unassisted writing growth noninferior to human feedback — no worse than a stated margin.

Current status
Insufficient evidenceEvidence stable · what the nine statuses mean
Why
No noninferiority evidence exists: the one head-to-head comparison is small, non-randomized (Serious ROBINS-I), and directionally favors the human tutor without significance. A nonsignificant difference does not establish the claim; as written it is untested.
What would change it
Adequately powered head-to-head trials whose confidence intervals exclude a stated meaningful deficit.
Linked evidence
1 link · 1 consistent, but doesn't test the claim
Last updated
Sep 18, 2026

The status is the collection's own judgment on its nine-label scale — not a certainty rating, not a GRADE level, and not advice. Every status change is dated, reasoned, and kept below under History.

01The evidence

The evidence, sorted by what it shows

Consistent, but doesn't test the claim (1)

  • AI-generated feedback on writing: insights into efficacy and ENL student preferenceQuasi-experimental study · 2023 · Study tier 3 · Secondary different · University direct

    No significant group-by-time difference in unassisted writing growth (p=.085, trend favoring the human tutor). A nonsignificant difference cannot test a noninferiority claim: no margin was stated or tested; N=48; ROBINS-I Serious. Rests on finding F1, judged at assessment v2.

02Certainty

Certainty by outcome

A claim can be broken into separate bodies of evidence — one per outcome, split by whether AI was available at assessment and when the outcome was measured. Each body carries four separate judgments: how confident we are (certainty), what the evidence points to (the conclusion), how directly it speaks to this claim (applicability), and who has stood behind the judgment. Confidence and conclusion are never merged into one word.

UA-GROWTH — Unassisted writing growth vs human feedback · AI at assessment: no · immediate post

Field Assembly certainty: Untested (not a GRADE rating — what our scale means) · Conclusion: Not estimable · Applicability: Partial · AI: two passes agreed

Rated against: noninferiority: CI excludes a stated meaningful deficit (no margin exists in the literature)

The claim as written is untested: no margin-based analysis exists

DomainJudgment and reasoning
Risk of biasvery serious · ROBINS-I Serious: unreported allocation, unadjusted baseline imbalance
Inconsistencynot serious · single study
Indirectnessserious · TA-mediated delivery; one institution
Imprecisionvery serious · N=48; interval consistent with meaningful differences in either direction
Reporting and publication biasnot serious · not assessable
Why untestedNo study performs a margin-based (noninferiority or equivalence) analysis; the one comparison reports a nonsignificant difference, which cannot answer the claim as written (Rule 6)

Population: university English-as-new-language learners
Comparator: weekly one-on-one human tutor feedback
Assessed by software, two independent passes · search: seed/corpus.yaml (verified inventory, proposal section 7) + data/searches/ · method: EVIDENCE-MODEL.md v2 + fa-certainty-scale v2

1 study in this body.

03History

How this assessment has changed

  • Insufficient evidenceSep 18, 2026 · Evidence stable

    initial assessed status (Q-009, owner-accepted IN-018)

  • Not yet assessedSep 18, 2026 · Evidence unclear

    initial curated status