About
A structured record of the evidence on generative AI in learning: what has been studied, in whom, under what conditions, with what result — and how much of it measures what learners can do once the AI is taken away.
How assessments are made
Evidence comes from peer-reviewed journals, preprint servers and working-paper series (both marked as preliminary until peer-reviewed), and official reports. Press releases, marketing, and news coverage are never used as evidence. Vendor-published efficacy material is recorded as a claim in circulation, under the same standards as everything else, and never carries a claim on its own.
For each study the collection keeps two things apart: the facts reported in the paper — participants, arms, effect sizes — and our judgments about them, such as how closely the population matches the two groups this collection covers, or how much a result can be relied on. Judgments are versioned; earlier versions stay on record.
The gating rule. The unit of evidence here is the finding: one outcome contrast, badged with five dimensions — whether AI was available at assessment, when the outcome was measured, what kind of measure it was, how far it sits from the practiced material, and over what durations. A finding measured while the AI was still available records assisted performance: a fact about the learner-plus-tool system. It can support statements about assisted performance and nothing else. Only findings measured without AI can support or contradict a learning claim, whatever the size of the assisted gain. Proxy outcomes — engagement, satisfaction, self-reported learning, the quality of AI-assisted work, process metrics — are recorded but can never carry a learning claim either.
Study tiers describe a single study, weighing design, randomization level and analysis, comparison quality — what the AI arm got that the comparator didn't — baseline equivalence, attrition, and sample size as precision, not virtue. Certainty ratings are different: they apply to a whole body of evidence for one outcome, per claim and per set of assessment conditions. This collection rates certainty on its own scale — well supported, supported, weak, or untested — using the same five domains a formal GRADE assessment uses: risk of bias, inconsistency, indirectness, imprecision, and reporting bias. Every rating shows its reasoning for all five; a rating whose reasoning is missing is withheld, because the reasoning is the rating.
Our scale is not GRADE, and we do not describe it as GRADE. Where someone else has published a formal GRADE rating for an outcome, we import it and show it as theirs, with the citation. The two are always labeled so you can tell which you are reading. Alongside certainty, each body separately states its conclusion (benefit, harm, no meaningful effect, varies by design, inconsistent, or not estimable), its applicability, and who has stood behind the judgment. Confidence and conclusion are never merged: "mixed" is not a rating here, because being confident that effects differ by design and being unsure of everything are different states of knowledge.
One level is worth explaining. Untested does not mean the evidence is poor. It means the question has not been properly asked: no study in the body uses a design that could answer it, or the studies that could answer it did not report the result. That is a fact about the literature, and it is different from a weak answer. Not yet assessed is different again — it means we have not done the work, and says nothing about the evidence.
Two more rules shape what you will read. "As good as" claims are held to a stated margin: a nonsignificant difference between AI and human feedback is not evidence of equivalence, and no claim of noninferiority is supported here without an analysis whose confidence interval excludes a stated meaningful deficit. And effects are never translated into intuitive units ("months of learning") unless a published, population- and test-appropriate benchmark is cited beside the conversion; at present none is approved, so no conversions appear.
The nine claim statuses
A claim status is the collection's own judgment across every linked body of evidence for one claim, in one population. It sits above outcome certainty, it is assigned only through two independent review passes plus an owner decision, and every change is dated and reasoned on the claim's page. The nine labels:
- Established
- Consistent support from multiple strong, adequately sized studies measured without AI. Overturning it would be surprising.
- Probable
- The direct evidence favors the claim, with meaningful residual uncertainty — typically breadth: few sites, few subjects, or few designs.
- Preliminary signal
- At least one well-conducted direct test supports the claim, but it stands nearly alone — a single study, site, or immediate-only outcome — so the result could still be an artifact of where and how it was measured.
- Plausible, unproven
- Nothing adequate tests the claim directly; what surrounds it makes it credible. This is a statement about absence, not a weak yes.
- Proxy outcomes only
- Everything linked to the claim measures proxies — engagement, satisfaction, assisted output quality, process metrics — which cannot carry a learning claim regardless of their size or consistency.
- Insufficient evidence
- Direct evidence exists but is too thin, imprecise, or methodologically limited to support any verdict, for or against. Distinct from untested: the question has been asked, just not answered.
- Contradictory
- Well-conducted direct tests point in opposite directions, and the rationale names the best candidate explanations for the split — what the studies held constant, and what they didn't.
- Not supported
- The claim's best direct tests found no support: adequately conducted studies measured without AI show no benefit, or point the other way.
- Evidence of harm
- The direct evidence indicates the practice worsens unassisted performance — the learner does worse afterward than they would have without it.
A tenth state, Not yet assessed, is not a status: it means the two-pass review has not finished, and a claim in that state is not published at all.
What the automation does and does not establish
This collection is produced with AI assistance, and says so. Every extraction and appraisal judgment is made by two independent AI passes over the same source; agreement publishes with the label "AI: two passes agreed", and disagreement goes to a human owner, whose decision is recorded. Claim statuses and certainty ratings additionally always require an owner decision, whatever the passes concluded.
Two-pass agreement is a reliability check: it catches slips and idiosyncratic readings. It cannot certify correctness, because two passes can share the same bias or the same misreading. Verbatim-quote verification establishes provenance — the quoted words exist contiguously in the source — not that the quote was interpreted correctly or is representative. Validity rests on the human review layer, on the published reasoning being open to challenge, and on the corrections channel below; the automation reduces error rates and cost, and is disclosed as exactly that.
Where two passes disagreed and nothing has settled the disagreement, the result is published as "Assessed, not settled" with both readings, rather than hidden — the held-back set would otherwise be exactly the contested judgments.
Errors: write to hello@fieldassembly.net with a link to the paper you are reading against. Corrections are made on the record, never silently.
The limits
This is a research record, not guidance. It does not recommend adopting, mandating, restricting or banning any tool or teaching practice, and nothing here should be read as advice about any particular learner, course, or institution. Assessments can be wrong; their review-status labels mean what they say. Always read the original source, which every study page links to.
The collection is independent of every AI vendor, publisher, journal, regulator and institution it mentions, accepts no vendor funding, and shows the funding of every study it records.