AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

Independent Engineering Transfer After Traceable Generative-AI-Assisted Learning: A Six-University Controlled Trial with Deterministic Cluster Allocation in Agricultural Engineering Education

Design
Quasi-experimental study · randomized at class level · 656 participants
Subject and task
Other
Engineering modelling, verification, and decision-making tasks in agricultural systems: problem formulation, dimensional checks, model construction, numerical implementation, machinery performance, power balance, energy use, optimization, sensor data, uncertainty, drone/IoT planning, integrated capstone. Each substantive task followed a five-stage workflow: (1) independent attempt (variables, units, assumptions, constraints, expected range, initial strategy); (2) condition-supplied provisional alternative; (3) critical comparison against a common checklist; (4) independent verification in a separate analytical or computational environment; (5) accept/correct/reject with a defensible engineering recommendation.
Exposure: 16-week module (3 February to 19 May 2025); planned dose 64 contact hours plus 32 laboratory hours
Population match
Secondary students: Different · University students: Direct
Attrition
656 assigned (330/326); 641 analyzed immediate (missingness 2.3%, withdrawal only), 589 delayed (10.2%; absence, medical, transfer, withdrawal); likelihood-based available-outcome mixed models, pattern-mixture stress test for delayed missingness (interval first includes zero at a 16-point adverse differential shift).
Study tier
Study tier 2The decisive design fact is that final class assignment was a deterministic constrained minimization computed after baseline covariates were observed — not randomization — so this is a quasi-experimental cluster allocation and the authors themselves state that it limits causal interpretation and that the intervals are model-based rather than randomization-based. Within that constraint, conduct and analysis are unusually strong for this literature: a matched active control equated on tasks, contact time, software, feedback, portfolio, and verification demands; a prespecified SAP predating database lock; cluster-appropriate REML mixed ANCOVA with Kenward-Roger small-sample inference across 28 clusters; masked duplicate scoring with high reliability; low and near-balanced attrition (2.3% immediate, 10.2% delayed) with an adverse pattern-mixture stress test; and consistent sensitivity estimates (leave-one-class-out range 2.60-3.08, class-level, per-protocol, institution-fixed-effects, attendance-adjusted all close to the primary 2.78). The endpoints are supervised, tool-free, condition-neutral transfer tests at 7 days and ~16 weeks, exactly the kind of outcome this evidence model privileges, and precision is good (CI half-width ~0.77 points on a 32-point scale). The residual threats are real: possible residual confounding from the data-dependent allocation rule, an author-developed instrument with reliability evidence but no content- or construct-validation, unblinded students and instructors, no prospective public registration, and bundling of adaptivity with the GenAI source so the contrast attaches to the whole configuration rather than GenAI per se. Those cap the rating below HIGH; the design discipline, active comparison, precision, and robustness keep it clearly above LOW. MODERATE.
Assessment
Version 1 · AI: two passes agreed · Oct 1, 2026

Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.

Across six Uzbek universities, 28 intact agricultural-engineering classes (656 students) were assigned — by a deterministic balancing algorithm, not randomly — to a 16-week module in which a provisional alternative solution after each student's independent attempt came either from a tightly bounded Microsoft Copilot (GPT-4o) dialogue (max four turns, tools disabled, full trace kept, mandatory independent verification) or from a matched curated teacher/textbook bank. On supervised tool-free transfer tests, the GenAI classes scored 2.78 points higher (of 32) seven days after the module and 2.34 points higher ~16 weeks later, with robust sensitivity analyses; the non-random allocation and bundled configuration limit causal and component-level interpretation.

01Findings

What the study measured

Arms

ArmWhat it got
A · Traceable five-stage GenAI-assisted configuration (bounded Copilot/GPT-4o dialogue)
n = 330
Intervention · teacher supervised
Microsoft 365 Copilot Chat in Microsoft Teams (institutional interface, Microsoft 365 Education A3 licences) · GPT-4o · Guardrailed tutor · chat interface · in class · used 2025-02-03..2025-05-19
B · Structured active control (version-controlled curated alternative)
n = 326
Active comparison · teacher supervised
Same 16-week module, same five-stage workflow, same tasks, instructional time, disciplinary software, portfolio requirements, instructor feedback, and verification criteria; the provisional alternative came from a version-controlled teacher/textbook bank matched to the same task, intended learning outcome, approximate length, complexity, and critique demand, with up to three predefined (non-adaptive) clarifications permitted. The conditions differed only in the source and adaptivity of the provisional alternative.

Findings

A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.

Finding and conditionsResult
F1 · Form A: supervised condition-neutral engineering-transfer assessment, eight-criterion rubric, 0-4 per criterion, total 0-32, unfamiliar integrated agricultural engineering scenarios (primary outcome)
Traceable five-stage GenAI-assisted configuration (bounded Copilot/GPT-4o dialogue) vs Structured active control (version-controlled curated alternative)Measured without AIDelayed post-test · 7 days after module completion (module ended 19 May 2025; Form A administered 26 May 2025). NOTE: the paper calls this the "immediate" assessment; it sits exactly on the IN-002 boundary (delayed_post = 7 days or more), so it is coded delayed_post here. The paper states no study-related instruction occurred in the interval and that the seven-day gap was a deliberate consolidation interval. Near transferResearcher-developed testPrespecification unclear
Benefit
+2.78 points on the 0-32 scale (adjusted means 18.77 GenAI vs 15.98 control; adjusted SE = 0.365; model-based d = 0.72; ICC = 0.014; 8.7% of the scale, ~0.35 point per criterion if spread evenly)
2.02 to 3.55 (Kenward-Roger 95% CI; df = 19.56)
p < 0.001
Author-developed parallel form scored masked in duplicate; pilot reliability high (ICC(A,2) 0.992; median weighted kappa 0.936) but no content- or construct-validation study. Condition-neutral and built on unfamiliar integrated scenarios (not practiced items), hence near rather than same_content; the eight-criterion rubric mirrors the trained modelling-verification workflow, a shared-alignment risk. Form A failed the prespecified +/-1.5-point cross-form tolerance against Forms P and B, so only the covariate-adjusted between-condition contrast (not raw change) is interpretable.
F2 · Form B: parallel supervised condition-neutral engineering-transfer assessment, same eight-criterion 0-32 rubric, unfamiliar integrated scenarios (delayed outcome)
Traceable five-stage GenAI-assisted configuration (bounded Copilot/GPT-4o dialogue) vs Structured active control (version-controlled curated alternative)Measured without AIDelayed post-test · approximately 16 weeks (112 days) after module completion and 15 weeks (105 days) after Form A: module ended 19 May 2025, Form B administered 8 September 2025Near transferResearcher-developed testPrespecification unclear
Benefit
+2.34 points on the 0-32 scale (adjusted means 18.61 GenAI vs 16.27 control; adjusted SE = 0.395; model-based d = 0.51; estimated class variance effectively zero, boundary estimate with optimizer precision-loss message reported)
1.54 to 3.14 (Kenward-Roger 95% CI; df = 36.23)
p < 0.001
Same author-developed instrument family and masked duplicate scoring as F1; Form B DID satisfy the +/-1.5-point cross-form tolerance against baseline Form P. Reliability strong, validity unverified; near-transfer within the trained engineering domain on unfamiliar scenarios.
Limitations
Stated by the authors: deterministic post-baseline allocation: residual confounding and data-dependent-allocation uncertainty remain; CIs are model-based, not randomization-based students and instructors could not be blinded to the source and interaction format of the alternative bounded GenAI dialogue was adaptive while control clarifications were predefined; the study cannot isolate the model, adaptivity, dialogue, or traceability components no prospective public registration forms developed by the author group; reliability established but no separate formal content-validity index or construct-validation study; reliability alone does not establish validity Form A did not meet the prespecified +/-1.5-point cross-form tolerance with Forms P and B; raw longitudinal change uninterpretable delayed missingness ~10%; deterministic stress test cannot replace formal multiple imputation per-protocol cohort much smaller (345) because the 80% attendance criterion was strict (287 below 80%); supportive estimates vulnerable to selection script order in scoring was not randomized optimizer precision-loss message; delayed class variance on the boundary (effectively zero) single-language (Uzbek), single-discipline, six-university Uzbekistan context bounds transferability unrecorded independent study during the seven-day interval before Form A cannot be excluded Observed by review: the outcome rubric is the same eight-criterion engineering-transfer rubric embedded in instruction for both arms; scenarios were unfamiliar, but criterion-level alignment with the trained verification workflow could inflate the measured contrast relative to a curriculum-independent measure (alignment risk, not condition-specific since the rubric was condition-neutral) mean attendance (78.48% / 76.82%) was below the intended dose in both arms, so the estimate reflects incomplete uptake of the 16-week configuration the historical protocol and SAP used the terms randomization and intention-to-treat; the manuscript reinterprets these post hoc as deterministic allocation and available-outcome analysis — transparent, but the analysis labels were re-framed after the fact allocation audit log (inputs, code, hashes) and signed protocol/SAP are held privately, so the deterministic allocation is not independently verifiable from the public package no scored process data (prompt quality, detected AI errors, verification behavior) in the locked database; mechanism claims rest on design intent only both arms received the full five-stage workflow, so the contrast is only the source and adaptivity of the provisional alternative — good for internal comparison, but the trial says nothing about the workflow itself versus ordinary instruction

Who was studied

Level
university
Years
second- and third-year undergraduate classes
Ages
18 years or older (adults; exact range not reported)
Country
UZ
Setting
Six universities in Uzbekistan (Karshi State Technical University; Samarkand State University of Veterinary Medicine, Livestock and Biotechnologies; Termez State University of Engineering and Agrotechnologies; Tashkent State Agrarian University; Andijan Institute of Agriculture and Agrotechnology; TIIAME National Research University), purposively selected for accredited agricultural-mechanization programmes, willingness, regional diversity, and computer/internet capacity. Nine volunteer instructors (>=3 years experience). All eligible intact classes of participating instructors recruited consecutively Sept-Oct 2023 until 28 classes enrolled. Instruction and assessment entirely in Uzbek.
Prior knowledge
baseline Form P engineering-transfer assessment administered to all students before allocation; baseline scores, GPAs, ages, sex similar across conditions (partly by construction of the minimization)

Methodological notes

Extracted from: https://www.mdpi.com/2076-3417/16/18/8949 Access: full text (gold OA, CC-BY), retrieved 2026-09-30

02Funding

Funding and conflicts

Funding
No external funding (stated).
Vendor funded
No
Vendor
None identified
Notes
Authors declare no conflicts of interest. Microsoft products used under institutional Education A3 licences with no reported vendor role; Gemini (Google) used post-lock for language refinement only, per the acknowledgment.

Funding is shown on every study and never used to score it.

03Claims

Claims this study bears on

  • Learning gains from AI-assisted practice persist on delayed (one week or longer) retention tests taken without AI.Supports · rests on finding F2

    Direct, gate-passing test of the claim at ~16 weeks: supervised tool-free transfer advantage of +2.34/32 [1.54, 3.14], d=0.51, robust to adverse missing-data shifts through 12 points, with no credible decay from the 7-day measurement. University population direct. Tempered in weight, not stance: allocation was deterministic (model-based CIs), the effect attaches to a complete guardrailed traceable configuration rather than AI access per se, and the subject (agricultural engineering) sits outside the concentration subjects — an indirectness note for the body, not a bar to the stance.

  • Learning gains from AI-assisted practice persist on delayed (one week or longer) retention tests taken without AI.Supports · rests on finding F1

    A second, earlier delayed timepoint: +2.78/32 [2.02, 3.55], d=0.72 at exactly 7 days — the claim's 'one week or longer' boundary, which IN-002 codes delayed_post. Same strengths and tempering as F2; linking both timepoints lets the claim page show persistence across the interval rather than a single snapshot.

04Source

The source, as retrieved

Abstract

Generative artificial intelligence (GenAI) can support engineering problem solving, but whether AI-assisted practice transfers to independent performance after the tool is removed remains unclear. This multicentre controlled trial evaluated a traceable five-stage GenAI-assisted learning configuration in agricultural engineering education. Twenty-eight second- and third-year classes from six universities in Uzbekistan were assigned within nine teacher blocks by deterministic constrained minimization to GenAI (14 classes) or structured active-control (14 classes) groups. Both groups completed the same 16-week module, tasks, software, contact time, feedback, and verification requirements. They differed in the source and adaptivity of a provisional alternative used after an independent attempt: bounded adaptive GenAI dialogue versus a version-controlled curated alternative. The full assigned cohort included 656 students; likelihood-based available-outcome analyses included 641 immediate and 589 delayed outcomes. Kenward–Roger analyses estimated an adjusted immediate difference of 2.78 points (95% confidence interval (CI) [2.02, 3.55]; p < 0.001; model-based d = 0.72) and a delayed difference of 2.34 points (95% CI [1.54, 3.14]; p < 0.001; d = 0.51). The results show a positive adjusted association for the evaluated traceable instructional configuration, but deterministic post-baseline allocation limits causal interpretation.

Where this record came from

SourceRetrievedIdentifier
monthly-sweepOct 1, 2026doi:10.3390/app16188949
Screened in 2026-09-30; metadata verified on the publisher page the same day; full text retrieved (CC-BY).