Independent Engineering Transfer After Traceable Generative-AI-Assisted Learning: A Six-University Controlled Trial with Deterministic Cluster Allocation in Agricultural Engineering Education
- Design
- Quasi-experimental study · randomized at class level · 656 participants
- Subject and task
- Other
Engineering modelling, verification, and decision-making tasks in agricultural systems: problem formulation, dimensional checks, model construction, numerical implementation, machinery performance, power balance, energy use, optimization, sensor data, uncertainty, drone/IoT planning, integrated capstone. Each substantive task followed a five-stage workflow: (1) independent attempt (variables, units, assumptions, constraints, expected range, initial strategy); (2) condition-supplied provisional alternative; (3) critical comparison against a common checklist; (4) independent verification in a separate analytical or computational environment; (5) accept/correct/reject with a defensible engineering recommendation.
Exposure: 16-week module (3 February to 19 May 2025); planned dose 64 contact hours plus 32 laboratory hours - Population match
- Secondary students: Different · University students: Direct
- Attrition
- 656 assigned (330/326); 641 analyzed immediate (missingness 2.3%, withdrawal only), 589 delayed (10.2%; absence, medical, transfer, withdrawal); likelihood-based available-outcome mixed models, pattern-mixture stress test for delayed missingness (interval first includes zero at a 16-point adverse differential shift).
- Study tier
- Study tier 2The decisive design fact is that final class assignment was a deterministic constrained minimization computed after baseline covariates were observed — not randomization — so this is a quasi-experimental cluster allocation and the authors themselves state that it limits causal interpretation and that the intervals are model-based rather than randomization-based. Within that constraint, conduct and analysis are unusually strong for this literature: a matched active control equated on tasks, contact time, software, feedback, portfolio, and verification demands; a prespecified SAP predating database lock; cluster-appropriate REML mixed ANCOVA with Kenward-Roger small-sample inference across 28 clusters; masked duplicate scoring with high reliability; low and near-balanced attrition (2.3% immediate, 10.2% delayed) with an adverse pattern-mixture stress test; and consistent sensitivity estimates (leave-one-class-out range 2.60-3.08, class-level, per-protocol, institution-fixed-effects, attendance-adjusted all close to the primary 2.78). The endpoints are supervised, tool-free, condition-neutral transfer tests at 7 days and ~16 weeks, exactly the kind of outcome this evidence model privileges, and precision is good (CI half-width ~0.77 points on a 32-point scale). The residual threats are real: possible residual confounding from the data-dependent allocation rule, an author-developed instrument with reliability evidence but no content- or construct-validation, unblinded students and instructors, no prospective public registration, and bundling of adaptivity with the GenAI source so the contrast attaches to the whole configuration rather than GenAI per se. Those cap the rating below HIGH; the design discipline, active comparison, precision, and robustness keep it clearly above LOW. MODERATE.
- Assessment
- Version 1 · AI: two passes agreed · Oct 1, 2026
Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.
Across six Uzbek universities, 28 intact agricultural-engineering classes (656 students) were assigned — by a deterministic balancing algorithm, not randomly — to a 16-week module in which a provisional alternative solution after each student's independent attempt came either from a tightly bounded Microsoft Copilot (GPT-4o) dialogue (max four turns, tools disabled, full trace kept, mandatory independent verification) or from a matched curated teacher/textbook bank. On supervised tool-free transfer tests, the GenAI classes scored 2.78 points higher (of 32) seven days after the module and 2.34 points higher ~16 weeks later, with robust sensitivity analyses; the non-random allocation and bundled configuration limit causal and component-level interpretation.
What the study measured
Arms
| Arm | What it got |
|---|---|
| A · Traceable five-stage GenAI-assisted configuration (bounded Copilot/GPT-4o dialogue) n = 330 | Intervention · teacher supervised Microsoft 365 Copilot Chat in Microsoft Teams (institutional interface, Microsoft 365 Education A3 licences) · GPT-4o · Guardrailed tutor · chat interface · in class · used 2025-02-03..2025-05-19 |
| B · Structured active control (version-controlled curated alternative) n = 326 | Active comparison · teacher supervised Same 16-week module, same five-stage workflow, same tasks, instructional time, disciplinary software, portfolio requirements, instructor feedback, and verification criteria; the provisional alternative came from a version-controlled teacher/textbook bank matched to the same task, intended learning outcome, approximate length, complexity, and critique demand, with up to three predefined (non-adaptive) clarifications permitted. The conditions differed only in the source and adaptivity of the provisional alternative. |
Findings
A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.
| Finding and conditions | Result |
|---|---|
| F1 · Form A: supervised condition-neutral engineering-transfer assessment, eight-criterion rubric, 0-4 per criterion, total 0-32, unfamiliar integrated agricultural engineering scenarios (primary outcome) Traceable five-stage GenAI-assisted configuration (bounded Copilot/GPT-4o dialogue) vs Structured active control (version-controlled curated alternative)Measured without AIDelayed post-test · 7 days after module completion (module ended 19 May 2025; Form A administered 26 May 2025). NOTE: the paper calls this the "immediate" assessment; it sits exactly on the IN-002 boundary (delayed_post = 7 days or more), so it is coded delayed_post here. The paper states no study-related instruction occurred in the interval and that the seven-day gap was a deliberate consolidation interval. Near transferResearcher-developed testPrespecification unclear | Benefit +2.78 points on the 0-32 scale (adjusted means 18.77 GenAI vs 15.98 control; adjusted SE = 0.365; model-based d = 0.72; ICC = 0.014; 8.7% of the scale, ~0.35 point per criterion if spread evenly) 2.02 to 3.55 (Kenward-Roger 95% CI; df = 19.56) p < 0.001 Author-developed parallel form scored masked in duplicate; pilot reliability high (ICC(A,2) 0.992; median weighted kappa 0.936) but no content- or construct-validation study. Condition-neutral and built on unfamiliar integrated scenarios (not practiced items), hence near rather than same_content; the eight-criterion rubric mirrors the trained modelling-verification workflow, a shared-alignment risk. Form A failed the prespecified +/-1.5-point cross-form tolerance against Forms P and B, so only the covariate-adjusted between-condition contrast (not raw change) is interpretable. |
| F2 · Form B: parallel supervised condition-neutral engineering-transfer assessment, same eight-criterion 0-32 rubric, unfamiliar integrated scenarios (delayed outcome) Traceable five-stage GenAI-assisted configuration (bounded Copilot/GPT-4o dialogue) vs Structured active control (version-controlled curated alternative)Measured without AIDelayed post-test · approximately 16 weeks (112 days) after module completion and 15 weeks (105 days) after Form A: module ended 19 May 2025, Form B administered 8 September 2025Near transferResearcher-developed testPrespecification unclear | Benefit +2.34 points on the 0-32 scale (adjusted means 18.61 GenAI vs 16.27 control; adjusted SE = 0.395; model-based d = 0.51; estimated class variance effectively zero, boundary estimate with optimizer precision-loss message reported) 1.54 to 3.14 (Kenward-Roger 95% CI; df = 36.23) p < 0.001 Same author-developed instrument family and masked duplicate scoring as F1; Form B DID satisfy the +/-1.5-point cross-form tolerance against baseline Form P. Reliability strong, validity unverified; near-transfer within the trained engineering domain on unfamiliar scenarios. |
- Limitations
- Stated by the authors: deterministic post-baseline allocation: residual confounding and data-dependent-allocation uncertainty remain; CIs are model-based, not randomization-based students and instructors could not be blinded to the source and interaction format of the alternative bounded GenAI dialogue was adaptive while control clarifications were predefined; the study cannot isolate the model, adaptivity, dialogue, or traceability components no prospective public registration forms developed by the author group; reliability established but no separate formal content-validity index or construct-validation study; reliability alone does not establish validity Form A did not meet the prespecified +/-1.5-point cross-form tolerance with Forms P and B; raw longitudinal change uninterpretable delayed missingness ~10%; deterministic stress test cannot replace formal multiple imputation per-protocol cohort much smaller (345) because the 80% attendance criterion was strict (287 below 80%); supportive estimates vulnerable to selection script order in scoring was not randomized optimizer precision-loss message; delayed class variance on the boundary (effectively zero) single-language (Uzbek), single-discipline, six-university Uzbekistan context bounds transferability unrecorded independent study during the seven-day interval before Form A cannot be excluded Observed by review: the outcome rubric is the same eight-criterion engineering-transfer rubric embedded in instruction for both arms; scenarios were unfamiliar, but criterion-level alignment with the trained verification workflow could inflate the measured contrast relative to a curriculum-independent measure (alignment risk, not condition-specific since the rubric was condition-neutral) mean attendance (78.48% / 76.82%) was below the intended dose in both arms, so the estimate reflects incomplete uptake of the 16-week configuration the historical protocol and SAP used the terms randomization and intention-to-treat; the manuscript reinterprets these post hoc as deterministic allocation and available-outcome analysis — transparent, but the analysis labels were re-framed after the fact allocation audit log (inputs, code, hashes) and signed protocol/SAP are held privately, so the deterministic allocation is not independently verifiable from the public package no scored process data (prompt quality, detected AI errors, verification behavior) in the locked database; mechanism claims rest on design intent only both arms received the full five-stage workflow, so the contrast is only the source and adaptivity of the provisional alternative — good for internal comparison, but the trial says nothing about the workflow itself versus ordinary instruction
Who was studied
- Level
- university
- Years
- second- and third-year undergraduate classes
- Ages
- 18 years or older (adults; exact range not reported)
- Country
- UZ
- Setting
- Six universities in Uzbekistan (Karshi State Technical University; Samarkand State University of Veterinary Medicine, Livestock and Biotechnologies; Termez State University of Engineering and Agrotechnologies; Tashkent State Agrarian University; Andijan Institute of Agriculture and Agrotechnology; TIIAME National Research University), purposively selected for accredited agricultural-mechanization programmes, willingness, regional diversity, and computer/internet capacity. Nine volunteer instructors (>=3 years experience). All eligible intact classes of participating instructors recruited consecutively Sept-Oct 2023 until 28 classes enrolled. Instruction and assessment entirely in Uzbek.
- Prior knowledge
- baseline Form P engineering-transfer assessment administered to all students before allocation; baseline scores, GPAs, ages, sex similar across conditions (partly by construction of the minimization)
Methodological notes
Extracted from: https://www.mdpi.com/2076-3417/16/18/8949 Access: full text (gold OA, CC-BY), retrieved 2026-09-30
Funding and conflicts
- Funding
- No external funding (stated).
- Vendor funded
- No
- Vendor
- None identified
- Notes
- Authors declare no conflicts of interest. Microsoft products used under institutional Education A3 licences with no reported vendor role; Gemini (Google) used post-lock for language refinement only, per the acknowledgment.
Funding is shown on every study and never used to score it.
Claims this study bears on
- Learning gains from AI-assisted practice persist on delayed (one week or longer) retention tests taken without AI.Supports · rests on finding F2
Direct, gate-passing test of the claim at ~16 weeks: supervised tool-free transfer advantage of +2.34/32 [1.54, 3.14], d=0.51, robust to adverse missing-data shifts through 12 points, with no credible decay from the 7-day measurement. University population direct. Tempered in weight, not stance: allocation was deterministic (model-based CIs), the effect attaches to a complete guardrailed traceable configuration rather than AI access per se, and the subject (agricultural engineering) sits outside the concentration subjects — an indirectness note for the body, not a bar to the stance.
- Learning gains from AI-assisted practice persist on delayed (one week or longer) retention tests taken without AI.Supports · rests on finding F1
A second, earlier delayed timepoint: +2.78/32 [2.02, 3.55], d=0.72 at exactly 7 days — the claim's 'one week or longer' boundary, which IN-002 codes delayed_post. Same strengths and tempering as F2; linking both timepoints lets the claim page show persistence across the interval rather than a single snapshot.
The source, as retrieved
Abstract
Generative artificial intelligence (GenAI) can support engineering problem solving, but whether AI-assisted practice transfers to independent performance after the tool is removed remains unclear. This multicentre controlled trial evaluated a traceable five-stage GenAI-assisted learning configuration in agricultural engineering education. Twenty-eight second- and third-year classes from six universities in Uzbekistan were assigned within nine teacher blocks by deterministic constrained minimization to GenAI (14 classes) or structured active-control (14 classes) groups. Both groups completed the same 16-week module, tasks, software, contact time, feedback, and verification requirements. They differed in the source and adaptivity of a provisional alternative used after an independent attempt: bounded adaptive GenAI dialogue versus a version-controlled curated alternative. The full assigned cohort included 656 students; likelihood-based available-outcome analyses included 641 immediate and 589 delayed outcomes. Kenward–Roger analyses estimated an adjusted immediate difference of 2.78 points (95% confidence interval (CI) [2.02, 3.55]; p < 0.001; model-based d = 0.72) and a delayed difference of 2.34 points (95% CI [1.54, 3.14]; p < 0.001; d = 0.51). The results show a positive adjusted association for the evaluated traceable instructional configuration, but deterministic post-baseline allocation limits causal interpretation.
Where this record came from
| Source | Retrieved | Identifier |
|---|---|---|
| monthly-sweep | Oct 1, 2026 | doi:10.3390/app16188949 Screened in 2026-09-30; metadata verified on the publisher page the same day; full text retrieved (CC-BY). |