Impact of AI assistance on student agency
- Design
- Randomized trial · randomized at student level · 1625 participants
- Subject and task
- Other
Peer review of student-created learning resources in RiPPLE (a learnersourcing platform): students rated resources on a 4-item rubric (alignment, correctness, difficulty, critical thinking) and wrote justification/feedback comments to the author. AI quality-control functions (rule-based suggestion detection, SBERT relatedness score, GLEU similarity score) flagged low-quality comments and prompted revision.
Exposure: 8 weeks total (first 8 weeks of semester 2, 2020). Phase 1 (weeks 1-4): all participants received AI prompts during peer review. Phase 2 (weeks 5-8): random assignment to four conditions (AI continued / NR no prompts / SR self-monitoring checklist / SAI checklist plus prompts). - Population match
- Secondary students: Different · University students: Direct
- Attrition
- not_reported as a flow statement; group counts are identical in the phase-1 and phase-2 summary tables (396/409/402/418, total 1625), indicating the same consenting cohort was analysed in both phases, but per-student dropout or reduced review activity is not reported (phase-2 review volume is lower than phase 1: 11,243 vs 16,007 reviews).
- Study tier
- Study tier 2Randomised controlled field experiment with individual-level assignment of 1625 students to four conditions after a shared four-week AI-scaffold phase, with demonstrated baseline equivalence on all six measures during that phase and large samples per arm giving good precision (one-way ANOVA plus Tukey HSD pairwise tests with CIs and Cohen's d). Rated MODERATE rather than HIGH because outcomes are automated process/quality proxies for peer-review quality (flag rate, SBERT relatedness, GLEU similarity, length, time, likes) partly generated by the same algorithms that powered the intervention prompts, not independent learning assessments; no preregistration is reported; and the design lacks a never-assisted arm, so "reliance" is inferred from the withdrawal contrast against a still-assisted control.
- Assessment
- Version 2 · AI: two passes agreed · Sep 18, 2026
Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.
In 10 Australian university courses, 1625 students wrote peer reviews with AI quality prompts for four weeks, then were randomly split into four groups: prompts kept, prompts removed, prompts replaced by a self-monitoring checklist, or checklist plus prompts. When the AI prompts were removed, review quality dropped (more flagged reviews, more generic/repetitive and less relevant, shorter comments), suggesting students had relied on the AI rather than learning from it. The self-monitoring checklist partly, but not fully, compensated for the removed AI, and adding the checklist on top of AI prompts was no better than AI prompts alone.
What the study measured
Arms
| Arm | What it got |
|---|---|
| A · AI (control) - AI-assistance prompts continued in weeks 5-8 n = 396 | Intervention · teacher independent RiPPLE AI-assistance prompts · SBERT/BERT embeddings plus rule-based NLP (suggestion detection; GLEU similarity); pre-LLM, not a chatbot · Scripted · feedback interface · unrestricted · used second semester 2020, weeks 1-8 |
| B · NR - Not Receiving assistance: AI prompts removed in weeks 5-8 n = 409 | Control · teacher independent |
| C · SR - Self-monitoring checklist replaces AI prompts in weeks 5-8 n = 402 | Active comparison · teacher independent |
| D · SAI - Self-monitoring checklist plus AI prompts in weeks 5-8 n = 418 | Intervention · teacher independent RiPPLE AI-assistance prompts (plus self-monitoring checklist) · SBERT/BERT embeddings plus rule-based NLP (suggestion detection; GLEU similarity); pre-LLM, not a chatbot · Scripted · feedback interface · unrestricted · used second semester 2020, weeks 1-8 |
Findings
A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.
| Finding and conditions | Result |
|---|---|
| F1 · Baseline equivalence on all six peer-review measures (flag rate, similarity, relatedness, length, time, likes) during weeks 1-4 when every group had AI prompts AI (control) - AI-assistance prompts continued in weeks 5-8 vs With AI at assessment — assisted performance, not learning evidenceDuring interventionSame contentProcess / behaviorPrespecification unclear | Null No significant group differences: flagged reviews F(3,1621)=0.866, similarity F(3,1621)=0.532, relatedness F(3,1621)=2.04, length F(3,1621)=1.92, time F(3,1621)=0.586, likes F(3,1621)=0.56; all eta-squared <= 0.004 not_reported p=0.458; p=0.660; p=0.106; p=0.125; p=0.624; p=0.641 System-logged behavioural proxies; establishes pre-withdrawal comparability of the randomised groups rather than a treatment effect. |
| F10 · All six peer-review measures for AI prompts plus self-monitoring checklist vs AI prompts alone (hybrid human-AI contrast), weeks 5-8 SAI - Self-monitoring checklist plus AI prompts in weeks 5-8 vs AI (control) - AI-assistance prompts continued in weeks 5-8With AI at assessment — assisted performance, not learning evidenceDuring interventionSame contentProcess / behaviorPrespecification unclear | Null No significant differences: flag rate SAI M=0.14 (SD=0.24) vs AI M=0.14 (SD=0.23), d=-0.01; similarity d=7e-5; relatedness d=0.04; comment length SAI M=25.67 (SD=14.75) vs AI M=23.38 (SD=12.41), t=-2.34, d=-0.16 (marginal, ns); time d=-0.003; like rate d=-0.11 flag rate 95% CI [-0.05, 0.05]; length [-4.80, 0.23] flag rate p=1.000; similarity p=1.000; relatedness p=0.959; length p=0.090; time p=1.000; likes p=0.408 Both arms assisted at assessment. SAI produced the longest comments of all four groups (marginal, ns) and the authors note descriptive gains in like-rate distribution; overall no significant advantage of the hybrid over AI alone (RQ3 answer). |
| F2 · Rate of AI-flagged (revision-needing) reviews after AI-prompt withdrawal, weeks 5-8 NR - Not Receiving assistance: AI prompts removed in weeks 5-8 vs AI (control) - AI-assistance prompts continued in weeks 5-8Measured without AIDelayed post-test · measures aggregated over weeks 1-4 after prompt withdrawal (phase 2, weeks 5-8 of semester)Same contentProcess / behaviorPrespecification unclear | Harm NR M=0.31 (SD=0.31) vs AI M=0.14 (SD=0.23); t=-8.70; Cohen's d=-0.61 (medium-to-large); omnibus F(3,1621)=39.63, eta-squared=0.068 95% CI for mean difference [-0.22, -0.12] p_tukey < .001 Withdrawal-phase (unassisted) performance for the NR arm - the learning-relevant dependence result, coded ai_available 'no' for that reason; comparison AI arm still had prompts at assessment. Flags are computed by the same AI quality-control algorithms that generated the intervention prompts. |
| F3 · Similarity (GLEU) of comments to the student's own previous comments (higher = more generic/repetitive) after withdrawal, weeks 5-8 NR - Not Receiving assistance: AI prompts removed in weeks 5-8 vs AI (control) - AI-assistance prompts continued in weeks 5-8Measured without AIDelayed post-test · measures aggregated over weeks 1-4 after prompt withdrawal (phase 2)Same contentProcess / behaviorPrespecification unclear | Harm NR M=0.32 (SD=0.17) vs AI M=0.28 (SD=0.12); t=-4.22; Cohen's d=-0.30 (small-to-medium); omnibus F(3,1621)=9.65, eta-squared=0.018 95% CI [-0.07, -0.02] p < .001 Automated GLEU n-gram metric; unassisted performance for NR arm, AI arm assisted at assessment. |
| F4 · Relatedness (SBERT) of comments to the resource under review after withdrawal, weeks 5-8 NR - Not Receiving assistance: AI prompts removed in weeks 5-8 vs AI (control) - AI-assistance prompts continued in weeks 5-8Measured without AIDelayed post-test · measures aggregated over weeks 1-4 after prompt withdrawal (phase 2)Same contentProcess / behaviorPrespecification unclear | Harm NR M=0.27 (SD=0.12) vs AI M=0.31 (SD=0.13); t=4.70; Cohen's d=0.33 (small-to-medium); omnibus F(3,1621)=12.24, eta-squared=0.026 95% CI [0.02, 0.06] p < .001 SBERT cosine-similarity proxy for feedback specificity/relevance; unassisted for NR arm, AI arm assisted. |
| F5 · Average comment length in words after withdrawal, weeks 5-8 NR - Not Receiving assistance: AI prompts removed in weeks 5-8 vs AI (control) - AI-assistance prompts continued in weeks 5-8Measured without AIDelayed post-test · measures aggregated over weeks 1-4 after prompt withdrawal (phase 2)Same contentProcess / behaviorPrespecification unclear | Harm NR M=19.43 (SD=13.15) vs AI M=23.38 (SD=12.41); t=4.03; Cohen's d=0.28 (small); omnibus F(3,1621)=14.14, eta-squared=0.026 95% CI [1.43, 6.49] p < .001 Word count as quality proxy; authors note longer comments do not necessarily guarantee higher quality but length is associated with feedback quality in prior work. |
| F6 · Time spent on reviews and rate of likes/helpfulness after withdrawal, weeks 5-8 NR - Not Receiving assistance: AI prompts removed in weeks 5-8 vs AI (control) - AI-assistance prompts continued in weeks 5-8Measured without AIDelayed post-test · measures aggregated over weeks 1-4 after prompt withdrawal (phase 2)Same contentProcess / behaviorPrespecification unclear | Null Time: NR M=219.36s (SD=866.23) vs AI M=265.44s (SD=988.78), t=0.79, d=0.06; Like rate: NR M=0.28 (SD=0.27) vs AI M=0.28 (SD=0.26), t=-0.01, d=-8e-4 Time 95% CI [-103.88, 196.03]; Likes 95% CI [-0.05, 0.05] time p=0.859; likes p=1.000 Engagement time and peer-perceived helpfulness unchanged despite quality-metric declines; like rate is peer-voted, not algorithmic. |
| F7 · Rate of flagged reviews with self-monitoring checklist replacing AI prompts (checklist-compensation contrast), weeks 5-8 SR - Self-monitoring checklist replaces AI prompts in weeks 5-8 vs AI (control) - AI-assistance prompts continued in weeks 5-8Measured without AIDelayed post-test · measures aggregated over weeks 1-4 after prompt replacement (phase 2)Same contentProcess / behaviorPrespecification unclear | Harm SR M=0.26 (SD=0.30) vs AI M=0.14 (SD=0.23); t=-6.36; Cohen's d=-0.45 (medium) - worse than continued AI, but a smaller deficit than NR's d=-0.61, which the authors read as partial compensation by the checklist 95% CI [-0.17, -0.07] p_tukey < .001 Unassisted (checklist-only) performance for SR arm, coded ai_available 'no'; AI arm assisted at assessment. SR-vs-NR is not reported as a direct pairwise test - partial compensation is inferred from relative effect sizes against the shared AI control. |
| F8 · Similarity and relatedness of comments with self-monitoring checklist replacing AI prompts, weeks 5-8 SR - Self-monitoring checklist replaces AI prompts in weeks 5-8 vs AI (control) - AI-assistance prompts continued in weeks 5-8Measured without AIDelayed post-test · measures aggregated over weeks 1-4 after prompt replacement (phase 2)Same contentProcess / behaviorPrespecification unclear | Harm Similarity: SR M=0.31 (SD=0.15) vs AI M=0.28 (SD=0.12), t=-3.20, d=-0.23; Relatedness: SR M=0.28 (SD=0.12) vs AI M=0.31 (SD=0.13), t=4.26, d=0.30 Similarity 95% CI [-0.06, -0.01]; Relatedness 95% CI [0.02, 0.06] similarity p=0.008; relatedness p < .001 SR remained worse than continued-AI on both automated quality metrics; the paper describes SR's improvements over NR on these scores as negligible. |
| F9 · Comment length, time on reviews, and like rate with checklist replacing AI prompts, weeks 5-8 SR - Self-monitoring checklist replaces AI prompts in weeks 5-8 vs AI (control) - AI-assistance prompts continued in weeks 5-8Measured without AIDelayed post-test · measures aggregated over weeks 1-4 after prompt replacement (phase 2)Same contentProcess / behaviorPrespecification unclear | Null Length: SR M=23.06 (SD=15.26) vs AI M=23.38 (SD=12.41), t=0.32, d=0.02; Time: SR M=220.23 (SD=481.87) vs AI M=265.44 (SD=988.78), t=0.77, d=0.06; Like rate: SR M=0.31 (SD=0.30) vs AI M=0.28 (SD=0.26), t=-0.55, d=-0.04 Length 95% CI [-2.22, 2.86]; Time [-105.39, 195.81]; Likes [-0.06, 0.04] length p=0.988; time p=0.867; likes p=0.947 Checklist maintained comment length at the AI-assisted level (unlike NR, which shortened), which the authors count as a success of self-monitoring. |
- Limitations
- Authors note: experiment limited to the first eight weeks of semester (no longer-term follow-up); the NR comparison group had received AI assistance for the first four weeks, so there is no never-assisted comparison; reliance on quantitative behavioural measures without qualitative interview/open-ended data; the null SAI result may reflect a design issue (AI prompts and checklist not tailored to the hybrid condition); biases of AI-generated feedback not thoroughly investigated; course content, instructional methods, motivation and instructor preferences vary across the 10 courses; generalisability beyond peer feedback tasks untested.
Who was studied
- Level
- university (undergraduate)
- Ages
- not_reported
- Country
- Australia (The University of Queensland; RiPPLE courses)
- Prior knowledge
- not_reported; all groups statistically equivalent on all six peer-review measures during the shared 4-week AI-prompt phase
- Selection
- Undergraduates in 10 courses using RiPPLE in second semester 2020, spanning Humanities & Social Sciences, Health & Behavioural Sciences, Medicine, Business, and Engineering; only students who consented to data use were included
Methodological notes
Extracted from: https://researchmgt.monash.edu/ws/portalfiles/portal/572678646/564026501_oa.pdf Access: full text (Monash University institutional repository, open-access publisher PDF, CC BY) Sample-size note: 1625 students across 10 courses; phase 2 group sizes AI n=396, NR n=409, SR n=402, SAI n=418; 16,007 peer reviews on 4,501 resources in phase 1 and 11,243 peer reviews on 3,573 resources in phase 2.
How much each result can be relied on
Risk of bias is judged per result, not per paper: the same study can report one outcome at low risk and another at high. Two reviewers assess each result independently against a written guide, and every domain judgment below carries its reasoning and the sentences it rests on.
Risk of bias is not the same as whether the numbers work. A trial can be well conducted and still print a standard deviation that is arithmetically impossible, or a table whose sign contradicts its own text. Risk of bias has no question for that, so where it was found it is recorded separately below, under Source integrity.
These words are not quality scores. On the RoB 2 scale, High risk of bias is the worst rating a result can receive and Low risk of bias is the best — the opposite of how the same words read on some other scales. Levels are always written out in full here for that reason.
WITHDRAW
High risk of bias · RoB 2 · Two reviewers agreed on every domain
How the overall was reached: Worst-domain rule: D4 is High because the headline measures for this result (flag rate, similarity, relatedness; length to a lesser degree) are computed by the intervention's own quality-control algorithms, so the AI arm could revise submissions until they passed the exact metrics used as outcomes while the NR arm could not - the effect estimate for NR vs AI on these measures is partly an artifact of measuring the intervention with itself. Randomization appears sound (baseline balance in weeks 1-4), deviations are unlikely given automated delivery, and reporting of all six measures is complete, but undescribed allocation procedures (D1), absent attrition/flow reporting (D3), and no preregistration (D5) each add Some concerns. Interpretation caveat noted by the authors themselves: NR had four weeks of prior AI exposure, so this contrast estimates withdrawal effects, not never-assisted performance.
| Domain | Judgment and reasoning |
|---|---|
| D1 · Randomisation process Some concerns about risk of bias | The study is described as a randomised controlled field experiment with phase-2 randomization of 1,625 students to four arms after four weeks of universal AI prompts, and Appendix A shows no significant between-group differences on any of the six measures during the baseline (first four weeks) period, which is consistent with successful randomization. However, the paper gives no information on how the allocation sequence was generated (e.g., simple vs stratified randomization, whether it was done within the RiPPLE platform) and no information on allocation concealment. With sequence generation and concealment methods not reported, RoB 2 supports "Some concerns" despite the reassuring baseline balance.The study was a randomised controlled field experiment, with the peer review condition as the independent variable (one control group and three experimental groups). Notably, statistical analysis revealed no significant differences in any of these measures between the groups during this period. |
| D2 · Deviations from intended interventions Low risk of bias | The interventions were delivered automatically through the RiPPLE peer-review interface (AI prompts present vs removed), so delivery was faithful to assignment by construction; there is no indication of crossover, contamination, or co-intervention differing between arms. Students and researchers could not plausibly alter which interface condition a student received. All randomized students appear to have been analysed in their assigned groups (Ns in Tables 1 and 2 match the randomized group sizes), consistent with an intention-to-treat style analysis of the effect of assignment. Participants were necessarily unblinded to the presence/absence of prompts, but no deviations from intended intervention arising from the trial context are apparent.Experiment 1 (NR): Not Receiving assistance group – the AI assistance prompts were removed from the peer review interface. |
| D3 · Missing outcome data Some concerns about risk of bias | No participant flow or attrition is reported. Exactly 1,625 students and identical per-arm Ns (AI 396, NR 409) appear in both the first-four-weeks appendix and the weeks 5-8 analysis, with no account of students who consented but stopped submitting peer reviews in the second period; per-user outcome measures require at least some submitted reviews, and the paper does not state how (or whether) students with no phase-2 reviews were handled. Review volumes are similar across arms (2,725 AI vs 2,757 NR), which offers indirect reassurance that engagement-related missingness did not differ grossly by arm, but the absence of any missing-data reporting prevents a "Low" judgment. Only consenting students were included, though consent preceded (and was independent of) arm assignment.In the second 4 weeks, 1625 students submitted 11,243 peer reviews on 3573 resources across the 10 courses. Only participants who consented to have their data used in the study were included. |
| D4 · Measurement of the outcome High risk of bias | Outcome ascertainment is automated and identical in mechanism across arms (the system computed flags, GLEU similarity, and SBERT relatedness for all groups, recording flags silently in NR), so classic assessor bias is absent. The decisive problem is circularity: three of the four measures in this result (flag rate, similarity, relatedness) are computed by the very quality-control algorithms that constitute the intervention. Students in the AI arm were prompted, at submission time, to revise any comment those same algorithms flagged, and could iterate until the flag cleared; NR students received no such opportunity. The AI arm's advantage on these metrics is therefore partly definitional - the intervention optimizes the outcome measure directly - making the measurement of the construct "feedback quality/agency" systematically favourable to the AI arm. Comment length is also directly targeted by the detail-oriented prompts. Notably, the two measures independent of the intervention's algorithms (time spent, like/helpfulness rate) showed no group differences, underscoring that the observed effects are concentrated in intervention-defined metrics. This is an outcome-measurement bias differential by arm, warranting "High".Rate of Flagged Reviews serves as an indicator of the extent to which students’ comments require revision based on the quality control functions. When presented with these prompts, students have the option of acting on their own to either edit their feedback independently in response to the suggestions or confirm that their input is appropriate as is. |
| D5 · Selection of the reported result Some concerns about risk of bias | No preregistration, registered protocol, or prespecified analysis plan is mentioned anywhere in the paper. The dependent variables were multiple (six measures), and the analysis involved multiple pairwise arm contrasts; without a prespecified plan it cannot be ruled out that measures or analyses were selected after seeing the data. Mitigating this, all six prespecified-seeming measures are reported in full for every contrast, including null results (time, like rate), with complete statistics (means, SDs, CIs, t, Cohen's d, Tukey p), so there is no visible evidence of selective reporting - but the absence of a plan keeps this at "Some concerns" rather than "Low".The dependent variables were the quality of the comments or feedback measured by different metrics, including the rate of flagged reviews (comments needed revision), similarity (rate of generic comments), relatedness (comment-resource pair), comment length, time spent on review, and likes/helpfulness. |
Funding and conflicts
- Funding
- Australian Government, Australian Research Council Industrial Transformation Training Centre for Information Resilience (CIRES), project IC200100022 (partial support)
- Vendor funded
- No
- Vendor
- None identified
- Notes
- No conflict-of-interest statement observed in the article. The AI-assistance functions were built by the research team into RiPPLE, an adaptive platform associated with author Khosravi's group at UQ, so the evaluated tool is researcher-developed. Data availability: "The authors do not have permission to share data."
Funding is shown on every study and never used to score it.
The source, as retrieved
Abstract
No abstract retrieved.
Where this record came from
| Source | Retrieved | Identifier |
|---|---|---|
| seed | Sep 18, 2026 | link first ingestion |