AI Learning Evidence

Not guidance. A record of published research on AI-assisted learning, and our assessments of it. The limits

Effective and Scalable Math Support: Experimental Evidence on the Impact of an AI Math Tutor in Ghana

Design
Cluster-randomized trial · randomized at school level · 477 participants
Subject and task
Math
Independent practice of numeracy and algebra micro-lessons (Global Proficiency Framework curriculum) via a WhatsApp chat tutor during supervised study hall
Exposure: Two 30-minute sessions per week during study hall, approximately 8 months (early February to late August 2023)
Population match
Secondary students: Partial · University students: Different
Attrition
160 of 637 baseline students (25.1%) did not complete the endline, attributed primarily to inconsistent school attendance. Dropouts had lower baseline scores (M=18.53, SD=7.70) than completers (M=22.26, SD=7.57); no significant differences by age or gender. Authors note the growth-score analysis uses completers only.
Study tier
Study tier 3Randomization was at the school level with only 11 clusters (5 treatment, 6 control), but all analyses (independent-samples t-tests on student-level growth scores) ignore clustering, so the reported p < 0.001 is likely overstated. Attrition was 25% and differential on baseline ability, the outcome was a researcher-developed 35-item test with acknowledged ceiling effects, and no CI or preregistration is reported. The unresolved discrepancies between the arXiv landing-page abstract (~1,000 students, grades 3-9, d=0.37) and the published/v2 text (~500, grades 3-8 analyzed, d=0.36) further reduce confidence in the exact reported quantities, though the direction of effect is consistent across versions. The authors themselves label the study a preliminary evaluation of year 1.
Assessment
Version 2 · AI: two passes agreed · Sep 18, 2026

Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers.

Eleven schools in the Rising Academies network in Ghana were randomly split so that students in five schools spent two 30-minute study-hall sessions per week chatting with Rori, an AI math tutor on WhatsApp, for about 8 months, while six schools continued normal schooling. Among the 477 students who took both the baseline and endline of a 35-question math test, the Rori group gained about 3 points more (Cohen's d = 0.36, p < 0.001). The study is promising but preliminary: schools, not students, were randomized yet the analysis treats students as independent, a quarter of students dropped out, and the test had ceiling effects. Reported figures also differ between the arXiv landing-page abstract and the published version.

01Findings

What the study measured

Arms

ArmWhat it got
A · Rori AI tutor supplement (two 30-minute weekly study-hall sessions)
n = 236
Intervention · teacher supervised
Rori (Rising Academies) · Unspecified; described as using NLP methods including specialized language models (LLMs); cost discussion references an LLM API · Guardrailed tutor · tutor scaffold interface · in class · used Early February to late August 2023
B · Business-as-usual control (regular math instruction, study hall without Rori)
n = 241
Control · teacher led
Continued regular math instruction with identical curricula and classroom hours; no access to Rori during study hall

Findings

A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.

Finding and conditionsResult
F1 · Growth score (endline minus baseline raw score) on a 35-question math assessment covering GPF grade 3-5 numeracy and algebra skills, administered to all grades at both timepoints
Rori AI tutor supplement (two 30-minute weekly study-hall sessions) vs Business-as-usual control (regular math instruction, study hall without Rori)AI availability at assessment not reported — cannot carry a learning claimImmediate post-testNear transferResearcher-developed testPrespecification unclear
Benefit
Cohen's d = 0.36 (Morris pooled-SD pretest-posttest-control formula); growth-score difference 3.01 points (treatment M = 5.13, SD = 7.03 vs control M = 2.12, SD = 6.30); authors equate it to roughly an extra year of learning
not reported
p < 0.001 (independent-samples t-test on student-level growth scores)
Researcher/implementer-developed test aligned to the same Global Proficiency Framework that Rori's curriculum is built on, so the outcome is near-transfer at best. Same 35-item test for grades 3-8 produced ceiling effects at baseline for older students. The paper does not state the assessment conditions explicitly; no AI assistance at testing is implied (control students never had Rori access and treatment access was limited to supervised study-hall sessions) but not reported. The t-test ignores the school-level (11-cluster) randomization, so the significance level is not clustering-adjusted; effect size 0.37 in the arXiv landing-page abstract vs 0.36 in the v2 full text and Springer version.
Limitations
Source-stated: possible Hawthorne effects at treatment schools; possible unobserved baseline differences in prior ability or socio-economic background; same assessment used for all grades produced ceiling effects (some higher-grade students scored perfectly at baseline, so their growth was unobservable and gains may be understated - or Rori may mainly help on easier topics); results cover only year 1; attrition warrants further examination; dose-response not analyzed. Observed: only 11 clusters randomized at school level with no clustering adjustment in the t-test analysis; completer-only analysis with 25% differential attrition; no CI, no preregistration, no funding/COI statement; the AI model powering Rori is not specified. Version discrepancies: arXiv landing-page abstract vs v2 PDF/Springer differ on N (~1,000 vs ~500), grade range (3-9 vs 3-8 in the analyzed sample; v2 body itself states both), effect size (0.37 vs 0.36), and session framing; v2 text vs Table 1 differ on treatment n (236 vs 237).

Who was studied

Level
mixed
Ages
Mean age 12.10 (SD 1.95) among students completing both tests; grades 3-8 per Participation section (grades 3-9 stated in the study overview and arXiv landing-page abstract)
Country
Ghana
Prior knowledge
Baseline mean approximately 20.2/35 on the study assessment in both groups; assessment covered grade 3-5 skills, and some higher-grade students scored at ceiling at baseline
Selection
Students from 11 Rising Academies network schools selected for similarity in geography, demographics, curricula, and teaching methodologies; participation defined by taking the baseline assessment

Methodological notes

Extracted from: https://arxiv.org/pdf/2402.09809 (v2 full text, PDF); https://arxiv.org/abs/2402.09809 (arXiv landing-page abstract and version history); https://link.springer.com/chapter/10.1007/978-3-031-64315-6_34 (Springer AIED 2024 chapter abstract) Access: Values extracted from the arXiv full text (v2 PDF, dated 5 May 2024; 10 pages, read page-by-page from the downloaded PDF). The Springer chapter abstract (AIED 2024, pp. 373-381, published 2 July 2024) was fetched for comparison and matches the v2 PDF abstract on all key figures. The arXiv LANDING-PAGE abstract still carries older (v1-era) figures that conflict with both the v2 PDF and the Springer version; version history shows v1 (15 Feb 2024) and v2 (5 May 2024). The v1 PDF itself was not fetched. Sample-size note: VERSION DISCREPANCIES. (1) N: the arXiv landing-page abstract says "approximately 1,000 students in grades 3-9 across 11 schools", but the v2 full-text PDF and the Springer abstract both say "approximately 500 students"; the full text reports 637 students at baseline (336 treatment, 301 control) and 477 analyzed (text: 241 control, 236 treatment). The ~1,000 figure appears to be a v1-era metadata holdover and appears nowhere in the v2 full text. (2) Grade range: arXiv landing-page abstract says grades 3-9; the v2 full text is internally inconsistent, saying "students in grades 3-9" in the study overview but "637 students in grades 3-8" in the Participation and attrition sections; the Springer abstract states no grade range. (3) Effect size: arXiv landing-page abstract says 0.37; the v2 full text and Springer abstract both say 0.36. (4) Session description: arXiv landing-page abstract says "two 30-minute sessions per week over 8 months"; the v2/Springer abstracts say "one hour a week" with a phone during study hall (Springer adds "low-cost smartphone" and "monitored"); the v2 body says two 30-minute weekly sessions - total dose is consistent (1 hr/week) but the framing differs. (5) Internal inconsistency in v2: text says treatment n=236 but Table 1 says Treatment (N=237). The recorded sample_size of 477 is the analyzed completer sample from the v2 full text.

01bRisk of bias

How much each result can be relied on

Risk of bias is judged per result, not per paper: the same study can report one outcome at low risk and another at high. Two reviewers assess each result independently against a written guide, and every domain judgment below carries its reasoning and the sentences it rests on.

Risk of bias is not the same as whether the numbers work. A trial can be well conducted and still print a standard deviation that is arithmetically impossible, or a table whose sign contradicts its own text. Risk of bias has no question for that, so where it was found it is recorded separately below, under Source integrity.

These words are not quality scores. On the RoB 2 scale, High risk of bias is the worst rating a result can receive and Low risk of bias is the best — the opposite of how the same words read on some other scales. Levels are always written out in full here for that reason.

GROWTH

High risk of bias · RoB 2 · Two reviewers agreed on every domain

How the overall was reached: Worst domain is D3 (High): 25% overall attrition, roughly 30% vs 20% differential loss across arms, and dropouts who scored markedly lower at baseline, with no sensitivity analysis. Independently of the domain judgments, the headline inference is undermined by the mismatch between design and analysis: randomization occurred at the school level across only 11 clusters, yet the reported result is a student-level independent-samples t-test on 477 students ("An independent samples t-test between the control (M = 2.12, SD = 6.30) and the treatment group (M = 5.13, SD = 7.03) revealed that the 3.01 difference in growth scores was highly statistically significant (p < 0.001)."). Ignoring intra-school correlation overstates precision, and with 5 vs 6 clusters any properly clustered analysis would have far wider uncertainty; the p < 0.001 and the claimed statistical significance are therefore not reliable as reported. The 0.36 effect-size point estimate may still be informative, but the trial as analyzed is at high risk of bias for this result.

DomainJudgment and reasoning
D1 · Randomisation process
Some concerns about risk of bias
Assignment was randomized at the school level (5 treatment, 6 control schools), but the paper gives no detail on how the random sequence was generated, whether any stratification or restriction was used, or whether allocation was concealed from those enrolling schools and students. With only 11 clusters, chance imbalance is a real risk and cannot be ruled out by the reported checks: baseline equivalence on scores, age, and gender was tested only among the 477 completers (t(475)), not the full randomized baseline sample of 637, and only at the student level. Schools were purposively pre-selected for similarity before randomization, which mitigates but does not eliminate cluster-level confounding. Timing of baseline assessment relative to assignment (relevant to identification/recruitment bias in cluster trials) is not reported. Observed baseline scores were well balanced among analyzed students, which keeps this at Some concerns rather than High. Note also that the analysis ignores the clustered design entirely (see overall_rationale).
Schools were randomly assigned to two groups: a control group, which included students from six schools; and a treatment group, which included students from five schools.
The assignment was done as the school level, as assignment at the student level was not feasible for a variety of reasons.
D2 · Deviations from intended interventions
Some concerns about risk of bias
The intervention could not be blinded: students, teachers, and school administrators knew which schools received Rori. Treatment students used Rori during supervised study-hall sessions with teachers present to resolve technical issues; control schools ran study hall without Rori. The authors themselves acknowledge that treatment-school staff may have changed their behavior due to increased observation (Hawthorne effect), i.e., deviations from intended conditions that could arise because of the trial context, and offer only a qualitative argument (routine testing culture) against it. Curricula, lesson plans, and timetables were reportedly identical across arms, and no crossover (control access to Rori) is reported, so concerns are moderate rather than severe. The analysis is of completers rather than a formal ITT, but no assignment-non-adherence is reported that would distort the estimate of the effect of assignment.
The control group continued regular math instruction and did not receive access to Rori during study hall.
Meanwhile, the students in the control group did not have access to Rori during their study hall period and continued to receive their regular math instruction.
D3 · Missing outcome data
High risk of bias
160 of 637 baseline students (25.1%) lacked endline data. Attrition was differential by arm: 336 treatment students at baseline vs 236-237 analyzed (about 30% loss) against 301 control at baseline vs 241 analyzed (about 20% loss). Dropouts scored substantially lower at baseline than completers (18.53 vs 22.26 raw points), so missingness is plausibly related to the true outcome (attendance-driven attrition concentrates among weaker students), and it was heavier in the treatment arm. Preferential loss of low-scoring students from the treatment group can inflate the treatment-control growth difference; the authors' claim that growth scores "would not be impacted" does not hold under differential attrition. No sensitivity analyses (e.g., bounds, imputation) were conducted; the authors defer the attrition question to future work.
Initially, 637 students in grades 3-8 participated in the baseline assessment and out of that group, 477 students also completed the endline.
The attrition of 160 students was primarily due to inconsistent school attendance; it is common for students to miss school periodically due to personal reasons.
D4 · Measurement of the outcome
Some concerns about risk of bias
The outcome was an objective 35-item written math assessment, identical for both arms, at both timepoints, and across all grade levels, which limits differential-measurement risk. However, the paper does not say who administered or scored the assessments or under what conditions; administration was presumably by school staff who were aware of assignment, and no proctoring or independent scoring procedures are described (nor whether treatment students could access Rori or phones during testing, though nothing suggests they could). Measurement validity is also weakened by documented ceiling effects: some students, especially in higher grades, scored perfectly at baseline, truncating observable growth. Ceiling truncation applies to both arms and would most plausibly bias toward the null, so this stays at Some concerns.
The same math assessment was used at both timepoints and all grade levels.
The assessment consisted of 35 questions, each worth one point, with a mixture of multiple-choice and open-response questions covering numeracy and algebra skills from grades 3 to 5 on the Global Proficiency Framework.
D5 · Selection of the reported result
Some concerns about risk of bias
No preregistration, trial registration, protocol, or pre-specified analysis plan is mentioned anywhere in the paper, so the reported result cannot be checked against pre-specified intentions. The study reports a single outcome (growth score on one assessment) with a single simple analysis, which limits the visible scope for selective reporting among multiple measurements; but analysis flexibility remains (growth score vs ANCOVA vs difference-in-differences, completers-only sample definition, unclustered vs clustered inference), and there is no way to verify the reported analysis was not selected from several. Quote below documents the analysis as reported; no quote about registration exists because the paper contains no registration or protocol statement.
To assess the impact of using Rori on learning, growth scores were computed by subtracting baseline raw scores, the number of questions answered correctly, from endline raw scores for each student who completed both tests.
02Funding

Funding and conflicts

Funding
Not stated; the paper contains no funding or acknowledgements statement.
Vendor funded
Partial
Vendor
None identified
Notes
Rori is a product of Rising Academies, and the study ran inside Rising Academies' own school network. Co-author Hannah Horne-Robinson is affiliated with Rising Academies (per the author list: Henkel - University of Oxford; Horne-Robinson - Rising Academies; Kozhakhmetova and Lee - J-PAL North America). No conflict-of-interest statement is provided in either version.

Funding is shown on every study and never used to score it.

03Claims

Claims this study bears on

04Source

The source, as retrieved

Abstract

No abstract retrieved.

Where this record came from

SourceRetrievedIdentifier
seedSep 18, 2026link
first ingestion