From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria
- Design
- Randomized trial · randomized at student level · 1328 participants
- Subject and task
- Writing language
Curriculum-aligned English language practice via dialogue with an LLM chatbot acting as a virtual tutor (students in pairs, teacher-provided starter prompts)
Exposure: 6 weeks (June-July 2024; up to twelve 90-minute after-school sessions, max two per week; mean attendance ~72% of sessions) - Population match
- Secondary students: Direct · University students: Different
- Attrition
- Substantial and differential: 422/657 (~64%) of treatment and 337/671 (~50%) of control completed the endline; the paper states the treatment-control difference in attrition is significant and addresses it with Lee bounds (effects remain positive and significant; e.g. total weighted score bounds 0.255-0.327) and inverse-probability weighting (estimates largely unchanged)
- Study tier
- PreliminaryIndividually randomized RCT (lottery among volunteer students within 9 schools) with a business-as-usual control, school fixed effects, baseline-score control, and robust standard errors; because randomization is at the student level there is no cluster-inference problem, but the design invites spillovers and documented control-group contamination (some control students attended sessions). The sample is self-selected (52% of eligible students volunteered) and endline attrition is large and differential (64% vs 50%), though Lee bounds and IPW checks hold. Effects are precisely estimated (SEs ~0.07). Rated PRELIMINARY because this is a non-peer-reviewed Policy Research Working Paper series entry authored by the implementing World Bank team, notwithstanding the verified reproducibility package.
- Assessment
- Version 2 · AI: two passes agreed · Sep 18, 2026
Study facts come from the paper. Population match and study tier are our judgments, made separately for the two populations this collection covers. This item has not been peer reviewed.
In Benin City, Nigeria, volunteer first-year senior secondary students were randomly chosen to attend up to twelve 90-minute after-school sessions over six weeks in which pairs of students chatted with Microsoft Copilot (GPT-4) as an English tutor, guided by trained teachers using curated prompts. On a pencil-and-paper endline test, selected students scored about 0.31 SD higher overall and 0.24 SD higher in English, and also did about 0.21 SD better on their regular third-term school English exam. Gains were larger for girls, for students with higher baseline scores, and for higher-SES students, and grew roughly linearly with days attended. The authors rate the program among the most cost-effective learning interventions, though the study is a non-peer-reviewed working paper with volunteer participants and substantial differential attrition.
What the study measured
Arms
| Arm | What it got |
|---|---|
| A · AI tutoring: after-school Copilot sessions n = 657 | Intervention · teacher supervised Microsoft Copilot · GPT-4 · Prompted pedagogy · chat interface · in class · used June 3 - July 11, 2024 (six-week pilot) |
| B · Control: business-as-usual schooling, no intervention n = 671 | Control · teacher not reported No after-school program; students continued their regular classroom learning only |
Findings
A finding is one outcome contrast: two arms, one measure, one set of conditions. A paper is never "positive" or "negative" here; each finding carries its own direction and its own conditions. The badge that matters most is the first one: whether the AI was still available when the outcome was measured. Why that gates everything.
| Finding and conditions | Result |
|---|---|
| F1 · Total score on endline assessment, weighted by expert-rated item difficulty (English + digital skills + AI knowledge) AI tutoring: after-school Copilot sessions vs Control: business-as-usual schooling, no interventionMeasured without AIImmediate post-testSame contentResearcher-developed testPrespecification unclear | Benefit ITT 0.310 SD (SE 0.068) not reported (SE given); Lee bounds 0.255-0.327, bounds 95% CI [0.115, 0.462] p<0.01 Multiple-choice test designed by experts on the Nigerian curriculum, multiple randomized-order versions, monitored pencil-and-paper administration (7/11-7/12/24, at intervention end); composite includes AI-knowledge and digital-skills items on content the treatment group was plausibly differentially exposed to |
| F10 · Dose-response: effect per additional day of attendance on total weighted endline score (attendance instrumented by treatment assignment) AI tutoring: after-school Copilot sessions vs Control: business-as-usual schooling, no interventionMeasured without AIImmediate post-testSame contentResearcher-developed testPrespecification unclear | Benefit IV/2SLS 0.033 SD per day attended (SE 0.007); 0.031 SD/day under conservative assumption of zero effect for the 3.5% never-attending non-compliers; extrapolated 1.2-2.23 SD for a full 36-week school year depending on attendance assumptions not reported (SE given) p<0.01 Mean attendance 72% in the treatment group; linear fit supported (quadratic term n.s.); full-year projections are model-based extrapolations requiring constant-effect and linearity assumptions, not experimental estimates |
| F2 · Total score on endline assessment, IRT-scaled AI tutoring: after-school Copilot sessions vs Control: business-as-usual schooling, no interventionMeasured without AIImmediate post-testSame contentResearcher-developed testPrespecification unclear | Benefit ITT 0.263 SD (SE 0.068) not reported (SE given) p<0.01 Same instrument as F1 rescaled by Item Response Theory using observed item difficulty; robustness alternative to the ex-ante weighted score |
| F3 · English skills section of endline assessment (stated main outcome of interest) AI tutoring: after-school Copilot sessions vs Control: business-as-usual schooling, no interventionMeasured without AIImmediate post-testSame contentResearcher-developed testPrespecification unclear | Benefit ITT 0.238 SD (SE 0.068) not reported (SE given); Lee bounds 0.202-0.238, bounds 95% CI [0.058, 0.384] p<0.01 Majority of test items; English topics aligned with the Nigerian first-year curriculum covered during the intervention period |
| F4 · Digital skills section of endline assessment AI tutoring: after-school Copilot sessions vs Control: business-as-usual schooling, no interventionMeasured without AIImmediate post-testNear transferResearcher-developed testPrespecification unclear | Benefit ITT 0.139 SD (SE 0.076) not reported (SE given); Lee bounds 0.138-0.145, bounds 95% CI [-0.018, 0.297] p<0.10 (marginal; significant at 10% level only) Basic digital-concepts knowledge; not a primary target of the program (paper interprets it as a spillover), and treatment students had differential hands-on computer exposure |
| F5 · AI knowledge section of endline assessment AI tutoring: after-school Copilot sessions vs Control: business-as-usual schooling, no interventionMeasured without AIImmediate post-testNear transferResearcher-developed testPrespecification unclear | Benefit ITT 0.309 SD (SE 0.077) not reported (SE given); Lee bounds 0.232-0.351, bounds 95% CI [0.078, 0.489] p<0.01 AI-knowledge items favor the treatment group by construction: the first session explicitly taught AI benefits/risks and sessions built AI familiarity; largest sub-score effect |
| F6 · Regular third-term English curricular exam (school-administered) AI tutoring: after-school Copilot sessions vs Control: business-as-usual schooling, no interventionAI availability at assessment not reported — cannot carry a learning claimImmediate post-testNear transferCourse examPrespecification unclear | Benefit ITT 0.206 SD (SE 0.067) not reported (SE given) p<0.01 Independent school exam covering the entire term's content beyond the six-week program (timeline: 7/12/24, one day after sessions ended, hence immediate_post); note internal inconsistency - the text says exam content was broader than and included intervention material, while Table 2's note calls it unrelated to the intervention's material; AI availability at this exam not stated, though it is a regular curricular exam |
| F7 · Heterogeneity by gender: Treatment x Female interaction on total weighted endline score AI tutoring: after-school Copilot sessions vs Control: business-as-usual schooling, no interventionMeasured without AIImmediate post-testSame contentResearcher-developed testPrespecification unclear | Benefit interaction 0.420 SD (SE 0.188); main treatment effect in this specification -0.039 (SE 0.172), n.s. not reported (SE given) p<0.05 (interaction) Larger effect for female students; authors caution the result appears influenced by inclusion of a girls-only school that performed worse than others at baseline |
| F8 · Heterogeneity by baseline performance: Treatment x Second Term Exam score interaction on total weighted endline score AI tutoring: after-school Copilot sessions vs Control: business-as-usual schooling, no interventionMeasured without AIImmediate post-testSame contentResearcher-developed testPrespecification unclear | Benefit interaction 0.151 (SE 0.072); main treatment effect 0.311 (SE 0.068) not reported (SE given) p<0.05 (interaction) Higher-baseline students benefited more; quantile regressions nonetheless show positive significant effects across all quantiles of the outcome distribution |
| F9 · Heterogeneity by socioeconomic status: Treatment x SES index interaction on total weighted endline score AI tutoring: after-school Copilot sessions vs Control: business-as-usual schooling, no interventionMeasured without AIImmediate post-testSame contentResearcher-developed testPrespecification unclear | Benefit interaction 0.113 (SE 0.047); main treatment effect 0.305 (SE 0.071) not reported (SE given) p<0.05 (interaction) SES index from PCA of household goods, internet access, study space, parental education; higher-SES students benefited more, consistent with anecdotal reports that poorer students were often first-time computer users |
- Limitations
- Source-stated: student-level randomization may cause spillovers to controls; some control students inadvertently gained access to sessions (teachers unwilling to enforce the distinction early on), and early implementation problems (account creation, internet disruptions, power outages) attenuate estimates; differential attrition at endline (addressed with Lee bounds and IPW); volunteer sample with no demographic data on non-interested students, limiting representativeness checks; the female heterogeneity result may be driven by one girls-only school with low baseline performance; urban schools selected for having computer labs, limiting external validity (rural settings untested); short 6-week duration with no long-term follow-up; year-long extrapolations (1.2-2.2 SD) rest on strong linearity and constant-effect assumptions. Observed: the endline instrument is researcher-commissioned and partly measures content (AI knowledge, digital skills) plausibly taught only to the treatment group; the paper is internally inconsistent about the third-term exam, describing its content in the text as broader than but including intervention-period material while Table 2's note calls its content "unrelated to the intervention's material"; treatment group also received extra instructional time and adult supervision, so the AI component is not isolated (authors acknowledge and argue against pure time effects); no pre-registration or pre-analysis plan is mentioned.
Who was studied
- Level
- secondary
- Ages
- first-year senior secondary (SS1), typically 15 years old
- Country
- Nigeria (Benin City, Edo State)
- Prior knowledge
- regular first-year English curriculum students; baseline measured by first- and second-term curricular exam scores; for poorer students often first experience using computers
- Selection
- 9 public schools selected for computer-lab availability; all SS1 students invited, 52% expressed interest within a 10-day window; randomization by lottery only among interested volunteers; guardians signed consent
Methodological notes
Extracted from: https://documents1.worldbank.org/curated/en/099548105192529324/pdf/IDU-c09f40d8-9ff8-42dc-b315-591157499be7.pdf Access: full text (official World Bank PDF downloaded and text-extracted locally; version note on title page says originally published May 2025, this version updated December 2025) Sample-size note: 1,328 randomized (657 treatment, 671 control); 759 completed the final assessment (422 treatment, 337 control); regression Ns 636-654 depending on outcome (654 for endline assessment scores, 636 for third-term exam)
How much each result can be relied on
Risk of bias is judged per result, not per paper: the same study can report one outcome at low risk and another at high. Two reviewers assess each result independently against a written guide, and every domain judgment below carries its reasoning and the sentences it rests on.
Risk of bias is not the same as whether the numbers work. A trial can be well conducted and still print a standard deviation that is arithmetically impossible, or a table whose sign contradicts its own text. Risk of bias has no question for that, so where it was found it is recorded separately below, under Source integrity.
These words are not quality scores. On the RoB 2 scale, High risk of bias is the worst rating a result can receive and Low risk of bias is the best — the opposite of how the same words read on some other scales. Levels are always written out in full here for that reason.
Assessed, not settled (1)
Recorded rather than omitted. Leaving these out would make the appraisal look cleaner than it is, and a disagreement between two careful readers is itself worth knowing.
ENDLINE
Assessed, not settled · RoB 2 · Two reviewers reached different judgments and nothing has settled it
This result has no overall judgment, because the two independent reviews did not reach one. Both readings are shown. The domains they did agree on are below.
| Where they differ | The two readings |
|---|---|
| D3 · Missing outcome data | One reviewer: Some concerns about risk of bias The other: High risk of bias |
| Domain | Judgment and reasoning |
|---|---|
| D1 · Randomisation process Low risk of bias | Individual-level lottery among students who volunteered after an expression-of-interest window, conducted by computerized simple random sampling without replacement and without stratification. Allocation results were documented and saved (though no fixed random seed was used). Table 1 shows balance on gender, age, SES index, and school; baseline first- and second-term exam score differences (0.131 SD and 0.096 SD, both SE = 0.073) favor treatment but are statistically insignificant, and the authors conservatively control for the second-term exam in all models. Volunteering precedes randomization, so it affects generalizability, not internal validity. No indication of allocation problems.Once the period to express interest closed, the randomization was carried out using simple random sampling without replacement among interested students to assign them either to the treatment group, which participated in the program, or to the control group, which did not receive any intervention but continued their regular learning in the classroom. These results confirm that the randomization process achieved balance in key characteristics, supporting the validity of subsequent comparisons between the treatment and control groups. |
| D2 · Deviations from intended interventions Some concerns about risk of bias | Participants and implementing teachers were necessarily aware of assignment, and deviations from the intended conditions arose in the trial context: monitoring data show some control students inadvertently attended the after-school sessions because some teachers did not enforce the distinction, especially in the early weeks; there were also early implementation problems with account creation and engagement in the treatment arm. The analysis is an appropriate ITT estimate of the effect of assignment (with a LATE/dose-response analysis reported separately; treatment-group attendance averaged 72%, non-compliers only 3.5%). Control-group crossover would most plausibly attenuate the ITT estimate toward zero rather than inflate it, so the effect on the reported positive result is likely conservative, but deviations did occur and their extent is not quantified, so Low cannot be assigned.Second, monitoring and evaluation data indicate that some control group students inadvertently gained access to the after-school sessions, due to the lack of willingness of some teachers to enforce the distinction, especially during the first weeks. All the results presented thus far are ITT estimates, which are based on an average attendance rate of approximately 72% among participants in the treatment group. |
| D3 · Missing outcome data Some concerns about risk of bias | Severe missing outcome data: of 1,328 randomized (657 treatment, 671 control), only 759 completed the final assessment (422 treatment = 64.2%; 337 control = 50.2%), i.e., about 43% of the randomized sample was not assessed, with a large (~14 percentage point) and statistically significant differential by arm — plausibly related to the outcome, since assignment itself likely motivated endline attendance and the paper shows attendance correlates with prior performance. The authors address this with Lee (2009) bounds tightened by school dummies, under which the ITT effects remain positive and significant for the total weighted score (bounds 0.255-0.327, 95% CI [0.115, 0.462]) and English skills (0.202-0.238, 95% CI [0.058, 0.384]), and with inverse-probability-weighted estimates (probit on second-term exam, sex, school) that leave the ITT effects largely unchanged (total weighted 0.272; English 0.194). These sensitivity analyses provide meaningful evidence that the qualitative result is not an artifact of attrition, but Lee bounds rest on a monotonicity assumption and IPW on missingness-at-random given a small set of observables; with attrition of this magnitude and significance, bias in the point estimates from outcome-dependent missingness cannot be fully excluded, so the domain cannot be rated Low, while the robust bounded results argue against High.However, only 422 students in the treatment group and 337 in the control group completed the final assessment, which constitutes the final sample used for the analysis. Finally, since the difference in attrition between the treatment and control groups is significant, we first provide Lee-bounds estimates of ITT effects on the outcome variables. |
| D4 · Measurement of the outcome Some concerns about risk of bias | The endline was an objective, multiple-choice pencil-and-paper assessment designed by experts based on the Nigerian curriculum, administered identically to both arms with monitors stationed at schools and multiple randomized-order versions to limit cheating; scoring of multiple-choice items leaves little room for assessor influence even though blinding of scorers is not described, and students' awareness of assignment is unlikely to bias an objective proficiency test materially. The concern is alignment of the instrument to the intervention: the total weighted score — one of the two outcomes in this result — includes AI-knowledge and digital-skills sections covering content that only the treatment group was systematically exposed to during sessions, so that composite partly measures exposure to the intervention's own materials and could overstate learning gains relative to a curriculum-only test (treatment effects were largest on AI skills, 0.309 SD). The English score, the stated main outcome, is curriculum-aligned for both groups and is corroborated by the independently administered third-term curricular exam (0.206 SD), which mitigates but does not remove the concern for the composite.The assessment was administered in a traditional pencil-and-paper format and consisted of multiple-choice questions designed by experts based on the Nigerian curriculum. To minimize the risk of cheating, multiple versions of the assessment were created, each with a randomized order of questions. |
| D5 · Selection of the reported result Some concerns about risk of bias | No trial registration, pre-analysis plan, or pre-specified analysis intentions are mentioned anywhere in the paper, so there is no basis to confirm that the reported result was selected from analyses specified before unblinded outcome data were available. Multiple outcome constructions existed (unweighted, expert-weighted, and IRT-scaled scores; domain subscores; third-term exam) and multiple analysis choices (baseline control, school fixed effects), creating scope for selective reporting, although the paper transparently reports the full battery, the estimates are consistent across specifications (0.21-0.31 SD, all significant), and a verified reproducibility package accompanies the paper (which supports computational reproducibility but is not a prospective pre-specification). Under RoB 2, absence of a pre-specified plan with multiple eligible measurements yields Some concerns; there is no positive evidence of result-selection that would warrant High.A verified reproducibility package for this paper is available at http://reproducibility.worldbank.org, click here for direct access. Furthermore, we estimated the models using alternative specifications of the dependent variable to assess the robustness of our findings. |
Funding and conflicts
- Funding
- Mastercard Foundation (financial support acknowledged); study conducted and authored by World Bank staff (Education Global Department) in collaboration with Edo State education authorities
- Vendor funded
- No
- Vendor
- None identified
- Notes
- The tutoring tool was Microsoft Copilot powered by GPT-4, but the paper reports no funding, involvement, or role of Microsoft or OpenAI; the tool was used as a free off-the-shelf product. Funder is the Mastercard Foundation (philanthropy). Authors are World Bank employees evaluating a pilot their own institution helped implement (evaluator-implementer overlap, though not commercial). A verified reproducibility package is available at reproducibility.worldbank.org.
Funding is shown on every study and never used to score it.
Claims this study bears on
- Unrestricted use of a general-purpose LLM chatbot during practice improves secondary students' subsequent performance on assessments taken without AI.No direct evidence · whole-study judgment
Positive unassisted endline, but teacher-supervised prompted-pedagogy Copilot use is not unrestricted chatbot use (intervention-class lint).
- Guardrailed AI tutors (hints, no direct answers, pedagogical prompting) improve secondary students' performance on assessments taken without AI, compared with business-as-usual instruction.Supports · rests on finding F3
English endline (pencil-and-paper, no AI): +0.238 SD for the supervised after-school Copilot program vs no program. Caveats: the comparator bundles no extra instruction time (the program adds it); PRELIMINARY working paper; attrition handling assessed-not-settled (D3).
- Secondary students with low prior knowledge gain less, or are harmed more, by AI assistance during practice than high-prior-knowledge peers.Supports · rests on finding F8
Treatment-by-baseline interaction +0.151 SD (p<0.05) on the unassisted endline: higher-baseline students gained more. Weak support: PRELIMINARY working paper; attrition assessed-not-settled.
The source, as retrieved
Abstract
No abstract retrieved.
Where this record came from
| Source | Retrieved | Identifier |
|---|---|---|
| seed | Sep 18, 2026 | link first ingestion |