The Only Randomised Trial on Classroom AI Says 25 Minutes a Week
An EEF-funded RCT found ChatGPT cut teacher lesson-planning time by 31 percent with no quality loss — modest, real evidence in a market that mostly runs on none.
A rare thing: an actual control group
Most claims about AI saving teachers time come from surveys, vendor case studies or self-reported estimates. The Education Endowment Foundation's Teacher Choices trial, evaluated independently by the National Foundation for Educational Research (NFER), is different: it is a randomised controlled trial, run with 259 teachers across 68 secondary schools in England.
Half the teachers — 129 of them, across 34 schools — were given access to ChatGPT plus a short implementation guide and asked to use it for lesson and resource preparation in Year 7 and 8 science. The other half planned as normal. Time spent was logged, and a sample of the resulting materials was independently assessed for quality.
The result: teachers using ChatGPT spent 56.2 minutes a week on lesson and resource planning, against 81.5 minutes in the comparison group. That is a saving of 25.3 minutes a week, a 31 percent reduction. Crucially, NFER's quality assessment found no noticeable difference between the materials produced by the two groups. Teachers mostly used the tool to generate quiz and activity ideas and to tailor existing resources to specific classes, not to write lessons from scratch.
Twenty-five minutes a week is not nothing, but it is also a long way from the "six weeks a year" and "5 to 10 hours a week" figures now circulating in vendor marketing and some US foundation-backed surveys. Those numbers come from self-report, not from a control group. This trial is one of the few pieces of evidence in the sector that was actually designed to rule out the possibility that teachers who adopt AI are simply the ones who were already faster, or that they are unconsciously inflating their own estimates.
Why the gap matters more than the number
The interesting story here is not really the 31 percent. It is the contrast between this trial and how the wider market is behaving. In the United States, spending on AI in education totalled roughly $2.5 billion last year and is forecast to exceed $15 billion by 2033, according to reporting by Stateline and the Baltimore Sun this month. District leaders quoted in that reporting describe an "asymmetry" between what vendors know about their products and what schools can verify before buying — echoing Mark Schneider of the American Enterprise Institute, who told Stateline that gap is baked into how the market operates.
Put the two stories side by side and the picture is stark: the evidence base that actually meets a recognised bar — pre-registration, randomisation, independent evaluation — points to modest, believable, single-digit-percent gains in a narrow task. The commercial market, meanwhile, is scaling on claims that mostly haven't been tested that way at all.
The EEF trial did not find that AI transforms teaching. It found that, used narrowly and with guidance, it saves some time on a specific task without making the output worse. That is a useful, boring, credible result — which is precisely why it gets drowned out.
What the trial doesn't tell you
It's worth being honest about the limits too. This was one subject (science), two year groups (7 and 8), one tool (ChatGPT), and one task (lesson and resource preparation) — not marking, not differentiation across a whole timetable, not SEND provision, not behaviour management. A 31 percent reduction in a 56 to 81 minute weekly task is not the same as reclaiming hours across a teacher's full workload. NFER's own writeup is careful to frame this as evidence about a specific, bounded use case, not a verdict on classroom AI in general.
That specificity is a feature, not a flaw. It is exactly the kind of claim school leaders should be asking vendors and consultants to produce before signing a procurement contract: what task, what population, what comparison group, what was actually measured.
What to do
- Treat this trial as a floor, not a ceiling, on plausible time savings — and be sceptical of any pitch quoting multi-hour weekly gains without a comparable study behind it.
- When evaluating a new AI tool for planning or resourcing, ask what task it was tested on and whether output quality was independently checked, not just time saved.
- Encourage staff piloting AI for planning to log actual time spent for two or three weeks before and after, rather than relying on impression alone.
What to watch
The EEF and NFER have signalled more work in this space is likely, given the interest the trial has generated. Watch for follow-up trials that test marking, differentiation or SEND-related tasks — areas where the workload case for AI is asserted most loudly but tested least. Until those exist, the science lesson-planning trial remains one of the only pieces of solid ground in an otherwise evidence-thin market.
Sources: Education Endowment Foundation and NFER Teacher Choices trial report; Stateline and Baltimore Sun reporting on school AI procurement, August 2026.