One in Five EdTech Tools Have Evidence They Work. Here's Why
Researchers argue randomised trials can't keep pace with AI tools that update monthly, leaving schools to buy on marketing rather than proof.
The gold standard has a timing problem
Every school leader evaluating an AI tool eventually asks the same question: where's the evidence it works? For most products, the honest answer is that there isn't much, and a new argument from Brookings suggests the standard method for producing that evidence is quietly breaking down.
In an August 2026 article for Brookings, Stacey Alicea of the Research Partnership for Professional Learning and Meghan McCormick of the Overdeck Family Foundation lay out the problem plainly: only around one in five edtech products carry any evidence that they actually improve learning. That figure isn't new to procurement teams who have sat through vendor pitches heavy on testimonials and light on data. What is new is the authors' explanation for why the gap persists even as AI tools multiply.
Why randomised trials can't keep up
The randomised controlled trial has long been treated as the gold standard for education research: assign similar groups of students to a tool or a control condition, run it for a term or a year, and compare outcomes. The method works when the intervention stays still.
AI tools don't stay still. Vendors push model updates, retune prompts, and change interfaces on a rolling basis, often monthly. Alicea and McCormick note that an RCT depends on "a stable intervention applied consistently across contexts" — an assumption that collapses when the product being tested is a different piece of software by the time the trial reports its findings. A two-year study that begins with one generation of underlying model can conclude after the model, the guardrails, and the interface have all changed twice.
By the time a rigorous trial finishes, the technology it evaluated may no longer exist in the form that was tested.
That is not an argument for abandoning rigour. It is an argument that the field has been applying the wrong kind of rigour to a moving target, and that schools relying on "properly evaluated" as a purchasing filter are applying a test that structurally cannot be passed on any useful timescale.
What the alternative looks like
The authors point to "implementation research and development" as a better fit: shorter, iterative testing cycles that use a tool's own usage data to check whether it changes behaviour in the ways it claims to, before anyone tries to measure whether it moves exam results. Teaching Lab, for example, is running an ongoing A/B test comparing two versions of an AI coaching tool to see which design choices actually shift coaching practice — evidence that is provisional and narrow, but current. The US Institute of Education Sciences has funded four generative AI research centres built around the same iterative model rather than a single terminal trial.
The authors set out three principles for anyone assessing a tool under this approach: build evidence in stages rather than demanding a finished verdict up front; ask "how does this change what people do" before asking "did outcomes improve"; and let the actual question drive the method, rather than defaulting to a large trial because it sounds more credible.
What this means for schools, not just researchers
None of this is abstract for anyone sitting through a procurement cycle this term. Vendors routinely cite "evidence-based" or "research-backed" in marketing copy, and school leaders have had few ways to interrogate what sits behind the claim. If the researchers making the case for AI tools are themselves saying the traditional proof point is often unreachable, that changes what a reasonable question to a vendor looks like.
It also raises the stakes on what schools already control: their own pilot data. A department running a six-week trial of an AI marking or planning tool, tracking teacher time and simple usage measures, is closer to what Alicea and McCormick describe as credible early evidence than waiting for a vendor to produce a study that may already be out of date by the time it lands in an inbox.
What to do
Ask vendors for implementation data, not just outcome claims: what changed in the product over the past six months, and what did their own usage telemetry show. Treat a lack of large-scale trial evidence as normal for AI tools rather than disqualifying, but weight it against small, current, honestly reported pilot data. Run short internal trials before wide rollout, and record simple behavioural measures — time saved, tasks completed, features actually used — rather than waiting to measure attainment change.
What to watch
Whether IES-funded research centres and similar rapid-cycle models produce evidence that regulators and school inspectorates are willing to accept as a substitute for RCT findings, and whether procurement frameworks in England and elsewhere start asking for iterative evidence rather than treating its absence as a red flag.
Sources: AI is rapidly changing education and research needs to keep up (Brookings)