Research Review: AI+Human Tutors Match Quality of Human-Only Tutors?
What is actually going on here?
Hey folks - this is where I write on special Wednesdays about edtech and math education. Today I’m going below the surface of some recent high-profile research that investigated how AI might offer students more help in math.
The Headline
The Reality
The Study
AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms
What They Did
The researchers embedded three kinds of support inside Eedi, a math software platform1.
Support #1: Static pre-written hints.
Support #2: A chat interface for accessing support from a human.
Support #3: A chat interface for accessing support from an LLM, with a human approving the messages.
The study was nicely controlled. Every condition got tested against the other. Support #3 is the focus of the study. “In sessions with LearnLM, a supervising tutor reviewed each message that LearnLM drafted. They could either edit the message, completely re-write it, or approve it without any changes.”
Their outcome measures were whether students corrected a mistake, resolved a misconception, or transferred their knowledge. (See p. 9 for more precise descriptions of those outcome measures.)
Can I See a Picture?
What Happened?
Both of the dynamic supports clobbered the pre-written hints, as you’d probably expect.
There wasn’t much difference between “human alone” and “human + AI” on mistake remediation or misconception resolution. The big headline: the human + AI saw a 5.5% increase over a human alone on knowledge transfer. The study reports the odds that this is a legit finding at 1.3.
Finally, we directly compare the two tutoring conditions. We estimate that receiving support from LearnLM improved a student’s odds of success by a factor of 1.3 [0.9, 1.7] relative to human tutors, corresponding to an ATE of +5.5% [–1.4%, +12.4%]. Based on this posterior distribution, we find a strong probability (93.6%) that LearnLM elicited greater knowledge transfer than human tutors alone.
If, like me, you aren’t particularly familiar with the Bayesian methods here (the odds ratio) and if you are likewise concerned that the confidence interval contains 0 and the odds intervals contains 1, you can read this response from one of the study contributors, who seems to say, “Yeah, the odds aren’t 100% but they’re a lot better than chance.”
The study stored the original AI responses and the edits from human tutors (a fun source of data) which led to the finding that 76% of the LLM responses were accepted in full by the human tutor.
LearnLM proved to be a trustworthy source of pedagogical instruction, with the supervising tutors approving over 76% of its messages without changes or with only minimal edits (changing one or two characters; e.g., deleting an emoji).
Reasons to Be Happy
It’s a responsible study. Humans supervising the LLM. Lower risk than just setting the LLMs onto the kids directly.
“No differences” is still good. I think we can chalk something of a win for AI even when there isn’t a significant difference between the human tutor and human + LLM. The LLM operates under different constraints than a human. It can work faster and support more students simultaneously. The mere fact that there wasn’t a degradation in student learning is reason enough to be happy.
Reasons to Be Real
It’s a small practical effect in a short intervention. The total time here was seven hours—one hour per week over seven weeks—for an increase of 5.5% on some proprietary measures. I would put this study in the basket of “point solutions that don’t congeal in any obvious way into a larger plan for school improvement.” Point solutions aren’t bad but we need to be real about the scope of the intervention here.
This study is trying to do the wrong thing the right way. If the best tutoring interventions have kids a) working with trusted community members (parents, paraprofessionals, etc) b) in person for c) multiple times per week, d) doing work that is aligned to their curriculum, this intervention has none of those hallmarks. In general, getting tutored by a human through a text messaging interface shaves off a huge amount of the value of getting tutored by a human while also playing to a text chatbot’s strengths. It’s as if we asked Michael Bublé to play the kazoo and then supervise a robot purpose-built to play a kazoo. I’m not sure we’d be surprised by the results. I’m also not sure we’d feel good about saying “AI+Michael Bublé Match Quality Of Michael Bublé Alone.”
Reasons for Surprise
It’s honestly quite strange to me that the Human + LLM showed its strengths with “Knowledge Transfer.” I would have expected to see the greatest difference on follow-up tasks that were closer to the tutoring moment. Wouldn’t you? Kids didn’t fix their answer on the next try (“mistake remediation”) or the try after that (“misconception resolution,” I guess) at different rates between the human and human+AI conditions. But apparently if you look beyond those moments, to a question in a different unit the next week, the effect of the AI really kicks in?
This isn’t what’s commonly called “Knowledge Transfer.” Doing the next topic is doing the next topic. Knowledge transfer is typically understood as taking your learning on a topic and transferring it to a different context. i.e. “Okay you learned about slope in the context of an airplane landing. Can you calculate it in the context of the unit price at a grocery store?”
I’m still surprised by this. If I were cynical, I’d wonder if the researchers found no significant differences on those first two much more intuitive measures and then went looking around at other measures, a frowned-upon process known as p-hacking. But thankfully I am not a cynic—a sweetie, in point of fact—and I will just process my cognitive disequilibrium on my own.
Implications
In these kinds of text-based tutoring environments, it seems possible AI could do some good or at least not do harm. Even if the LLM did fine in 76% of exchanges, though, I suspect there is some real frustration in the long tail of the distribution. For example, one human tutor said they had to overrule the LLM’s tendency to question students to death:
Tutors often found it necessary to step in when LearnLM’s Socratic questions, while pedagogically sound, persisted longer than a student’s patience. One tutor described a common scenario where “[LearnLM] will go, ‘Okay, you’ve got the answer. Let’s dig a little deeper about why you’ve got that answer.’ And the child is just like, ‘No, I’ve got it. I know what I’m doing. Can I go now?”’ (T1).
If the questions persist past their usefulness, I don’t think we can call them “pedagogically sound.”
Odds & Ends
¶ Over on LinkedIn, Shelbi Cole posted the image above and wrote:
The only reliable way to create strategic groups for Tiers 2 & 3 math instruction is to understand what children are doing.
The kid gets one out of six problems correct for a score of 17%. But “one out of six problems correct” is a pretty limited understanding of the child’s understanding and actively unhelpful for prescribing an intervention for the child. That’s a challenge for digital systems that reduce student thinking down to a binary correct / incorrect mark and that don’t make room for student handwriting.
¶ I was a guest on the last EdTechnical podcast of 2025. Libby Hills and Owen Henkel have a great vibe, know the space backward and forward, and keep a brisk conversation. We chatted about discovery learning v. explicit instruction and about AI Guys, how they’re made and how they limit their own effectiveness in edtech.
¶ Dylan Kane describes High Accountability Teaching in a post I liked quite a lot. Here’s the tease:
The core of high accountability teaching is to get as many of the swingers on board as possible.
Lately, I think a lot about “mattering” in classrooms. Do kids feel like they matter in their class? Do they feel like their ideas matter? Do they feel like the work they are doing matters? When kids complete work in class that a) isn’t seen by other kids, b) isn’t seen by grownups, c) isn’t commented on, d) doesn’t count for a grade, when they start to wonder, “Does this work actually matter to anyone?” that spins up a pretty negative cycle for learning. Dylan’s post added a lot to my thinking here.
¶ University of Washington freshman Luran Yang wrote a beautiful op-ed in the Seattle Times responding to the near-constant use of AI on her campus. “Not all hope is lost, however — I’ve been in one class, and only one class, with a near-zero rate of AI usage.”
¶ Education Week reports on Austin ISD’s use of artificial intelligence to support the annual creation of their master schedule, a task that historically requires lots of manual shuffling of teachers and classrooms and students.
… it saved [principals and their scheduling team] at least 50 hours. And the feedback that we got from those schedulers was that it increased their job satisfaction. They were so excited they got their summer back. They felt like they could actually manage master scheduling, and the master scheduling wasn’t managing them.
I don’t hold up “time savings” as a first-order goal for teachers, administrators, or schools. There are plenty of ways to save time that result in a worse education for kids. But I am much more optimistic that AI can support back-office operational tasks like this rather than the in-classroom relational work of teaching.






Cool:
"Because we could not cleanly isolate tutor throughput in the main trial, we conducted a supplementary operational simulation with several of the tutors. Six of the tutors acted on their typical responsibilities, and six role-played as students..."
"Tutors took longer to complete the average supervised session (5.1 minutes) than they did to complete the average session on their own (3.9 minutes)."
That feels like it might be relevant to interpreting the results. Maybe longer sessions help?
I love the Michael Bublé example LOL