Skip to main content
ExplainerAI TutoringExplainer· 4 min read· in Education

Why AI Tutors Like Khanmigo Require Human Oversight to Boost Student Achievement

Recent large-scale trials reveal that while AI tutoring platforms can improve math scores, their effectiveness is severely limited by low student engagement. Without human teachers to enforce usage, most students simply do not interact with the tools enough to benefit.

By Amelie Rousseau

Education Researchers 40%Ed-Tech Developers 35%Classroom Educators 25%
Education Researchers
Emphasize that the technology's effectiveness is currently bottlenecked by low student engagement and behavioral friction.
Ed-Tech Developers
Argue that AI tutoring provides scalable, personalized practice that drives measurable gains when implemented correctly.
Classroom Educators
Maintain that human relationships and accountability are irreplaceable components of student motivation and learning.

Perspectives this story doesn't cover

  • Students who struggle with help-seeking
  • School district procurement officers

AI tutors like Khanmigo require human oversight to boost student achievement because, without a teacher enforcing engagement, most students simply do not log in or interact deeply enough to learn. The promise of a personalized AI tutor for every child has met the reality of middle school behavior: when given the option to seek help from a chatbot, most students choose not to.[1][3]

The evidence comes from a wave of new large-scale studies published in August 2026, which provide the first rigorous look at how generative AI performs in actual classrooms rather than controlled lab settings. The consensus is clear: the technology works when students use it, but access alone does not equal engagement.[1][5]

A two-year randomized controlled trial conducted by the National Bureau of Economic Research (NBER) tracked students across 18 middle schools in Hamilton County, Tennessee. The students were assigned to use Khan Academy's AI tutor, Khanmigo, during their daily remedial mathematics sessions.[1]

The NBER researchers found that assignment to the AI tool raised math achievement by 1.3 national percentile ranks per term, or roughly 0.06 to 0.08 standard deviations over a school year. For students who actively participated for a full year, the implied effect reached 0.14 standard deviations—a respectable gain comparable to traditional software practice.[1][2]

However, the log files revealed a massive engagement gap. While 96 percent of students tried Khanmigo at least once, the median student messaged the AI on only a third of the days they practiced. More critically, students engaged the tutor in only 17 percent of the exercise sessions where they made a mistake.[1]

Data from the NBER trial reveals a steep drop-off in student engagement with the Khanmigo AI tutor.
While 96 percent of students tried Khanmigo at least once, the median student messaged the AI on only a third of the days they practiced.

When students did interact with the chatbot, the dialogue was rarely substantive. The messages consisted mostly of bare answers or clicks on suggested prompts, rather than the deep mathematical reasoning the tool was designed to foster. "The binding constraint appears to be engagement: realizing the promise of AI tutoring will require getting students to use it, not just giving them access," the researchers wrote.[1]

Sal Khan, founder of Khan Academy, acknowledged these findings in an August 2026 review of the study. He noted that the 0.14 standard deviation gain is a "genuinely strong result" for a historically difficult-to-move population, but agreed that the platform's non-proactive interface and students' reluctance to seek help limited its impact. In response, Khan Academy is redesigning the tool to embed the AI directly into the practice workflow.[2]

The engagement problem extends beyond Khanmigo. A comprehensive research brief published by Stanford University's SCALE Initiative evaluated the broader landscape of AI tutoring. The report, titled "AI Tutoring is Not a Monolith," separated platforms based on their level of human involvement.[3][4]

The Stanford researchers found that fully automated, AI-only tutoring currently has "unknown effectiveness" due to severe usage drop-offs. In two separate school districts analyzed by the team, between 40 and 47 percent of students never used the independent AI platform at all.[3][4]

Even among the students who did log in, engagement was fleeting. Students used the platform for an average of just two to five minutes per week over a multi-month intervention. This falls drastically short of the 30 minutes of weekly use typically recommended by software providers to achieve measurable reading or math gains.[3][5]

Stanford researchers found that students left to use AI tutors independently fell drastically short of recommended usage times.

Conversely, the Stanford brief highlighted that the strongest evidence for AI in education currently comes from models where the technology supports a human tutor rather than replacing them. Systems that provide real-time pedagogical suggestions to adult tutors have shown significant improvements in student mastery, particularly for lower-rated instructors.[3][5]

As states and districts spend millions to roll out AI platforms, the research underscores a fundamental reality of education: kids do their best for people, not robots. "We don't have solid research showing that AI tutoring can work in the U.S. at scale," said Susanna Loeb, executive director of the National Student Support Accelerator. While artificial intelligence can lower costs and personalize practice, it cannot replicate the relationship, accountability, and motivation that a human teacher provides to keep a student engaged in the productive struggle of learning.[3][4][5]

Key points

  • A two-year NBER trial found that Khan Academy's AI tutor improved math scores by up to 0.14 standard deviations for active users.
  • However, engagement was severely limited, with the median student messaging the AI on only a third of their practice days.
  • Stanford researchers found that 40 to 47 percent of students never used independent AI platforms when left to their own devices.
  • The strongest evidence for AI in education currently supports models where the technology assists a human tutor rather than replacing them.

Key terms

Standard Deviation (SD) Gain
A statistical measure used in education research to quantify how much a student's test scores improved relative to the average spread of scores.
Generative AI Tutor
An educational software tool powered by large language models that is designed to coach students through problems via conversational dialogue rather than simply providing the correct answer.
Randomized Controlled Trial (RCT)
A scientific study design where participants are randomly assigned to either receive an intervention or be part of a control group, considered the gold standard for measuring effectiveness.

Sources

Source coverage

5 outlets

3 viewpoints surfaced

Education Researchers 40%Ed-Tech Developers 35%Classroom Educators 25%
  1. [1]National Bureau of Economic ResearchEducation Researchers

    One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment

    Read on National Bureau of Economic Research
  2. [2]Khan Academy BlogEd-Tech Developers

    What I found compelling in a new randomized trial of Khan Academy in math intervention

    Read on Khan Academy Blog
  3. [3]The 74Classroom Educators

    AI Tutors Not Yet a Replacement for Humans, Research Says

    Read on The 74
  4. [4]LA School ReportClassroom Educators

    AI Tutors Not Yet a Replacement for Humans, Research Says

    Read on LA School Report
  5. [5]Factlen Editorial TeamEducation Researchers

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get Education stories with full source coverage and perspective breakdowns delivered to your inbox.