Skip to main content
Factlen AnalysisAI in EducationResearch FindingAug 13, 2026, 9:45 PM· 6 min read· in education

Stanford Study Finds AI Tutoring Fails to Boost Reading Scores Due to Critically Low Student Engagement

A new randomized controlled trial reveals that elementary students given independent access to an AI reading tutor used it for just two to five minutes a week. Researchers concluded that while artificial intelligence can personalize instruction, it requires human support to motivate actual student usage.

By Nabil Faris

Implementation Researchers 50%Education Reporters 30%EdTech Advocates 20%
Implementation Researchers
Argue that technology only works when paired with structured human support and adequate dosage.
Education Reporters
Focus on the practical realities and failures of rolling out new tech in understaffed schools.
EdTech Advocates
Believe AI can scale personalized instruction to solve educational gaps.

Why it matters

School districts nationwide are investing millions in AI tutoring platforms to solve pandemic-era learning loss at scale. This research demonstrates that purchasing software licenses is insufficient without also funding the human staff required to keep young students on task.

When two unnamed school districts carved out dedicated time for roughly 350 elementary students to use a well-known artificial intelligence reading tutor, the results were startlingly clear: nearly half of the students assigned to work independently never logged on to the platform at all. Across the entire independent group, average weekly usage hovered between a mere two to five minutes. The randomized controlled trial, conducted by researchers at Stanford University, sought to understand how young learners interact with highly touted educational technology when left to their own devices. Instead of embracing the personalized learning software, the students largely ignored it, demonstrating a massive gap between the theoretical capabilities of generative AI and the practical realities of an elementary school classroom. The findings serve as a reality check for an industry that has aggressively pitched digital tutors as a seamless, scalable solution to nationwide reading deficits.[1][2]

The actionable takeaway for school administrators, policymakers, and parents is unambiguous: providing access to educational technology does not automatically equate to student engagement. The software provider behind the AI reading tutor explicitly recommended that students complete at least two 30-minute sessions per week to see any measurable gains in their literacy skills. Because the vast majority of students failed to reach even a fraction of this 60-minute threshold, the intervention ultimately produced no meaningful improvement in reading achievement across either of the participating school districts. This disconnect highlights a critical flaw in how educational technology is often procured and deployed; districts frequently invest heavily in software licenses while underestimating the human infrastructure required to ensure those tools are actually utilized. Without a mechanism to keep easily distracted young learners on task, the software remains an untapped resource.[1][4]

To test whether human interaction could bridge this engagement gap, the Stanford researchers from the SCALE Initiative introduced a secondary variable into the trial. They paired a subset of the students with human tutors—afterschool program staff in one district, and high-achieving middle school students in the other. Crucially, these human aides were not instructed to teach reading concepts or provide direct academic instruction. Instead, their sole mandate was to offer motivation, enforce accountability, and provide basic technical troubleshooting to keep the younger students focused on the AI platform. The experiment was designed to isolate the value of human encouragement from the actual delivery of the curriculum, testing whether a supportive presence could coax young children into engaging with a digital interface that they otherwise ignored when left alone.[3][4]

Students averaged just 2 to 5 minutes of weekly usage, falling far short of the 30 minutes recommended by developers.

Adding this layer of human support did successfully move the needle on student participation, though the absolute gains remained remarkably small. Students who were paired with a human tutor increased their average weekly platform usage by one to four minutes compared to their independent peers. Furthermore, this human-backed group saw their overall engagement—measured by the total number of reading stories completed on the platform—jump by 71 to 80 percent. However, because the baseline usage was so critically low to begin with, this added exposure still amounted to less than two extra hours of total platform time across the entire 14-to-31-week intervention window. While the percentage increase in engagement looks impressive on paper, the actual time spent reading was still far too low to generate the dosage required to meaningfully impact standardized test scores.[1][4]

Adding this layer of human support did successfully move the needle on student participation, though the absolute gains remained remarkably small.

Beyond the overall lack of engagement, the data revealed a concerning equity gap that could have profound implications for how AI is deployed in public schools. Among the students who were assigned to work independently, those receiving special education services and those with lower baseline academic achievement were the least likely to log on to the platform. Conversely, the students who did manage to rack up minutes on the AI tutor were generally higher-performing individuals who already possessed the intrinsic motivation to read. This dynamic suggests that deploying AI learning tools without structured, in-person human oversight could inadvertently widen existing educational disparities. If only the most self-directed, high-achieving students take advantage of personalized digital instruction, the technology will fail the exact vulnerable populations it is most often purchased to help.[1][5]

These disappointing usage metrics align closely with early reports from other high-profile artificial intelligence education initiatives across the country. Sal Khan, the founder of Khan Academy and one of the earliest champions of generative AI in the classroom, recently acknowledged that student engagement with their proprietary AI tutor, Khanmigo, was significantly lower than initially anticipated. EdTech companies routinely pitch these platforms as a highly scalable, cost-effective solution to the ongoing national teacher shortage, promising that algorithms can deliver the kind of one-on-one attention that human staffing levels simply cannot support. However, the Stanford data underscores a fundamental truth about early childhood education: technology cannot replace the interpersonal connection, gentle nudging, and emotional regulation that human educators provide to keep young minds focused on difficult tasks.[1][2][5]

Pairing students with a human tutor increased their engagement with the AI platform by up to 80 percent.

For school districts that are currently piloting or preparing to launch AI tutoring programs this fall, the immediate next step is to audit actual student usage logs rather than simply tracking the number of licenses activated. Researchers strongly advise that schools must budget for paraprofessionals, classroom aides, or trained community volunteers to facilitate these digital sessions. The software must be treated as a supplemental curriculum tool that enhances a human relationship, rather than a standalone, set-it-and-forget-it fix for learning loss. Administrators are being urged to view AI not as a replacement for staffing, but as a resource that requires its own dedicated human infrastructure to function effectively in an elementary school environment.[3][4][5]

Ultimately, the Stanford study proves that while artificial intelligence has achieved the technical capability to generate highly personalized, adaptive lesson plans, it cannot manufacture the intrinsic motivation required to complete them. Reading is a difficult, cognitively demanding task for young children, and the friction of learning to decode text often requires the empathetic encouragement of a human being. The future of artificial intelligence in the elementary classroom will likely depend on hybrid models where the technology handles the heavy lifting of curriculum differentiation, while human educators handle the complex, irreplaceable work of managing the student. Until that balance is struck, even the most advanced AI tutors will remain unused.[4][5]

What to know

  • Elementary students given independent access to an AI reading tutor averaged just 2-5 minutes of weekly use.
  • Nearly half of the students in the independent control group never logged onto the platform.
  • Pairing students with a human tutor increased engagement by up to 80 percent, though total usage remained low.
  • The critically low dosage resulted in no measurable improvement in student reading scores.

Where opinion splits

EdTech Developers

Proponents who view AI as the ultimate tool for scaling personalized education.

Technology developers argue that AI platforms offer a level of individualized feedback and pacing that is impossible to achieve in a traditional classroom of thirty students. They maintain that the underlying technology is sound and capable of adapting to various learning styles. From this perspective, the current challenge is a user-interface and implementation problem, not a failure of the AI itself, suggesting that gamification and better software design could eventually bridge the engagement gap.

Education Researchers

Academics focused on the empirical evidence of implementation and dosage.

Researchers emphasize that the theoretical capability of an educational tool is irrelevant if students do not use it. They point to the Stanford study as proof that 'access' is a flawed metric for equity or success. This camp argues that school districts must shift their focus from acquiring software licenses to developing structured implementation plans, insisting that any AI rollout must be accompanied by funding for the human staff required to monitor and motivate young learners.

Classroom Educators

Teachers who prioritize the relational aspect of early childhood learning.

Veteran teachers argue that elementary education is fundamentally a relational enterprise. They note that young children rarely possess the executive function or intrinsic motivation to sit independently and work through challenging reading exercises, regardless of how sophisticated the software is. For this camp, the study validates what they have long known: human connection, encouragement, and accountability are the true engines of early literacy, and technology can only ever serve as a supplement to that bond.

Sources

Source coverage

5 outlets

3 viewpoints surfaced

Implementation Researchers 50%Education Reporters 30%EdTech Advocates 20%
  1. [1]ChalkbeatEducation Reporters

    Research on AI tutoring ran into a problem: Most students wouldn't use it

    Read on Chalkbeat
  2. [2]The 74Education Reporters

    Study: Giving Kids Access to AI Tutors Doesn't Mean They'll Use Them

    Read on The 74
  3. [3]Stanford NSSAImplementation Researchers

    Access is Not Enough: Human Support Improves Engagement with AI Tutoring

    Read on Stanford NSSA
  4. [4]EdWorkingPapersImplementation Researchers

    Access is Not Enough: Human Support Improves Engagement with AI Tutoring

    Read on EdWorkingPapers
  5. [5]Factlen Editorial TeamImplementation Researchers

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get education stories with full source coverage and perspective breakdowns delivered to your inbox.