Human Still Required / Field Notes

Students Skipped the AI Tutor. That’s the Real Lesson.

Blue screens of instant answers beside educators and students solving math; headline Who Did the Thinking? AI tutoring

Two new pieces of AI tutoring research landed this week, and together they tell a story every school leader should hear. When AI gave students the answers, their homework improved and their exam scores fell. When AI refused to give the answers and asked questions instead, most students simply stopped using it.

Neither result is really about the technology.

Both are about what students have learned school is asking of them.

Key takeaways

  • When AI supplies answers, students can finish work faster without learning more. One large study found homework scores rose while exam scores dropped.
  • When an AI tutor coaches instead of answering, many students disengage. In a two-year Tennessee trial, the tutor added nothing beyond Khan Academy practice alone.
  • The core problem is not access to AI. It is whether students choose to do the thinking, and whether school makes that choice worth it.
  • Leaders should judge AI tutoring by engagement and learning, not by adoption, and redesign tasks so completion is not the goal.

What’s happening with AI tutoring research

On September 21, Jill Barshay’s Proof Points column in The Hechinger Report covered a two-year study of Khanmigo, Khan Academy’s AI math tutor. The underlying National Bureau of Economic Research working paper by Philip Oreopoulos and Nina Low of the University of Toronto is a cluster randomized trial in 18 Tennessee middle schools. Students who were behind in math used Khan Academy with Khanmigo, configured to coach rather than give answers, during their daily remedial math sessions.

The students did improve. According to the paper, assignment raised math achievement by 1.3 national percentile ranks per term. But those gains resembled what Khan Academy practice produces without AI help. The tutor itself added nothing measurable.

The reason is the part that stuck with me. Almost every student, 96 percent, tried Khanmigo at least once. Hechinger reports that many tried to get it to spit out the answer, and when it refused and offered hints or questions instead, they largely stopped using it. The median student messaged the tutor on only a third of the days they practiced, and in only 17 percent of the sessions in which they made a mistake. The researchers put it plainly: “We observe how students actually used the AI tutor, and the answer is: not much.” Their conclusion is that the binding constraint is engagement, not access.

A day earlier, on September 20, the Australian Broadcasting Corporation reported on scientists’ concerns about AI’s impact on deep thinking. The article describes a study published this year of more than 26,000 students in grades 7 to 12. AI use lifted homework scores by 18 percent and improved completion times by 30 percent. But it lowered monthly exam scores by 20 percent within six months.

University of Tasmania neuroscientist Matt Kirkcaldie described the concern this way: students might read AI-generated text and understand what it says, “but they never went through the struggle” of working through the ideas. Kathy Mills of Australian Catholic University told the ABC that students and teachers want AI tools designed to guide learning, not just tell answers, and noted that take-home assignments make it harder for teachers to see how much AI support students are using.

Put the two together and you get one picture. Give students answers and they take the shortcut. Take the shortcut away and many walk away from the tool.

What this reveals: the real risk is atrophy

In Chapter 5 of Human Still Required, “AI as a Thinking Scaffold (Not a Thinking Replacement),” I make an argument that this week’s research lines up with almost exactly.

The real risk isn’t cheating. It’s atrophy.

Human Still Required

The homework study the ABC described is atrophy in numbers. Students looked more productive. The work got done faster and scored higher. And the learning quietly slipped away. That is the difference between activity and learning that I keep coming back to.

The Khanmigo finding is the harder one, and I think it is the more important one for leaders. The tool was designed the right way. It was built as a scaffold, not a shortcut. It asked questions. It offered hints. It tried to keep the student doing the thinking. And students declined.

That is not a failure of students. It is information about the system they are in. In Chapter 2, “When ‘Good Enough’ Became the Enemy,” I describe how students learn to play the game of school, and how AI became a cheat code for average work.

If school rewards completion over cognition, students will optimize for completion.

Human Still Required

If the goal of a remedial math session, as students experience it, is to get through the problems, then a tutor that slows you down with questions feels like an obstacle, not a help. The students were being rational. They were optimizing for what they believed counted.

This is why I don’t think the lesson here is “AI tutoring doesn’t work.” The lesson is that no tool can supply the motivation to think. A scaffold only helps if someone is willing to climb it. And the willingness to climb is shaped by the culture, the task design and the relationships around the student, all of which are human work.

The goal was never the scaffold. The goal was independence.

Human Still Required

What leaders can do now

  • Measure engagement, not adoption. If you have purchased or are piloting an AI tutor, ask for usage data that goes beyond log-ins. How often do students actually ask questions? What happens after a mistake? The Tennessee study found almost everyone tried the tool and few used it meaningfully.
  • Ask what the task rewards. Pick one common assignment in each grade band and ask your team: if a student finished this with AI in five minutes, what would they have missed? If the honest answer is “nothing we grade,” the task is rewarding completion.
  • Put thinking where you can see it. Shift some weight from take-home products to in-class explanation, conferencing and revision. Students who know they will be asked to explain their reasoning have a reason to do it.
  • Keep a human in the tutoring loop. An AI tutor running in the background of a remedial block is not the same as a teacher who notices a student is stuck and nudges them to use the help. Build that expectation into how the tool is used.
  • Talk with students about the struggle. Tell them directly why the hard part matters. Many students already sense that the shortcut costs them something. Give them language and permission to choose the harder path.

If your team wants a structured way to work through these questions together, the 5-Week Book Study is a good place to start.

Frequently asked questions

Does AI tutoring improve student learning?

The evidence so far is mixed. In the two-year Khanmigo trial in 18 Tennessee middle schools, students improved modestly, but the gains matched Khan Academy practice without AI, and the researchers traced that largely to low engagement with the tutor. How students use a tool matters as much as how it is designed.

Why do students stop using AI tutors that don’t give answers?

According to Hechinger Report coverage of the Tennessee study, many students tried to get the tutor to give them the answer and largely stopped using it when it offered hints or questions instead. When tasks reward finishing, a tool that slows students down can feel like friction rather than help.

Can AI homework help hurt test performance?

It can. The ABC reported on a study of more than 26,000 students in grades 7 to 12 in which AI use raised homework scores by 18 percent but lowered monthly exam scores by 20 percent within six months. Faster, better-looking homework is not the same as learning.

The part only humans can do

I find this week’s research oddly hopeful. It confirms that the technology is not the deciding factor. A well-designed tutor without engaged students added nothing. An answer machine without purposeful tasks made students look better while learning less.

What decides the outcome is whether students believe the thinking is the point. That belief is built by teachers, by task design and by the leaders who protect both.

So here’s my question for you this week: in your school, what would a student say actually counts, finishing the work or understanding it?

Sources

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.