Human Still Required / Field Notes

Before You Pilot AI, Decide What Success Means

Educators at a library table beside blue data screens, headline What Does Success Mean? on school AI pilots

School AI pilots are everywhere this fall. A chatbot for social-emotional check-ins here. A humanoid robot proposal there. A careful, grade-by-grade rollout somewhere else. What most of them have in common is not the tool. It’s the missing answer to a simple question: what would success look like?

That is not a technology problem. It is a leadership problem. And it is one we can solve before the next contract renews.

Key takeaways

  • Many school AI pilots are running without an agreed definition of success, so districts end up measuring how people feel about a tool rather than what students learn.
  • The research base on AI in K-12 is still thin, which puts more weight on local judgment, not less.
  • Leaders should define what they want students to know, do and become before choosing a tool, and decide in advance what evidence will count.
  • Evidence of learning in an AI era looks like explanation, critique and transfer, not faster output.

What’s happening with school AI pilots

On September 28, NPR’s Lee V. Gaines reported that schools are experimenting with AI with little evidence or policy to guide them. The story walks through several districts. Dayton Public Schools in Ohio paid the ed tech company SchoolAI $72,600 last academic year, including for a social-emotional chatbot named Jordan used at an alternative school, and renewed the contract at $78,650 for this year (that makes the Private AI Server I’m pushing seem cheap!). A sophomore there described building “a little relationship” with the chatbot and said it helped them work through their own problems.

Then comes the harder part. Mike Kentz, an outside consultant evaluating the work, told NPR: “Every test, every experiment that I try to measure in some sort of concrete way, I can only get as far as perception.”

In Wichita, AI specialist Katelyn Schoenhofer was even more direct: “We wrestle with what does success really mean when it comes to AI use? And what metric do we want to use in order to evaluate that? And we have not come to an answer yet.” Wichita is moving ahead with caution, teaching younger students about AI and giving older students more limited access to tools.

The story also notes that the superintendent in Salamanca, New York, planned to bring in an AI humanoid robot but halted after community concerns about data privacy (listen to the podcast I did with the Superintendent here), and that New York City and Los Angeles have restricted student-facing AI. It cites a Stanford review of more than 800 academic papers that found very little research on how AI can help students while avoiding potential harms, with some studies suggesting carefully designed tools show more promise than general-purpose chatbots. Robin Lake of the Center on Reinventing Public Education called the landscape “very Wild West in terms of the randomness.” Rebecca Winthrop of the Brookings Institution said educators “should really be at the forefront” and that experimentation “should be [for] a problem they’re facing.”

The same day, District Administration published a piece by Jennifer Womble, chair of the FETC conference, titled AI cannot lead your district. Now, only people can. Her diagnosis lines up with NPR’s reporting: “Schools do not have a technology adoption problem. They have a human adoption problem.” She argues district leaders “must first define what they want students to know, do and become in tandem with local communities” and then decide where technology helps. Her warning: “The biggest risk for school districts is not that they will move too slowly. It is that the speed of technology will begin to set their direction for them.”

What this reveals: the pilot is not the plan

Put those two pieces side by side and a pattern shows up. Districts are not short on tools. They are short on a shared answer to why a tool is in the room and how anyone will know if it worked.

When success isn’t defined up front, perception becomes the metric by default. Did students like it? Did teachers use it? Did the dashboard show logins? Those are fair questions. They are not the same as, “Did students learn to think better?”

In Chapter 10 of Human Still Required, I argue that the new leadership skill is discernment, and that one of the most useful questions a leader can ask in any meeting is this:

What are we optimizing for here?

Human Still Required

That question belongs at the start of every pilot, not the end. If a team can’t answer it in a sentence, the pilot is running on momentum. I call that momentum without meaning, and it is how districts end up renewing contracts on the strength of good feelings.

Chapter 7 makes a related point. A thin research base doesn’t let leaders off the hook. It puts more of the decision on them. National evidence can tell you what might work. Only local judgment can tell you what fits your students, your staff and your community.

Wise leaders don’t ask, ‘Is this recommended?’ They ask, ‘Is this right for us?’

Human Still Required

That’s why I’d push back gently on waiting for the research to settle everything. The research matters. But no study will decide what your community values. Womble says it plainly: AI “cannot decide what a community owes its children.” That decision is ours.

There is also a measurement question hiding in all of this. In Chapter 6, I lay out the D.E.C.I.D.E.R. framework for what still counts as evidence of learning when output is cheap: Decision-making, Explanation, Critique, Inference, Development, Evaluation and Relocation (transfer to a new context). If a district wants to know whether an AI pilot is helping students, those are the places to look. Not how quickly work gets done. Whether students can explain it, defend it and use it somewhere new.

And for a tool like an SEL chatbot, the Chapter 8 question matters too: is the tool connecting students back to trusted adults, or standing in for them? I’m not judging Dayton’s choice from a distance. I’m saying that is exactly the kind of question a district should decide it will measure before the renewal, not after.

What leaders can do now

  • Write the one-sentence purpose. Before any new pilot, or any renewal, name the specific problem the tool is meant to address and who has that problem. Winthrop’s point is a good test: is this solving a problem educators are actually facing?
  • Decide what evidence will count, in advance. Pick two or three signals tied to student thinking, such as whether students can explain their reasoning, critique an answer or transfer a skill. Usage data and satisfaction surveys can sit alongside them, not replace them.
  • Set a decision date. Put a date on the calendar when the team will look at that evidence and decide to continue, change or stop. A pilot without an end point is just a purchase.
  • Bring the community in early. Salamanca’s robot plan stopped after community concerns. Ask families and staff what must remain human before the rollout, not after the headline.
  • Ask the optimization question out loud. In your next cabinet or leadership team meeting, ask, “What are we optimizing for here?” about one current AI initiative. Listen for whether the answers match.

If your leadership team wants a structured way to work through these questions together, the 5-Week Book Study is built for exactly that kind of conversation.

Frequently asked questions

How should a school district measure whether an AI pilot is working?

Start by defining the problem the tool is meant to solve and what student learning should look like if it works. Then choose a small number of measures tied to thinking, such as explanation, critique and transfer, and set a date to review them. Usage and satisfaction data are useful context but shouldn’t be the whole story.

Should we wait for more research before piloting AI in schools?

Not necessarily. The research base is still limited, but waiting doesn’t remove the need for judgment. A better approach is to pilot narrowly, around a real problem educators are facing, with clear evidence criteria and a planned decision point.

Who should decide what success looks like for school AI pilots?

District leaders should own the decision, but not alone. Teachers, families and students should help define what must remain human and what benefits matter most. The tool vendor can supply data; the community decides what counts.

The question behind the question

Every AI pilot is a small bet about what school is for. That’s why the success question feels so hard. It isn’t really about metrics. It’s about values.

Tools will keep arriving faster than research can study them. That’s the reality. The districts that do well won’t be the ones that adopted first or banned first. They’ll be the ones that knew what they were trying to protect and grow, and measured that.

So here’s my question for you: pick one AI tool running in your schools right now. Could you tell your community, in one sentence, what success looks like and how you’ll know?

Sources

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.