Skip to content
Saturday, August 29, 2026
Education FameEducation · EdTech
Research · Learning · Evidence
EdTech

How Districts Should Evaluate AI Tutors Before Buying

A due-diligence guide: the questions about evidence, data, accuracy, and pedagogy districts should put to every AI tutoring vendor.

Evaluation checklist chart for AI tutoring products by category

AI tutors — conversational systems that answer student questions, walk through problems, and adapt to responses — moved from novelty to procurement pipeline in under three years, and districts now face sales materials that outpace the research base. The disciplined response is not to ban the category or to trust it, but to subject it to the same due diligence applied to any instructional material, plus a set of questions unique to generative systems. Federal guidance has encouraged exactly this posture: the U.S. Department of Education's 2023 report on artificial intelligence in teaching and learning called for alignment to a human-centered approach, and its subsequent work on safe and transparent model design stressed evaluation and transparency. This guide organizes the questions districts should ask before signing.

What is actually being sold?

The label AI tutor covers at least three different products: a wrapper around a general-purpose chatbot with an education prompt; a system grounded in a specific curriculum that confines its answers to that content; and an intelligent tutoring architecture, with decades of learning-science history, augmented by language models. These differ enormously in evidence and risk. The first procurement question is therefore definitional: what model or models power the system, what content constrains its answers, and what happens when a student asks something outside the curriculum? Vendors that cannot answer plainly are asking districts to buy an undescribed artifact.

What evidence questions should districts ask?

  1. What controlled studies exist, who ran them, on what populations, and with what comparison conditions?
  2. Were outcomes measured with independent assessments or the product's own metrics?
  3. What were the effect sizes and durations, and were results published anywhere subject to review?
  4. How does the vendor monitor whether outcomes hold across student subgroups?
  5. What usage levels did the studies assume, and does the district's schedule support them?

The research base for large language model tutoring remained thin relative to adoption as of 2025 — early experimental studies showed both promise on specific tasks and persistent weaknesses on accuracy — so districts should treat vendor claims of proven results skeptically and weight pilot designs heavily.

What data and privacy questions matter most?

AI tutors are conversation engines, and conversations reveal a great deal. Districts should establish, in writing, what student inputs are collected; whether prompts and responses are stored, for how long, and where; whether data is used to train or improve models, including through human review of transcripts; and whether the vendor commits to not using student data for any commercial purpose. FERPA and COPPA obligations described earlier still apply — the school official exception for FERPA and, for products serving students under 13, the FTC's COPPA framework, which the FTC tightened with an updated rule finalized in 2025. Ask for the data flow diagram: a vendor that cannot draw where student speech goes should not hear it.

What accuracy risks need assessment?

Language models generate plausible text that can be wrong — hallucination is the industry term — and in tutoring the errors that matter are pedagogical, not only factual: a confidently wrong step in an algebra derivation, or a hint that completes the problem for the student. Districts should ask what guardrails detect errors, what independent accuracy testing the vendor has run and at what measured error rate, how errors discovered after deployment are corrected and communicated, and whether the system indicates uncertainty to students. A live test during evaluation is worth any slide deck: have staff and students ask calibrated questions, including common misconceptions in the relevant grade level, and record what happens.

What pedagogy questions separate tutors from answer machines?

Question to askWeak answerStrong answer
Does it give answers?Immediate full solutions on requestScaffolded hints; solution only after attempts
How does it handle struggle?Repeats the explanationDecomposes the step, changes representation
Teacher controls?NoneConfigurable restrictiveness, topic limits
Curriculum grounding?General knowledge onlyConfined to adopted course content
Equity monitoring?Not discussedOutcome reporting by subgroup

What should the pilot look like?

An AI tutor pilot needs more structure than a typical software trial. Define the instructional purpose narrowly — eighth-grade algebra support, for example — and choose a measurable outcome on an independent assessment. Randomize or match comparison classrooms where feasible, since enthusiasm effects are real in the first months of any novel tool. Capture teacher observation systematically, because teachers will see misuse patterns the logs miss: students offloading work, asking for answers, or testing the system with jokes. Set a decision date and a written standard for what evidence would justify districtwide purchase, and let the standard exist before the pilot begins.

What contract terms are specific to AI tools?

Beyond standard edtech terms, districts should require: explicit language that student data will not train vendor models without separate written authorization; disclosure of material changes to the underlying model, since behavior can shift with updates; accuracy and incident notification obligations; academic-honesty features described concretely; and an exit clause covering deletion of stored conversations. Vendors update models frequently, and a product evaluated in spring may behave differently in fall — change disclosure is the contract term that addresses that reality.

Should districts start at all, or wait?

A reasonable district can choose either path, but the choice should be deliberate. Waiting has costs: teachers and students already use general-purpose chatbots with no safeguards, no data protections, and no pedagogical design, so a governed district tool can reduce risk rather than add it. Moving early has costs too, since products, models, and state rules are still shifting. The middle path many districts have settled on is small, tightly scoped pilots with strong contract terms and explicit board awareness — enough to build institutional judgment without betting instruction on an immature market. Whichever path a district chooses, writing down why and revisiting the decision annually keeps the question from being settled by default.

How should districts involve teachers and families?

Teachers should be in the evaluation from the first demo, because they will predict misuse better than any rubric. Families deserve plain-language notice that an AI system will interact with their children, what it collects, and how to opt out where the district can honor it — several states began requiring disclosure or consent for AI in schools through legislation enacted in 2024 and 2025, and requirements vary by state. Districts that communicate early tend to keep the community with them when a glitch makes the news, which in this category it eventually will. The districts that navigate AI tutoring well will be the ones that bought slowly, measured honestly, and kept teachers in charge of what happens in the classroom.

Frequently Asked Questions

What should districts ask AI tutor vendors about evidence?
Ask for controlled studies, who ran them, on which populations, with what comparison conditions, and whether outcomes were measured with independent assessments rather than the product's own metrics.
Can student conversations with an AI tutor be used to train models?
Only if the contract allows it. Districts should require explicit language that student data will not train or improve vendor models without separate written authorization.
What accuracy risks do AI tutors carry?
Language models can generate confident but wrong explanations, so districts should ask about guardrails, independently measured error rates, and how post-deployment errors are corrected and communicated.
How should an AI tutor pilot be designed?
Narrowly: a defined instructional purpose, a measurable outcome on an independent assessment, comparison classrooms where feasible, systematic teacher observation, and a written decision standard set before the pilot begins.