Most AI tutoring evaluations are decided by a scripted demonstration, and a scripted demonstration is exactly the environment in which every product in this category performs well. What follows is the list we would use if we were buying rather than selling, including the items where our own answer is incomplete.
On what the tutor knows
Where does an answer come from, and can the learner see the page it came from? What happens when the course material does not cover the question — does the tool say so, or does it answer anyway? Who decides what the tutor may read, and how long does a change take to reach students?
Ask for a live question outside the syllabus during the demo rather than after it. Vendors prepare the happy path; the boundary is where the architecture shows.
On data
Is one institution’s data separated from another’s by application code or by the database? The distinction is not academic: application-level separation fails the first time a query is written without the right filter, and that failure is silent. Ask to see the isolation exercised as a signed-in user of one institution attempting to read another’s, not as an administrator asserting that it cannot happen.
Then ask the unglamorous questions. Can a learner export everything held about them? Can they have it erased? Where is guardian consent recorded if the institution serves under-18s? Ask which of these are implemented and which are policy documents.
On the model itself
Which providers are in use, what happens when one is unavailable, and is there a declared timeout per modality? A product with a single hard-wired provider is one vendor incident away from an outage during finals week. Ask whether cost and latency are recorded per call — not because the buyer cares about the number, but because a vendor that cannot produce it does not have the telemetry to notice a regression.
Ask, too, whether the institution can constrain which models are used. Approved-model lists are a normal procurement control everywhere else and are unusual in this category.
On evidence, and the answer you should distrust
Ask what outcome evidence exists. If a vendor answers this question confidently and quickly, read the study. Most of this category is too young to have real efficacy data, and a confident answer usually describes engagement metrics wearing the vocabulary of learning.
Our own answer is that we do not have outcome evidence. We have a measurement chain built and no cohort has produced enough data to make a claim from, and we would rather say that than present activity counts as attainment.
On what happens after the pilot
Who owns the material an instructor uploads, and what happens to it if the contract ends? What does the tool do on a course with nothing published? How is a roster loaded, and what happens if the same roster is loaded twice? The last one sounds trivial until a re-import creates a duplicate enrolment for every student in a programme.