Skip to content
BlogProcurementFor institutions

The demo is not the product. These are the questions that find out.

A procurement checklist for AI tutoring, written from the side of the table that has to answer it — including the questions we answer badly.

9 min read

Most AI tutoring evaluations are decided by a scripted demonstration, and a scripted demonstration is exactly the environment in which every product in this category performs well. What follows is the list we would use if we were buying rather than selling, including the items where our own answer is incomplete.

On what the tutor knows

Where does an answer come from, and can the learner see the page it came from? What happens when the course material does not cover the question — does the tool say so, or does it answer anyway? Who decides what the tutor may read, and how long does a change take to reach students?

Ask for a live question outside the syllabus during the demo rather than after it. Vendors prepare the happy path; the boundary is where the architecture shows.

On data

Is one institution’s data separated from another’s by application code or by the database? The distinction is not academic: application-level separation fails the first time a query is written without the right filter, and that failure is silent. Ask to see the isolation exercised as a signed-in user of one institution attempting to read another’s, not as an administrator asserting that it cannot happen.

Then ask the unglamorous questions. Can a learner export everything held about them? Can they have it erased? Where is guardian consent recorded if the institution serves under-18s? Ask which of these are implemented and which are policy documents.

On the model itself

Which providers are in use, what happens when one is unavailable, and is there a declared timeout per modality? A product with a single hard-wired provider is one vendor incident away from an outage during finals week. Ask whether cost and latency are recorded per call — not because the buyer cares about the number, but because a vendor that cannot produce it does not have the telemetry to notice a regression.

Ask, too, whether the institution can constrain which models are used. Approved-model lists are a normal procurement control everywhere else and are unusual in this category.

On evidence, and the answer you should distrust

Ask what outcome evidence exists. If a vendor answers this question confidently and quickly, read the study. Most of this category is too young to have real efficacy data, and a confident answer usually describes engagement metrics wearing the vocabulary of learning.

Our own answer is that we do not have outcome evidence. We have a measurement chain built and no cohort has produced enough data to make a claim from, and we would rather say that than present activity counts as attainment.

On what happens after the pilot

Who owns the material an instructor uploads, and what happens to it if the contract ends? What does the tool do on a course with nothing published? How is a roster loaded, and what happens if the same roster is loaded twice? The last one sounds trivial until a re-import creates a duplicate enrolment for every student in a programme.

What this rests on

  • Answers cited back to your own material, with the instructor one tap away when the material does not cover itOn our go/no-go benchmark every answer retrieved the right passage, carried a valid citation and stayed inside the material, and every out-of-scope question routed the learner to the instructor rather than a guess — against 54.2% for handing the model the whole document.
  • Each institution’s data is isolated at the database level, not by application codeTen checks run as a real signed-in user, so the database’s own access rules are genuinely exercised rather than assumed.
  • Data export, erasure and guardian consent are implemented and testedExport, erasure and guardian consent each ship today and are driven through a real browser before release.
  • Model routing with a declared fallback and timeout per modality, and cost and latency recorded on every callText, vision and voice each pinned to a primary, a named fallback and a timeout across four providers, with cost and latency recorded per call rather than estimated monthly.