Every institution evaluating AI tutoring is shown the same demo: a question goes in, a fluent paragraph comes out. The paragraph is almost always correct, which is exactly why the demo is not evidence of anything. The question a dean is actually asking is narrower — when this tool is wrong, how will anybody find out?
A general assistant cannot answer that question, because it has no notion of which course it is in. It draws on everything it has read and presents the result in one voice. A course-grounded tutor is a different architecture: it retrieves from the documents your instructor published, answers from those passages, and attaches a reference to the page it used.
A citation is a falsifiable claim
When an answer names the document and page it came from, a student or an instructor can check it in seconds. That single property changes what the tool is. An unsourced paragraph has to be trusted or discarded; a sourced one can be verified, and a wrong one can be traced back to the material that produced it.
It also changes what happens when the material is thin. A retrieval system knows the difference between "the course covers this" and "the course does not", because that difference is visible in what came back from the index. A general assistant has no such boundary and will answer either way, in the same confident register.
The refusal is the feature, not the failure
The behaviour institutions underrate in evaluation is the refusal. Ask a course-grounded tutor something the syllabus does not cover and the right answer is to say so and point the learner at the person who can help. We treat that as a first-class screen rather than an error state, because a learner who is told plainly that the material stops here goes to office hours, and a learner given a plausible improvisation does not.
In our Phase 2 acceptance gate, out-of-scope refusal measured 100% against a 54.2% baseline for dumping whole documents into the prompt. The baseline number is the honest part of that comparison: passing a course pack into context is a common vendor architecture, and it is where most unsourced confidence comes from.
Grounding is an instructor workflow before it is a retrieval problem
The retrieval half is well understood. The half that decides whether grounding is real is editorial: who decides what the tutor may read, and when does it take effect. An instructor attaches documents to a course, reviews them as a draft, and publishes when ready. A draft is invisible to students, and that invisibility is enforced by the database rather than by the page that renders it.
That ordering matters for a procurement conversation. It means the answer to "what is the tutor telling our students" is a list of published documents an instructor can read, not an opaque model state nobody at the institution controls.
What to ask in an evaluation
Ask for a question the course material does not cover, and watch what happens. Ask for the source of an answer and check the page. Ask what a student sees while a document is still in draft. Ask what the tutor does on the first day of term, when a course has been created and nothing has been published yet.
Those four questions separate a grounded product from a general assistant with an institutional logo on it, and none of them require a technical reviewer to answer.