One model going down is not an outage.
Text, vision and voice each route through a primary with a declared fallback and a timeout, and every call records what it cost and how long it took.

Four providers behind one interface; swapping a model is configuration, not a deployment.
Cost, latency and fallback events are recorded per call rather than estimated monthly.
A declared timeout per modality, so a slow provider degrades instead of hanging the lesson.
What this rests on
- Model routing with a declared fallback and timeout per modality, and cost and latency recorded on every callText, vision and voice each pinned to a primary, a named fallback and a timeout across four providers, with cost and latency recorded per call rather than estimated monthly.
- Each institution’s data is isolated at the database level, not by application codeTen checks run as a real signed-in user, so the database’s own access rules are genuinely exercised rather than assumed.