Multilingual support is usually asserted rather than measured, and the assertion is usually true of conversation and false of anything technical. We measured ours and the gap was larger than we expected.
The measurement
Against an English grounded-STEM score of 87%, the same evaluation scored 15.8% in Spanish and 6.3% in Polish. Those are not rounding differences; they describe a capability that works in one language and does not yet work in two others we tested.
The failure is concentrated where retrieval, technical vocabulary and symbolic notation meet. Conversational tutoring in the same languages behaves far better, which is precisely why a single "multilingual" claim is misleading — it averages two very different things.
Why we publish it
An institution serving a multilingual cohort will discover this in week two of a pilot. Discovering it from the vendor in week zero is a better relationship and a shorter procurement, and it lets the pilot be scoped to where the tool is strong rather than failing broadly and being blamed generally.
It is also the only way our own marketing rule can mean anything. A page may only cite a capability that names the gate that proved it. The same discipline has to apply when the gate returns a number nobody wants.
What we are not claiming
We are not claiming parity, and we are not claiming a date. The work to close it is retrieval and evaluation work in each target language, and until an evaluation says otherwise the honest position is that grounded STEM is an English capability with conversational support elsewhere.