Category
Model drift, and how you would know
A benchmark score tells you which model won. Modelometer tells you whether the model you chose is still behaving like the model you tested.
Drift is when the model behind an endpoint stops behaving the way it did when you chose it. The name covers several different events with one word: a provider ships a new build behind the same name, a serving stack changes, a system instruction is added upstream, or traffic is split between two backends. From the outside they look identical. What you see is that answers moved.
Why it is hard to notice
Nobody is notified. The endpoint name does not change. Any single odd answer is indistinguishable from ordinary sampling variation, so the only way to separate weather from climate is to have measured the same fixed questions, the same way, before and after, and to have fixed the threshold in advance.
What we do about it
The same frozen items go to the same endpoints on a published schedule. Every run is hash-chained. Two paraphrases of each core item are held back and never published, and their hash commitments are, so a reader can check at any freeze that the held-back wording was fixed in advance rather than chosen afterwards. The published set also carries smoke and calibration questions, which the reserve does not cover; the questions page names the split. A gap that widens between published and held-back phrasings is the signature of memorization rather than reasoning.
What we do not do
We do not say why a model changed. We observe outputs; we do not see training runs, and a claim about a provider's internal conduct is not ours to make. We also do not call a change good or bad. We say what moved, in which direction, on which question, between which dates, and we let you decide whether that matters for what you built.
A finding here always carries four things
What changed, over what window, by how much on how many items, and how to reproduce it. Anything missing one of them is not a finding. No finding has been published yet. The confirming step that gates a published behavioral change is specified and not yet built. Until it exists we do not publish behavioral change alerts at all.
The record began on 2026-07-17. Current readings are at analytics and change cards at diffs.