Category
Benchmarks and what they do not tell you
A benchmark score tells you which model won. Modelometer tells you whether the model you chose is still behaving like the model you tested.
A benchmark is a snapshot of capability. It asks how well a model performs a task on the day the test was run, and it answers that well. Three things it is not built to tell you: whether the model you deployed is still the model you evaluated, what it will refuse, and where it sits on a question that has no correct answer.
Modelometer measures the third and watches the first. It runs a frozen protocol against a fixed set of contested questions on a published schedule, records where each model lands on a five-point axis from most permissive to most restrictive, records what it refuses and how, and hash-chains every run so any number can be recomputed from public data by a stranger with a standard-library script.
What is deliberately out of scope, permanently
Capability, accuracy, quality, and jailbreak resistance. Other instruments measure those and measure them better. Adding them would make this a benchmark, and a benchmark cannot do the job this record exists to do.
No model judges another here
Scoring is deterministic: fixed items, fixed parsing, fixed thresholds set before the data was seen. That is the property the whole design is built around, and it is why a change in a number means a change in a model rather than a change in an opinion.
The record began on 2026-07-17. Everything it holds is at questions, data and analytics, and the method is at methodology.