Category

Benchmarks and what they do not tell you

A benchmark score tells you which model won. Modelometer tells you whether the model you chose is still behaving like the model you tested.

A benchmark is a snapshot of capability. It asks how well a model performs a task on the day the test was run, and it answers that well. Three things it is not built to tell you: whether the model you deployed is still the model you evaluated, what it will refuse, and where it sits on a question that has no correct answer.

Modelometer measures the third and watches the first. It runs a frozen protocol against a fixed set of contested questions on a published schedule, records where each model lands on a five-point axis from most permissive to most restrictive, records what it refuses and how, and hash-chains every run so any number can be recomputed from public data by a stranger with a standard-library script.

§

What is deliberately out of scope, permanently

Capability, accuracy, quality, and jailbreak resistance. Other instruments measure those and measure them better. Adding them would make this a benchmark, and a benchmark cannot do the job this record exists to do.

§

No model judges another here

Scoring is deterministic: fixed items, fixed parsing, fixed thresholds set before the data was seen. That is the property the whole design is built around, and it is why a change in a number means a change in a model rather than a change in an opinion.

The record began on 2026-07-17. Everything it holds is at questions, data and analytics, and the method is at methodology.