A benchmark asks whether a model can do something. A behavioral measurement asks what it will do: the position it takes on a contested question, what it declines, and whether either has changed. These are different questions, and the second one is mostly unmeasured.

Capability leaderboards answer a real question well: can this model solve the problem? Score the math set, the coding set, the reasoning set, rank the models, done. But “can it” is not the only thing that matters once a model is deployed and making judgment calls millions of times a day. The other question is “what does it hold”: where it lands when a question has no single correct answer, what it refuses, how firmly, and whether any of that drifts over time.

What capability scores leave out

A model can ace every benchmark and still shift, quarter to quarter, on whether it will take a firm stance on a contested policy question, or how often it declines a sensitive-but- legitimate request. None of that is a capability. It is disposition, and it does not have a right answer to grade against, so it falls outside what a benchmark is built to see. Two models with near-identical benchmark profiles can behave quite differently in the seat.

Why behavior needs different methods

You cannot grade a stance the way you grade a math answer. There is no key. So the method changes: ask the same contested question many times, record the distribution of positions rather than a single pick, keep the axis fixed and anchored so the answers are comparable across models and across time, and treat refusals as their own measured category instead of as wrong answers. The output is not a rank of who is best; it is a map of who holds what, with uncertainty attached, and a delta when it moves.

The two are complementary

This is not an argument against benchmarks. Capability and behavior are different axes, and a serious picture of a model needs both: what it can do, and what it does. Modelometer only takes the second axis, on purpose. It publishes no capability or quality scores and never ranks models by how good they are; it keeps the behavioral record and reports change.

?

Common questions

Is Modelometer a benchmark?

No. A benchmark scores capability against a known answer key. Modelometer measures behavior, the positions a model takes and what it refuses, which have no answer key, and reports how they change over time.

Why not just use an existing benchmark like MMLU?

Those measure capability, which is a different axis. They cannot tell you whether a model's stance or refusal behavior shifted, because there is no correct stance to grade against.

Do you score which model is best?

No. There is no quality ranking. The unit of a finding is a delta with an error bar, not a verdict on which model is better.