About · the bureau
An instrument, not an opinion.
Modelometer keeps an independent, citable, hash-chained record of how frontier AI models behave, the positions they take on contested questions, what they decline, and when their behavior silently changes.
The archive is the point. Deprecated model builds can never be re-measured, so the value compounds only if the record starts early and continues without silent gaps. That is why the pipeline going live and staying live matters more than any feature, and why gaps, outages, and aborted runs are published rather than papered over.
No capability or quality scores, ever. The meter tracks disposition and its change, what a model holds, not how smart it is.
Findings are numbers with uncertainty. No vendor name is placed near an evaluative adjective, the pipeline enforces this mechanically.
What a model declines to engage, and how, is measured as carefully as what it asserts. Elevated refusal on hard items is signal, not failure.
Every number carries n and SE. The site banners its own staleness. The validation study measures the meter before the meter measures anything else.
Who runs it
Modelometer is operated as an automated measurement bureau, gated by a named human owner who holds the keys and signs off on anything that pairs a vendor with a judgment. Model outputs are treated as data: rendered escaped, never executed. It is not a product of any AI lab; downstream consumers are disclosed.
(Full operator-disclosure and founder pages are drafted for the v1.0 launch.)
Status & data
This is the pre-series pilot. A validation study fixes the meter's own error bars and freezes the protocol as v1.0; until then, battery numbers are provisional and no result is treated as final.
Aggregates, charts, and site text are offered under CC BY 4.0 (license text pending ratification). Raw trial-level data is free for research with attribution; commercial use or redistribution is by license. The methodology and the full run archive are public, machine-readable index · llms.txt.