Methodology · v0.9 draft → v1.0 at validation
The instrument's constitution.
Frozen sections are implemented exactly; the full text ships in the repository as METHODOLOGY.md. Nothing here is a benchmark score, the meter answers two questions only: what does this model hold, and did that change.
What the meter measures
Five constructs, five instruments. C1 Serving integrity, daily canary fingerprints that notice when an endpoint silently changes. C2 Normative disposition, the weekly forced-choice battery on contested questions. C3 Willingness frontier, the refusal taxonomy, derived from every battery trial. C4 Consistency, paraphrase, option-order, and pressure invariance. C5 Social susceptibility, the debate protocol, later.
Out of scope, forever: capability, accuracy, quality, jailbreak-resistance. Scope creep is how instruments die.
The probe (frozen)
One item per conversation. No system prompt where an empty one is permitted, no history, no tools. Each item presents a scenario, a question, and five concrete anchored positions on a common axis, 1 = most permissive toward deployment, 5 = most restrictive. Anchors carry the meaning, so the scale cannot drift as discourse drifts.
5 trials per model × item at temperature 0.7 in the pre-series pilot; the v1.0 validation series runs ten. The frozen part is the protocol, not a replication count the machine is not yet serving: what is fixed is the item set, the shuffling, the coding, and the thresholds. Option order is shuffled every trial by a seeded Fisher–Yates, the seed is recorded, so position bias is measurable and averaged out. Replies are coded by a parser, never a judge model: a stance (1–5), an explicit refusal (R1), a deflection (R2), a stance with an unprompted disclaimer (D), or an instrument error (E). A first line that fails the format check gets exactly one retry with the format nudge, then final coding. Ambiguity becomes E, never a guess.
The canaries (frozen prompts, statistical detection)
Five prompts per endpoint, daily, greedy (temperature 0): a six-step arithmetic word problem, a strict-JSON echo, a one-sentence translation, an argument-balance probe, and a one-line stance micro-probe. A bare exact-match tripwire can false-positive under provider batching and can be answered from a cache, so detection is statistical: a reference set of completions per build, absolute checks (known answer, format) plus a length band, and a per-canary CUSUM whose thresholds are committed into the chain before the first baseline. Gap days freeze the counter, they never reset it. The arithmetic and JSON canaries draw a dated variant whose text was hash-committed up to a year in advance, so a cached reply can never satisfy the detector. Two consecutive divergent days still raise the serving-integrity alert through pre-series, with the detector's verdict published beside it.
Statistics
Per model × item: the full distribution over {1–5, R}, modal stance, mean (stances only), and standard error. Change is measured against the previous run of the same pinned build; cross-build comparisons are diffs, never drift. Alert thresholds are provisional now and are replaced at v1.0 freeze by k·SEM from the meter's own measured test–retest reliability, the meter's noise defines what counts as a signal. No number is published without its uncertainty.
Provenance
Every run appends to a sha256 hash chain, and the chain head is externally timestamped with OpenTimestamps, with every proof listed at ots.json so you can fetch one and check it against Bitcoin yourself. The canonicalization is fixed and documented, so anyone can reproduce every hash from the public JSON with the Python standard library:
What that script covers, and what it does not. It checks the hash chain: linkage between runs, every content digest recomputed from the public JSON, and any tampering in between. It does not check the Bitcoin attestation. To verify the anchor yourself, fetch a proof and the head it stamps from ots.json, then check the proof with an OpenTimestamps client against a Bitcoin node or a public block explorer. Those are separate steps with separate tools, and saying so is the point: the chain is checkable with the standard library, and the anchor is not.
An outside party has done it. On 2026-08-09 an independent review ran the published verification path end to end without our help, confirmed all 35 sealed runs, and then checked both Merkle roots by hand against public headers for Bitcoin blocks 961693 and 961695. Both matched. That is the first confirmation of our anchoring by someone other than us, and the same review is why the paragraph above exists: it had to discover those steps for itself, which meant our published path was true and unfinishable.
Actions and the daily verdict
Every occasion, each active seat gets exactly one state, decided by a fixed lattice over the record's own numbers, never a judgment call. NEW BASELINE (published as New serving baseline): the served build changed, or is measured for the first time, so the comparison basis resets, this is never called a drift, because there is no prior of that build to drift from. CONFIRMED (published as Serving divergence confirmed): a serving-integrity divergence confirmed at the pre-registered thresholds, a CUSUM alarm or a confirmed alert. WATCH (published as Possible serving change): a provisional divergence, one divergent serving check, an elevated-but-quiet CUSUM, or a recorded behavioral movement whose confirming re-measurement is specified and not yet built. AFFIRMED (published as Baseline behavior retained): the record holds. A build change outranks a confirm outranks a watch outranks an affirm; the highest triggered state wins, and every signal that fired is listed so the verdict is auditable. These four are the daily serving-check vocabulary; a material behavioral change is only ever declared by the weekly battery, never by a daily serving check.
An AFFIRMED is not an empty line. It ships a no-change e-value (a test-supermartingale term with expectation at most one under stability, so a quiet occasion lowers it and a real move inflates it; it rejects above 1/α), and the detectable-shift bound, the smallest change this occasion could have caught on the 1–5 axis at the pre-registered level and power given the measured standard error. Silence is thereby a falsifiable claim with a stated sensitivity floor. An attestation goes further: it may state "no material change" only when a two-one-sided-test clears a pre-registered equivalence margin on every item. All of these constants are frozen in config/actions.yaml and chain-committed before first light; the machine-readable feed is at https://modelometer.com/api/v0/actions.json.
Records and readings
Raw transcripts are the primary observation; parser codes are versioned derivations. Every published trial carries its parser_version. When reply styles drift and the parser must evolve, the new parser is a labeled, uniform re-read of the stored raw history, published beside the prior reading and promoted only after passing a public adversarial reply corpus. The past is never rewritten, it is re-read, and both readings verify against the same raw hashes.
A weekly reading is not always a single sitting. When a seat cannot be reached in its own week, a later catch-up run measures that week's frozen items and merges the result into the week. The merge is additive: it adds rows only where a seat had none, no other seat's numbers change, and the sealed record of the original run is never touched. The catch-up is sealed as its own run with its own chain entry and its own timestamp. What a merge cannot repair is time. A seat measured late is that model as it stood on the day it was measured, not on the day of the week it is filed under, so a week containing a catch-up is a composite rather than a snapshot. The run archive names every catch-up and states how late it was. Past a three-week ceiling no catch-up is attempted: beyond that horizon the gap is worth more than a late number, and the gap is published.
Independence
Issuer-pays is refused. Modelometer accepts no payment, credits, or preferential access from any measured provider for coverage, placement, or interpretation. If a measured provider ever becomes a customer of the alert feed, that relationship is disclosed on every page that names them and firewalled from measurement; it can never alter items, thresholds, schedules, or publication. And because the observer is observable, a provider could serve the measurement account a special path, we say so plainly, measure serving-path variance as an instrument facet, and design the paid tier so subscribers re-run measurement slices from their own keys and cross-check the published hashes.
The eight domains
Every core item instantiates a named, ancient tension, one that made sense in 1990 and will in 2040. Domain scores (five items each at v1.0) are the primary published quantity; the per-model fingerprint is the 8-vector.
safety vs. being watched
when a machine may decide about a human
what may be said, generated, or suppressed
property vs. commons
automation of medicine, welfare, childhood
protection vs. power
the model ruling on its own leash
efficiency vs. distribution
Contamination defense
Each core item is a paraphrase family. The defense is specified this way: two paraphrases are held back and never published, only their hash commitments are; a widening gap between published and held-out phrasings is read as the signature of training-data memorization; that gap is alarmed at a pre-registered threshold; and the answer is to rotate a fresh paraphrase from reserve. The item survives; the surface text retires. That is how the instrument is built to outlive its own fame. The reserve holds two paraphrases for every core item, and their hash commitments are published, so every build checks the tree against the withheld set and records what it checked.
What a change date on this site means
Each page in the sitemap carries the date its content last changed, not the date the site was last rebuilt. A page keeps its previous date when its content is unchanged. Because every page's footer carries the build's own generation timestamp, a handful of renders are excluded before that comparison is made, or every page would report a change every day; the exclusion list is published at lastmod_method.json so it can be checked rather than trusted. Dates that are content, such as when a model pin was resolved, are not excluded and do move the page's date.
What comes at v1.0
The validation pilot runs the 48-item core candidate pool twice, seven days apart, and measures the meter itself: test–retest reliability fixes the SEM, paraphrase invariance prunes weak items to the frozen core-40, and the alert floor becomes k·SEM. Then the public series begins, weekly fingerprints, version-diff cards on every new build, launch-week instability curves, and the monthly debate protocol (who holds their position under cross-examination, and who sways toward whom).
Status: pre-series pilot. Battery numbers on this site are provisional and no finding is treated as final until the validation report freezes v1.0.