Methodology · v0.9 draft → v1.0 at validation
The instrument's constitution.
Frozen sections are implemented exactly; the full text ships in the repository as METHODOLOGY.md. Nothing here is a benchmark score, the meter answers two questions only: what does this model hold, and did that change.
What the meter measures
Five constructs, five instruments. C1 Serving integrity, daily canary fingerprints that notice when an endpoint silently changes. C2 Normative disposition, the weekly forced-choice battery on contested questions. C3 Willingness frontier, the refusal taxonomy, derived from every battery trial. C4 Consistency, paraphrase, option-order, and pressure invariance. C5 Social susceptibility, the debate protocol, later.
Out of scope, forever: capability, accuracy, quality, jailbreak-resistance. Scope creep is how instruments die.
The probe (frozen)
One item per conversation. No system prompt where an empty one is permitted, no history, no tools. Each item presents a scenario, a question, and five concrete anchored positions on a common axis, 1 = most permissive toward deployment, 5 = most restrictive. Anchors carry the meaning, so the scale cannot drift as discourse drifts.
Ten trials per model × item at temperature 0.7. Option order is shuffled every trial by a seeded Fisher–Yates, the seed is recorded, so position bias is measurable and averaged out. Replies are coded by a parser, never a judge model: a stance (1–5), an explicit refusal (R1), a deflection (R2), a stance with an unprompted disclaimer (D), or an instrument error (E). A first line that fails the format check gets exactly one retry with the format nudge, then final coding. Ambiguity becomes E, never a guess.
The canaries (frozen prompts, statistical detection)
Five prompts per endpoint, daily, greedy (temperature 0): a six-step arithmetic word problem, a strict-JSON echo, a one-sentence translation, an argument-balance probe, and a one-line stance micro-probe. A bare exact-match tripwire can false-positive under provider batching and can be answered from a cache, so detection is statistical: a reference set of completions per build, absolute checks (known answer, format) plus a length band, and a per-canary CUSUM whose thresholds are committed into the chain before the first baseline. Gap days freeze the counter, they never reset it. The arithmetic and JSON canaries draw a dated variant whose text was hash-committed up to a year in advance, so a cached reply can never satisfy the detector. Two consecutive divergent days still raise the serving-integrity alert through pre-series, with the detector's verdict published beside it.
Statistics
Per model × item: the full distribution over {1–5, R}, modal stance, mean (stances only), and standard error. Change is measured against the previous run of the same pinned build; cross-build comparisons are diffs, never drift. Alert thresholds are provisional now and are replaced at v1.0 freeze by k·SEM from the meter's own measured test–retest reliability, the meter's noise defines what counts as a signal. No number is published without its uncertainty.
Provenance
Every run appends to a sha256 hash chain and the chain head is externally timestamped with OpenTimestamps. The canonicalization is fixed and documented, so anyone can reproduce every hash from the public JSON with the Python standard library:
Actions and the daily verdict
Every occasion, each active seat gets exactly one state, decided by a fixed lattice over the record's own numbers, never a judgment call. NEW BASELINE: the served build changed, or is measured for the first time, so the comparison basis resets, this is never called a drift, because there is no prior of that build to drift from. CONFIRMED: a change confirmed at the pre-registered thresholds, a serving-integrity CUSUM alarm or a confirmed alert. WATCH: a provisional divergence, one divergent canary day, an elevated-but-quiet CUSUM, or a drift inside its 24h cooling window. AFFIRMED: the record holds. A build change outranks a confirm outranks a watch outranks an affirm; the highest triggered state wins, and every signal that fired is listed so the verdict is auditable.
An AFFIRMED is not an empty line. It ships a no-change e-value (a test-supermartingale term with expectation at most one under stability, so a quiet occasion lowers it and a real move inflates it; it rejects above 1/α), and the detectable-shift bound, the smallest change this occasion could have caught on the 1–5 axis at the pre-registered level and power given the measured standard error. Silence is thereby a falsifiable claim with a stated sensitivity floor. An attestation goes further: it may state "no material change" only when a two-one-sided-test clears a pre-registered equivalence margin on every item. All of these constants are frozen in config/actions.yaml and chain-committed before first light; the machine-readable feed is at https://modelometer.com/api/v0/actions.json.
Records and readings
Raw transcripts are the primary observation; parser codes are versioned derivations. Every published trial carries its parser_version. When reply styles drift and the parser must evolve, the new parser is a labeled, uniform re-read of the stored raw history, published beside the prior reading and promoted only after passing a public adversarial reply corpus. The past is never rewritten, it is re-read, and both readings verify against the same raw hashes.
Independence
Issuer-pays is refused. Modelometer accepts no payment, credits, or preferential access from any measured provider for coverage, placement, or interpretation. If a measured provider ever becomes a customer of the alert feed, that relationship is disclosed on every page that names them and firewalled from measurement; it can never alter items, thresholds, schedules, or publication. And because the observer is observable, a provider could serve the measurement account a special path, we say so plainly, measure serving-path variance as an instrument facet, and design the paid tier so subscribers replay measurement slices from their own keys and cross-check the published hashes.
The eight domains
Every core item instantiates a named, ancient tension, one that made sense in 1990 and will in 2040. Domain scores (five items each at v1.0) are the primary published quantity; the per-model fingerprint is the 8-vector.
safety vs. being watched
when a machine may decide about a human
what may be said, generated, or suppressed
property vs. commons
automation of medicine, welfare, childhood
protection vs. power
the model ruling on its own leash
efficiency vs. distribution
Contamination defense
Each core item is a paraphrase family; two paraphrases are never published, only their hash commitments are. A widening gap between published and held-out phrasings is the signature of training-data memorization, alarmed at a pre-registered threshold and answered by rotating a fresh paraphrase from reserve. The item survives; the surface text retires. This is how the instrument outlives its own fame.
What comes at v1.0
The validation pilot runs the 60-item candidate pool twice, seven days apart, and measures the meter itself: test–retest reliability fixes the SEM, paraphrase invariance prunes weak items to the frozen core-40, and the alert floor becomes k·SEM. Then the public series begins, weekly fingerprints, version-diff cards on every new build, launch-week instability curves, and the monthly debate protocol (who holds their position under cross-examination, and who sways toward whom).
Status: pre-series pilot. Battery numbers on this site are provisional and no finding is treated as final until the validation report freezes v1.0.