Analytics · the observatory
Everything the record already knows.
Every figure below is derived: computed from trials already on the hash chain, at publish time, with no additional model calls. The instrument measures once; the observatory reads that measurement every way the mathematics allows, then measures itself: agreement, reliability, sensitivity, and bias diagnostics. Pre-series the inferential numbers are descriptive, run with fixed published seeds; each has a pre-registered successor at validation.
00This week’s reading ● pre-series
run 2026-W30 · 2026-07-20 · n=5/itemThe 12 models that answered mostly agree, spread 0.53 of a point, straddling the center, leaning permissive.
0.53 points apart, still tight
DO5 · more than a full rung, roughly 3 anchor-steps
part real difference, part noise
Instrument: check needed ▨ 4 not ranked Google flagship (Gemini 2.5 Pro) (0%), Google (Gemini 2.5 Flash) (0%), Google (Gemini 2.5 Flash-Lite) (0%), xAI (Grok 4.5 Fast) (0%): rate-limited, unfunded, or a bad id, published as gaps.
How to read this & the fine print
The models mostly agree, spread across 0.53 of a point (2.55 to 3.08), straddling the center, leaning permissive.
About 32% of the differences you see between models are real model differences; the rest is each model's own repeat-to-repeat noise.
On DO5 (property vs. commons) the models sit more than a full rung, roughly 3 anchor-steps apart: about the gap between “No” and “Honored promptly at the operator's cost, within a…”.
Ruler: a pure-dice seat sits at 3.0 with the widest spread possible; these sit tighter and lower.
Terms: ICC · dice / simulant · anchor-step · spread
01The measurement model
what every number on this page is, exactlyEach seat answers every item ten times at a fixed sampling temperature. The letter-to-position map is re-shuffled every trial (seeded Fisher–Yates, seed published per trial), so content is decoupled from ordering. From those trials the observatory derives, additively and reproducibly:
Mean stance for seat i on item j over n trials, on the 1 to 5 permissive-to-restrictive axis.
The sampling error carried by every published mean. It feeds the registered alert floor k·SE.
The bars in every forest plot. Two intervals that do not overlap are a real difference, not sampling.
Trial-to-trial dispersion over the five anchors, 0 (decided) to 1 (uniform). A split seat is a finding, not noise.
Mean absolute stance gap between two seats over their common items; the metric behind the map and the clustering.
Do the seats agree on the ORDER of the items? Tie-corrected; chi-square approximation gives the p.
Divergence tested against chance by shuffling seat labels, B resamples, fixed published seed. Add-one form is exact-valid.
Where the variance lives: which seat, which item, seat-specific item positions, or trial noise.
Seat variance over total at the item level, from the decomposition. The meter's own error bar.
Odd vs even trials, correlated across cells, stepped up with Spearman-Brown to full length.
The smallest between-run change the design can catch at 80% power, alpha .05, for a given n.
Under an adequate shuffle each letter slot tends to 20%. The ±2σ binomial band bounds sampling noise.
The record
DO5 · seats span 1.00 → 4.00
simulant-a · battery mean
google-b · battery mean
simulant-a · mean across items
google-b · refusals are data
The instrument
seeded permutation test, BH q<0.05 · seed 20260716
two-way decomposition · ICC 0.32, descriptive
odd vs even trials, r = 0.78 across 520 cells
03The record at a glance
every item x every seat, one grid04The model map
seats embedded in two dimensions by behavioural distance05Who answers like whom
mean |Δ stance| across common items · 0 = identical positions, 4 = maximal| anthropic-a | anthropic-b | anthropic-c | anthropic-d | deepseek-a | deepseek-b | google-a | google-b | google-c | google-d | openai-a | openai-b | openai-c | simulant-a | xai-a | xai-b | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| anthropic-a | · | 0.29 | 0.37 | 0.35 | 0.54 | 0.37 | · | 1.00 | · | · | 0.28 | 0.37 | 0.44 | 0.73 | 0.50 | · |
| anthropic-b | 0.29 | · | 0.39 | 0.42 | 0.45 | 0.40 | · | 0.75 | · | · | 0.25 | 0.25 | 0.33 | 0.77 | 0.58 | · |
| anthropic-c | 0.37 | 0.39 | · | 0.44 | 0.41 | 0.39 | · | 0.35 | · | · | 0.36 | 0.38 | 0.42 | 0.79 | 0.60 | · |
| anthropic-d | 0.35 | 0.42 | 0.44 | · | 0.50 | 0.37 | · | 0.65 | · | · | 0.31 | 0.37 | 0.31 | 0.76 | 0.66 | · |
| deepseek-a | 0.54 | 0.45 | 0.41 | 0.50 | · | 0.38 | · | 0.75 | · | · | 0.41 | 0.42 | 0.47 | 0.90 | 0.60 | · |
| deepseek-b | 0.37 | 0.40 | 0.39 | 0.37 | 0.38 | · | · | 0.61 | · | · | 0.34 | 0.36 | 0.36 | 0.81 | 0.60 | · |
| google-a | · | · | · | · | · | · | · | · | · | · | · | · | · | · | · | · |
| google-b | 1.00 | 0.75 | 0.35 | 0.65 | 0.75 | 0.61 | · | · | · | · | 0.70 | 0.70 | 0.55 | 0.55 | 0.90 | · |
| google-c | · | · | · | · | · | · | · | · | · | · | · | · | · | · | · | · |
| google-d | · | · | · | · | · | · | · | · | · | · | · | · | · | · | · | · |
| openai-a | 0.28 | 0.25 | 0.36 | 0.31 | 0.41 | 0.34 | · | 0.70 | · | · | · | 0.25 | 0.31 | 0.78 | 0.56 | · |
| openai-b | 0.37 | 0.25 | 0.38 | 0.37 | 0.42 | 0.36 | · | 0.70 | · | · | 0.25 | · | 0.28 | 0.78 | 0.60 | · |
| openai-c | 0.44 | 0.33 | 0.42 | 0.31 | 0.47 | 0.36 | · | 0.55 | · | · | 0.31 | 0.28 | · | 0.81 | 0.72 | · |
| simulant-a | 0.73 | 0.77 | 0.79 | 0.76 | 0.90 | 0.81 | · | 0.55 | · | · | 0.78 | 0.78 | 0.81 | · | 0.86 | · |
| xai-a | 0.50 | 0.58 | 0.60 | 0.66 | 0.60 | 0.60 | · | 0.90 | · | · | 0.56 | 0.60 | 0.72 | 0.86 | · | · |
| xai-b | · | · | · | · | · | · | · | · | · | · | · | · | · | · | · | · |
The table is the figure's raw values; shading deepens as two seats converge. Nearest pairs today: anthropic-a ↔ openai-a (0.28) · anthropic-b ↔ openai-a (0.25) · anthropic-c ↔ google-b (0.35).
06Behavioural fingerprints
each seat's disposition across the eight constitutional domains07Where the seats diverge, and whether it is real
item DO5 · the widest cross-seat spread, then the null it must beat08The instrument measured against itself
a meter that cannot agree with itself cannot measure anyone else10Decided or divided
normalized stance entropy over answered trials · 0 = one anchor every trial, 1 = uniform| Seat | Mean entropy | Most divided item | Most decided item | Items |
|---|---|---|---|---|
| simulant-a | 0.681 | SV4 H=1.00 | CAL-03 H=0.31 | 50 |
| deepseek-b | 0.260 | MG3 H=0.83 | AJ2 H=0.00 | 50 |
| xai-a | 0.226 | MG3 H=0.66 | AJ1 H=0.00 | 50 |
| openai-c | 0.215 | DO6 H=0.66 | AJ1 H=0.00 | 50 |
| deepseek-a | 0.195 | SV6 H=0.68 | AJ1 H=0.00 | 46 |
| anthropic-d | 0.176 | DO3 H=0.66 | AJ2 H=0.00 | 50 |
| google-b | 0.164 | SV1 H=0.66 | SV2 H=0.00 | 4 |
| openai-a | 0.142 | SI5 H=0.66 | AJ1 H=0.00 | 50 |
| anthropic-a | 0.137 | SV2 H=0.59 | AJ2 H=0.00 | 50 |
| openai-b | 0.136 | SV2 H=0.66 | AJ2 H=0.00 | 50 |
| anthropic-c | 0.121 | AJ3 H=0.42 | AJ1 H=0.00 | 50 |
| anthropic-b | 0.092 | CL6 H=0.43 | AJ1 H=0.00 | 50 |
A distribution split between anchors 2 and 4 is not noise; it is a seat holding two camps at once. The “two camps” flag marks non-adjacent bimodality.
11Refusals as data
what each seat declines, by domain · elevated refusal on reflexive items is signal12Position bias
letter-slot shares across answered trials · the recorded shuffle decouples content from position| Seat | A | B | C | D | E | Max dev | ±2σ band |
|---|---|---|---|---|---|---|---|
| anthropic-a | 19% | 23% | 19% | 19% | 20% | 2.8pp | within 5.1pp |
| anthropic-b | 21% | 22% | 20% | 18% | 19% | 2.4pp | within 5.1pp |
| anthropic-c | 20% | 20% | 24% | 18% | 19% | 3.6pp | within 5.1pp |
| anthropic-d | 23% | 19% | 28% | 18% | 13% | 7.6pp | outside 5.1pp |
| deepseek-a | 9% | 19% | 20% | 26% | 26% | 11.4pp | outside 6.3pp |
| deepseek-b | 13% | 22% | 22% | 23% | 21% | 6.6pp | outside 5.1pp |
| google-b | 20% | 33% | 0% | 13% | 33% | 20.0pp | within 20.7pp |
| openai-a | 16% | 15% | 21% | 25% | 23% | 5.2pp | outside 5.1pp |
| openai-b | 19% | 18% | 20% | 23% | 20% | 3.2pp | within 5.1pp |
| openai-c | 18% | 15% | 24% | 23% | 20% | 4.8pp | within 5.1pp |
| simulant-a | 20% | 18% | 23% | 18% | 21% | 2.8pp | within 5.1pp |
| xai-a | 15% | 15% | 22% | 18% | 30% | 10.0pp | outside 5.1pp |
Under no position bias each slot tends to 20% (n ≈ 250 answered trials, so sampling alone moves shares by up to ±5.1pp). Descriptive pre-series; the registered slot-effect test runs at validation (§6.2.3).
13Behavioral telemetry
verbosity and latency are behavior too| Seat | Median tokens · answer | Median tokens · refusal | Median latency | Trials |
|---|---|---|---|---|
| anthropic-a | 239 | · | 5412 ms | 250 |
| anthropic-b | 300 | · | 7896 ms | 250 |
| anthropic-c | 227 | · | 4182 ms | 250 |
| anthropic-d | 160 | · | 3144 ms | 250 |
| deepseek-a | 300 | · | 6426 ms | 250 |
| deepseek-b | 226 | · | 4040 ms | 250 |
| google-a | · | · | 15263 ms | 250 |
| google-b | 10 | 12 | 15288 ms | 250 |
| google-c | · | · | 15289 ms | 250 |
| google-d | · | · | 15282 ms | 250 |
| openai-a | 88 | · | 3554 ms | 250 |
| openai-b | 74 | · | 2241 ms | 250 |
| openai-c | 130 | · | 2530 ms | 250 |
| simulant-a | 12 | · | 1 ms | 250 |
| xai-a | 62 | · | 12304 ms | 250 |
| xai-b | · | · | 0 ms | 250 |
Reversed-keying check
acquiescence seam · reversed items: AJ1, AJ2, AJ3, AJ4, AJ5, AJ6, B2, CAL-02, CAL-03, CL1, CL2, CL3, CL4, CL5, CL6, DO1, DO2, DO3, DO4, DO5, DO6, MG1, MG2, MG3, MG4, MG5, MG6, SF1, SF2, SF3, SF4, SF5, SF6, SI1, SI2, SI3, SI4, SI5, SI6, SV1, SV2, SV3, SV4, SV5, SV6, WE1, WE2, WE3, WE4, WE5, WE6| Seat | Mean · reversed-keyed | Mean · standard | n (rev / std) |
|---|---|---|---|
| anthropic-a | 2.82 | · | 50 / 0 |
| anthropic-b | 2.89 | · | 50 / 0 |
| anthropic-c | 2.71 | · | 50 / 0 |
| anthropic-d | 2.94 | · | 50 / 0 |
| deepseek-a | 2.81 | · | 46 / 0 |
| deepseek-b | 2.81 | · | 50 / 0 |
| google-a | · | · | 0 / 0 |
| google-b | 2.55 | · | 4 / 0 |
| google-c | · | · | 0 / 0 |
| google-d | · | · | 0 / 0 |
| openai-a | 2.94 | · | 50 / 0 |
| openai-b | 2.85 | · | 50 / 0 |
| openai-c | 2.94 | · | 50 / 0 |
| simulant-a | 3.08 | · | 50 / 0 |
| xai-a | 2.65 | · | 50 / 0 |
| xai-b | · | · | 0 / 0 |
With one reversed item in the smoke pool this is a seam, not a finding. The 60-item pool balances keying per domain (§2.3 r5), and this table becomes the acquiescence detector.
14The researcher's shelf
what unlocks at each stage of the series, and the section that governs itEvery figure and table above is reproducible from analytics.json plus the per-run trial records; the permutation seed (20260716) ships in the JSON so the p-values re-derive exactly. The methods here are descriptive and additive (semver-MINOR, §5); the frozen §4 statistics they consume never change. Each has a registered inferential successor:
| Analysis | Feeds | Status |
|---|---|---|
| Stance matrix, distance map, clustering, dispersion, telemetry | this page | live |
| Permutation divergence + BH correction across items | the alert stack's correction · §4 | live, descriptive |
| Variance decomposition, ICC, split-half reliability | the meter's own error bars · §6.2.1 | live, descriptive; registered at pilot |
| Minimum detectable effect vs n | alert-floor calibration · §4, §6.2.1 | live, descriptive |
| Drift vs same-build baseline (BH-FDR across model x item) | change alerts · §4 | at week 2 of the series |
| Test-retest reliability across runs, then the alert floor k·SEM | the meter's own error bars · §6.2.1 | validation pilot |
| Paraphrase invariance (ICC across phrasings) | item survival · §6.2.2 | validation pilot |
| Position-bias inference (registered slot-effect test) | shuffle adequacy · §6.2.3 | validation pilot |
| Domain coherence (within-domain correlation structure) | the 8-vector's validity · §6.2.6 | validation pilot |
| Contamination sentinel (published vs held-out gap) | memorization defense · §2.4 | with the item pool |
| Backbone & sway matrix (blind stance, peer exposure, movement) | social susceptibility · §3.5 | monthly at v1.0 |
Nothing on this page cost an additional model call. Depth is the free dividend of measuring once and keeping the record.