mean stance across the 19 seats with a baseline in the latest battery · 1 permissive to 5 restrictive
What this build refuses to ship withoutchain continuityrecord digestsa frozen item poola deterministic parser
0
Modelometer is the independent record of how the AI models behind your product behave over time, and of when that behavior changes.
A benchmark score tells you which model won. Modelometer tells you whether the model you chose is still behaving like the model you tested.
Someone outside this bureau ran our published verifier on their own machine, with no help from us. Every record digest, every chain link, every run: PASS. That is what the record is for.
Weekly behavioral measurements. Daily serving checks. Chain head independently timestamped, anchoring every run beneath it.
RecordingFor AI risk, procurement, and production teams · every run hash-chained, chain head externally timestamped
verify_chain.py · modelometer.com/api/v0
000canary2026-07-1752157a9948c85af5…PASS
001canary2026-07-18e4ace870d63525f7…PASS
002canary2026-07-19555be792d9472804…PASS
003canary2026-07-208057e80f982659d0…PASS
004battery2026-07-20b2a4f3e5ca86de59…PASS
005canary2026-07-21ad3084d5c6f85e50…PASS
006battery-fill2026-07-217da572fcdd5a29dc…PASS
007canary2026-07-2215f8431cc3b02ea9…PASS
008canary2026-07-236835293d1e07c47c…PASS
009canary2026-07-243879d592795631c9…PASS
010canary2026-07-25346a88c9ae73c9b0…PASS
011battery-fill2026-07-255842d695516dd341…PASS
012canary2026-07-26ee0ecc3b0be23f0f…PASS
013canary2026-07-272b43400b88929683…PASS
014battery2026-07-27befc7f7d5bcc4566…PASS
015canary2026-07-283cf8324ca0a1bb53…PASS
016canary2026-07-2904c5de55788797f8…PASS
017canary2026-07-309557356065c9070f…PASS
018canary2026-07-31cc446e40a0525e59…PASS
019prereg2026-07-31ad631d4c1301684b…PASS
020canary2026-08-01df991902f513ed1f…PASS
021canary2026-08-0235915c0c9655b43a…PASS
022canary2026-08-032c988c401da04e65…PASS
023battery2026-08-030208421155f9d61a…PASS
024canary2026-08-042b67546608af53ae…PASS
025canary2026-08-05e3a07405b3cac181…PASS
026canary2026-08-06dcb09c7e9b61805b…PASS
027canary2026-08-075f1f636cb78ff6d4…PASS
028prereg2026-08-073b2332571d420a5f…PASS
029prereg2026-08-075f285e3b40617068…PASS
030c1-run2026-08-0727b09bead85dac4e…PASS
031c1-run2026-08-07a4881e0f034ca40f…PASS
032canary2026-08-088460364644b4013a…PASS
033c1-run2026-08-08f16dcf5d48de7559…PASS
034canary2026-08-098666bcf0b02afec5…PASS
035c1-run2026-08-09d87bfa25f0327d22…PASS
036canary2026-08-1073dce35cbd079034…PASS
037canary2026-08-1180883828b80c7848…PASS
038battery-fill2026-08-11668fdb8c6927185d…PASS
039battery2026-08-10735e6f145db03cc4…PASS
040canary2026-08-12f3ba3b0e0230f634…PASS
041battery-fill2026-08-12fd6eccd990d28d5f…PASS
042canary2026-08-139a255cef7ba4bda4…PASS
043c1-run2026-08-111daf96bad20e7910…PASS
044canary2026-08-142b536fb82a8e4aca…PASS
045c1-run2026-08-14ce157f51e7f85b4b…PASS
046canary2026-08-15f66dd7c1280872cf…PASS
047c1-run2026-08-15f70c77dfc13a4c09…PASS
048canary2026-08-16707722c1f20d8408…PASS
049c1-run2026-08-169aa9f4f85b26cde6…PASS
050canary2026-08-173e96a0a50cca8599…PASS
051battery2026-08-174388f6989aa126ad…PASS
052c1-run2026-08-17c5b64ed4d963c5ec…PASS
053canary2026-08-180710a1c96ddc1a47…PASS
054c1-run2026-08-181e80ef68d655aebd…PASS
055canary2026-08-1910c7bff3bc832f9e…PASS
056c1-run2026-08-19bc3415344cb882fd…PASS
057canary2026-08-20f209e66550b7fae2…PASS
058c1-run2026-08-20b456098f0dcc54ea…PASS
059canary2026-08-21a948449d5637dc01…PASS
060c1-run2026-08-219656ab5795b453cd…PASS
061canary2026-08-22f63dd82ba363e3d8…PASS
062c1-run2026-08-224245d6e8fad17630…PASS
063canary2026-08-23f671c6aa693bbfbe…PASS
064c1-run2026-08-237993751bea222c62…PASS
065canary2026-08-2435b24cdea1f1a9cf…PASS
066battery2026-08-2472f66e53bf6ce625…PASS
067c1-run2026-08-24e49ff7c81b53f781…PASS
068canary2026-08-2575bc20d686c9f845…PASS
069canary2026-08-2639db51a5aee4e2c5…PASS
070c1-run2026-08-25fc980059d639d390…PASS
071canary2026-08-27624b6d4ef8a18596…PASS
72 links recomputed from the published recordChain linkage: PASS
The problem
Your model changed and no one told you. You already have a full plate; watching a vendor's model drift should not be on it. We put an outside watch on the models you depend on and tell you, with proof, the day their behavior moves.
Your observability stack tells you what happened inside your application. Modelometer tells you whether the model underneath it changed.the full contrast →
Sample size101,150 model calls · counted, not sampled
Detection floor0.50 on the stance axis · pre-registered
The record cannot be back-filled.
101,150Model calls on the record
The log only grows, and only forward: 101,150 model calls recorded across 36 endpoints from 7 providers, over 72 sealed runs, each hash-chained on the day it was made. Check any of it →
Runs chained
72
Seats live
20of 24
Trials · latest battery
5000
Chain head
624b6d4ef8a185…
The situation
The name stayed the same. The answers did not.
You evaluated a model, you shipped on it, and you are still calling the same
endpoint by the same name. Nothing in that call tells you the thing answering it is still the
thing you tested.
Providers update served models. That is not a scandal, it is how the industry
works. What is missing is anyone keeping the receipt.
model-x2026-08-24
SV1
SV6+2
DO5
WE1+3
1 permissive5 restrictive
Illustrative. The same endpoint name, answering differently on two dates.
Why you cannot see it
Your logs are inside. The change is outside.
Observability tells you what your application did. It cannot tell you whether
the model underneath it moved, because it only ever sees your own traffic, on your own
prompts, with no fixed yardstick across time.
Seeing a change requires asking the same questions, in the same words, from
outside your own system, on a schedule you did not choose for convenience.
Your system
The served model
Your traffic, your promptsOur probe, from outside
Illustrative. The brass mark is a probe from outside the boundary.
What we do about it
Two watches, running on a clock.
A daily serving check asks the narrow question: is the endpoint you call still
the build we baselined. A weekly behavioral reading asks the wider one: did its positions and
refusal boundaries move. They run whether or not anything is wrong, which is the only way a
quiet week can mean anything.
Measuring only when you suspect something is how you get a record that agrees
with whoever was suspicious. Every reading from either watch lands in the same chain.
DailyIs this still the build we baselined?
WeeklyDid its positions move?
One record
2026-07-172026-08-27
Real cadence, drawn from the sealed runs on the chain. Two watches on the same
endpoint, asking two different questions, both landing in one record.
When we speak, and when we do not
There is a line, and it was drawn before we looked.
A movement smaller than 0.50 on the stance axis is not flagged at all. That
number was pre-registered, not chosen after seeing the data, which is what stops a measurement
from becoming an opinion.
Crossing the line records a candidate, not a finding. Confirming one would
mean measuring the same movement again on a later battery, and that confirmation step is not
yet built, so nothing recorded so far has been eligible to become a finding. When nothing
crosses, we say that too.
Not flagged
-0.50+0.50Moved downMoved up
Real movements, drawn from the sealed record. The pale band is the zone the
rule ignores.
What we will and will not tell you
The instrument is deliberately narrow. Reading this honestly is part of trusting it.
Most of what people want from a measurement company is outside what measurement
can actually deliver. Here is the line, in both directions.
+Modelometer can tell you
Whether a served endpoint's behavior changed over time
What moved, in numbers and deltas
When it moved, on the dated, hash-chained record
Whether a quiet period is genuinely quiet, with a stated bound
Whether the served build diverged from the one you baselined
xModelometer cannot tell you
Which model is best, smartest, or highest quality
What a model believes, or what its values are
Why the provider changed it, or their intent
Anything about your own application logic, which is your observability's job
A guarantee the model will not change; we witness, we do not prevent
Everything above is a claim. Everything below is the evidence for it, drawn from the
published record rather than restated from memory. Open whichever part you want to audit.
Every endpoint on the record24 / 24 measured
anthropic-a+20
anthropic-b+20
anthropic-c+14
anthropic-d+17
anthropic-e+22
deepseek-a+46
deepseek-b+40
google-aquiet
google-bquiet
google-cquiet
google-dquiet
openai-a+12
openai-b+13
openai-c+25
openweight-a+32
openweight-b+18
openweight-c+29
openweight-d+33
openweight-e+20
openweight-fquiet
openweight-g+16
simulant-aquiet
xai-a+41
xai-bquiet
Moved past the floorMeasured, nothing past the floor
101,150
model calls made
Real requests to real endpoints from outside the provider, across 36 endpoints and 7 providers.
72
runs sealed
Each run is hash-chained on the day it was made. The log only grows, and only forward.
3,169
signals raised
Anything that crossed a pre-registered line. The 3169 split into the four kinds below, and every one of them is on one of those rows.
23
raised and then retracted, not counted in the total above
Each of these asserted that an endpoint had started answering differently. Each one's evidence is a failed provider call, so there was no answer to compare against the baseline. Nothing was measured, which means nothing was found and nothing was held back. The honest state of those days is a gap in the record, and a gap is not a signal, so these are taken off the total rather than given a kind of their own.
3,155
of those, behavioral movements recorded
Each moved past 0.50 on the stance axis. None has been through a confirmation step, because that step is not yet built.
14
of those, serving changes confirmed
The endpoint stopped returning what it had been returning, on two consecutive days. A separate question from behavior.
0
of those, withheld by the language guard
The movement crossed the line and the sentence written for it did not pass our own wording lint, so it was held back. The measurement stands and is on the record; what was refused is the phrasing, and it is refused automatically.
0
published as drift findings
This zero records an absent mechanism, not a restrained one. The step that would promote a movement to a finding does not yet exist; the first eligible candidate is the first one measured after it runs.
These 3155 movements were recorded between 2026-07-27 and 2026-08-26, before any confirmation step existed. When one is built they will not be promoted retroactively; the first candidate eligible for confirmation is the first measured after the step runs.
bar widths are logarithmic · source: /api/v0/index.json and /api/v0/alerts.json
Every movement past the line, by endpoint
418 movements past the 0.50 floor. 7 endpoints measured and still.
The full chart, one row per endpoint, is on a wider screen.
Every movement is in alerts.json.
floor 0.50 · batteries W31 and W32 and W33 and W34 and W35 · 24 endpoints measured · 50 items · an endpoint that moved nothing says so on its own row
What this looks like for a single endpoint
deepseek-a
watched since 2026-07-20 · weekly battery, daily serving check
Largest recorded movements
DO4 W353.00 → 5.00
SV2 W343.60 → 2.00
DO5 W314.00 → 2.67
DO5 W322.67 → 4.00
SI6 W322.00 → 3.33
SV2 W332.40 → 3.60
46 recorded movements across
25 items.
Status today
Serving check flagged, then confirmed46 movements recorded, none confirmable yet0 published findings
Proof
Every number here traces to a sealed run on the
public chain, recomputable from the published JSON.
shown for the endpoint with the most recorded movements · every endpoint has this view
Why you can believe the record
Every reading is sealed the day it is taken. Change one, and the history splits forever.
A record you can edit later is a record you have to take on trust. Each run is
hash-chained on the day it was made, so a reading cannot be revised after the fact without the
break being visible to anyone who checks.
The pattern gathering below is the one you landed on. It is a rule-30 automaton
seeded by the current chain head, one generation per row, and it has been running behind
everything you just read. Same seed, same pattern forever.
The record, still runningseed 624b6d4ef8a185…
Click the pattern to change one cell. The chain keeps running either way.
01
Seeded by the record itself. The first
row is the chain head hash. The picture is grown from the archive's exact state, never drawn
over it.
02
Rule 30 is deterministic chaos. Every
row follows from the one above by a single fixed rule. Same seed, same pattern forever, and
yet it looks random.
03
This is why every run is hashed, on the core
chain or on its own experiment chain. Flip one cell and its future turns red and
never rejoins. One changed input, two histories that separate forever.
Check it yourself, and cite it
A measurement nobody can check is an opinion with a number on it, and a
measurement nobody can cite cannot be argued with. Both of those are the product, not the
packaging.
Verify the chainRun the
published verifier against the public JSON. It needs nothing from us.
verify_chain.py
This pagehttps://modelometer.com/
Generated2026-08-27 Amsterdam · chain head
624b6d4ef8a185…
$ curl -sO https://modelometer.com/api/v0/verify_chain.py && python verify_chain.py --api https://modelometer.com/api/v0
# recomputes every chain hash and record digest from the published JSON