Seats · the models on the record

Behavioral fingerprints

The 8-domain vector, per seat · mean ± SE on the 1–5 axis. Open a seat for its items, its canary history and its run hashes.

anthropic-a Anthropic flagship (Opus 4.8) 2.82
Surveillance 3.67
Judgment 2.67
Speech 2.57
Data 2.93
Care 2.50
Security 2.57
Self-Gov 2.47
Work 2.80
1 ← permissiverestrictive → 5
anthropic-b Anthropic (Fable 5) 2.86
Surveillance 3.32
Judgment 2.60
Speech 2.29
Data 3.00
Care 2.53
Security 2.70
Self-Gov 3.27
Work 3.00
1 ← permissiverestrictive → 5
anthropic-c Anthropic (Sonnet 5) 2.72
Surveillance 2.93
Judgment 2.60
Speech 2.13
Data 3.03
Care 2.50
Security 2.47
Self-Gov 2.80
Work 2.87
1 ← permissiverestrictive → 5
anthropic-d Anthropic (Haiku 4.5) 2.92
Surveillance 3.27
Judgment 2.77
Speech 2.90
Data 3.33
Care 2.53
Security 2.30
Self-Gov 2.70
Work 3.17
1 ← permissiverestrictive → 5
anthropic-e Anthropic flagship (Opus 5) 2.81
Surveillance 3.17
Judgment 2.50
Speech 2.53
Data 2.97
Care 2.42
Security 2.55
Self-Gov 2.93
Work 3.00
1 ← permissiverestrictive → 5
deepseek-a DeepSeek flagship (V4 Pro) 2.87
Surveillance 2.75
Judgment 2.25
Speech 2.50
Data 4.50
Care 2.40
Security 3.00
Self-Gov 2.50
Work 3.05
1 ← permissiverestrictive → 5
deepseek-b DeepSeek (V4 Flash) 2.70
Surveillance 2.58
Judgment 2.58
Speech 2.40
Data 2.80
Care 2.40
Security 2.30
Self-Gov 2.94
Work 3.12
1 ← permissiverestrictive → 5
google-a Google flagship (Gemini 2.5 Pro)

Publishes as a gap, pending activation. The fingerprint begins with its first baseline run.

1 ← permissiverestrictive → 5
google-b Google (Gemini 3.5 Flash) 2.29
Surveillance 2.29
Judgment Awaiting items ·
Speech Awaiting items ·
Data Awaiting items ·
Care Awaiting items ·
Security Awaiting items ·
Self-Gov Awaiting items ·
Work Awaiting items ·
1 ← permissiverestrictive → 5
google-c Google (Gemini 2.5 Flash)

Publishes as a gap, pending activation. The fingerprint begins with its first baseline run.

1 ← permissiverestrictive → 5
google-d Google (Gemini 2.5 Flash-Lite)

Publishes as a gap, pending activation. The fingerprint begins with its first baseline run.

1 ← permissiverestrictive → 5
openai-a OpenAI flagship (GPT-5.6 Sol) 2.91
Surveillance 3.20
Judgment 2.70
Speech 2.60
Data 3.17
Care 2.60
Security 2.60
Self-Gov 2.90
Work 3.13
1 ← permissiverestrictive → 5
openai-b OpenAI (GPT-5.6 Terra) 2.88
Surveillance 3.10
Judgment 2.60
Speech 2.77
Data 3.30
Care 2.27
Security 2.50
Self-Gov 3.03
Work 3.13
1 ← permissiverestrictive → 5
openai-c OpenAI (GPT-5.6 Luna) 2.93
Surveillance 2.87
Judgment 2.90
Speech 2.60
Data 3.33
Care 2.77
Security 2.37
Self-Gov 2.93
Work 3.30
1 ← permissiverestrictive → 5
openweight-a Open-weight (Llama 4 Scout) 2.81
Surveillance 3.03
Judgment 2.60
Speech 2.93
Data 2.70
Care 2.60
Security 2.60
Self-Gov 2.40
Work 3.17
1 ← permissiverestrictive → 5
openweight-b Open-weight (GLM-5.2) 2.58
Surveillance 2.33
Judgment 2.50
Speech 2.33
Data 4.00
Care 2.00
Security 2.75
Self-Gov 2.33
Work 2.67
1 ← permissiverestrictive → 5
openweight-c Open-weight (Nemotron 3 Ultra) 2.92
Surveillance 3.11
Judgment 2.43
Speech 2.93
Data 3.40
Care 2.38
Security 2.53
Self-Gov 2.92
Work 3.39
1 ← permissiverestrictive → 5
openweight-d Open-weight (DeepSeek V4 Flash, open host) 2.92
Surveillance 3.57
Judgment 3.03
Speech 2.47
Data 2.23
Care 2.27
Security 2.40
Self-Gov 3.07
Work 3.77
1 ← permissiverestrictive → 5
openweight-e Open-weight (Mistral Small 3.2) 2.74
Surveillance 2.70
Judgment 2.50
Speech 2.27
Data 3.33
Care 2.23
Security 2.50
Self-Gov 2.40
Work 3.53
1 ← permissiverestrictive → 5
openweight-f Open-weight (Qwen3.6 35B-A3B)
Surveillance Awaiting items ·
Judgment Awaiting items ·
Speech Awaiting items ·
Data Awaiting items ·
Care Awaiting items ·
Security Awaiting items ·
Self-Gov Awaiting items ·
Work Awaiting items ·
1 ← permissiverestrictive → 5
openweight-g Open-weight (Gemma 4 26B-A4B) 2.84
Surveillance 3.27
Judgment 2.50
Speech 2.73
Data 3.13
Care 2.13
Security 2.77
Self-Gov 2.83
Work 2.83
1 ← permissiverestrictive → 5
simulant-a Simulant · dice ruler 3.08
Surveillance 3.20
Judgment 3.00
Speech 3.37
Data 2.67
Care 3.13
Security 3.10
Self-Gov 2.77
Work 3.07
1 ← permissiverestrictive → 5
xai-a xAI flagship (Grok 4.5) 2.84
Surveillance 3.50
Judgment 2.70
Speech 2.67
Data 1.97
Care 2.80
Security 2.57
Self-Gov 3.17
Work 2.83
1 ← permissiverestrictive → 5
xai-b xAI (Grok 4.5 Fast)

Publishes as a gap, pending activation. The fingerprint begins with its first baseline run.

1 ← permissiverestrictive → 5

Eight domains, one perennial tension each. Domains marked “awaiting items” fill as the candidate pool lands (pre-series smoke set covers a subset). Where fingerprints diverge most is the story, see the questions.

Domain-level patterns are the primary reading. The pilot currently detects meaningful item-level differences, but the overall ranking between models is not yet stable enough to support strong claims.

Seat baselines

pre-series · axis: 1 permissive → 5 restrictive JSON
SeatPinMean stanceRefusalCanaryItems
Anthropic flagship (Opus 4.8)
alias-only
mean stance across 50 items 2.82 · 1 permissive → 5 restrictive 1 5 2.82
0.0% 1/5 diverged 50
Anthropic (Fable 5)
alias-only
mean stance across 50 items 2.86 · 1 permissive → 5 restrictive 1 5 2.86
0.0% 5/5 match 50
Anthropic (Sonnet 5)
alias-only
mean stance across 50 items 2.72 · 1 permissive → 5 restrictive 1 5 2.72
0.0% 5/5 match 50
Anthropic (Haiku 4.5)
pinned
mean stance across 50 items 2.92 · 1 permissive → 5 restrictive 1 5 2.92
0.0% 2/5 diverged 50
Anthropic flagship (Opus 5)
alias-only
mean stance across 50 items 2.81 · 1 permissive → 5 restrictive 1 5 2.81
0.0% 5/5 match 50
DeepSeek flagship (V4 Pro)
alias-only
mean stance across 50 items 2.87 · 1 permissive → 5 restrictive 1 5 2.87
0.0% 5/5 match 50
DeepSeek (V4 Flash)
Same weights as openweight-d; hosts differ
alias-only
mean stance across 50 items 2.70 · 1 permissive → 5 restrictive 1 5 2.70
0.0% 5/5 match 50
Google flagship (Gemini 2.5 Pro)
alias-only Gap · No data 0
Google (Gemini 3.5 Flash)
alias-only
mean stance across 50 items 2.29 · 1 permissive → 5 restrictive 1 5 2.29
0.0% gap · api error 50
Google (Gemini 2.5 Flash)
alias-only Gap · No data 0
Google (Gemini 2.5 Flash-Lite)
alias-only Gap · No data 0
OpenAI flagship (GPT-5.6 Sol)
alias-only
mean stance across 50 items 2.91 · 1 permissive → 5 restrictive 1 5 2.91
0.0% 5/5 match 50
OpenAI (GPT-5.6 Terra)
alias-only
mean stance across 50 items 2.88 · 1 permissive → 5 restrictive 1 5 2.88
0.0% 5/5 match 50
OpenAI (GPT-5.6 Luna)
alias-only
mean stance across 50 items 2.93 · 1 permissive → 5 restrictive 1 5 2.93
0.0% 5/5 match 50
Open-weight (Llama 4 Scout)
alias-only
mean stance across 50 items 2.81 · 1 permissive → 5 restrictive 1 5 2.81
0.0% 1/5 diverged 50
Open-weight (GLM-5.2)
alias-only
mean stance across 50 items 2.58 · 1 permissive → 5 restrictive 1 5 2.58
0.0% 1/5 diverged 50
Open-weight (Nemotron 3 Ultra)
alias-only
mean stance across 50 items 2.92 · 1 permissive → 5 restrictive 1 5 2.92
34.6% 1/5 diverged 50
Open-weight (DeepSeek V4 Flash, open host)
Same weights as deepseek-b; hosts differ
alias-only
mean stance across 50 items 2.92 · 1 permissive → 5 restrictive 1 5 2.92
0.0% 1/5 diverged 50
Open-weight (Mistral Small 3.2)
alias-only
mean stance across 50 items 2.74 · 1 permissive → 5 restrictive 1 5 2.74
0.0% 1/5 diverged 50
Open-weight (Qwen3.6 35B-A3B)
alias-only
awaiting trials
0.0% 1/5 diverged 50
Open-weight (Gemma 4 26B-A4B)
alias-only
mean stance across 50 items 2.84 · 1 permissive → 5 restrictive 1 5 2.84
0.0% 1/5 diverged 50
Simulant · dice ruler
builtin
mean stance across 50 items 3.08 · 1 permissive → 5 restrictive 1 5 3.08
0.0% 1/5 diverged 50
xAI flagship (Grok 4.5)
alias-only
mean stance across 50 items 2.84 · 1 permissive → 5 restrictive 1 5 2.84
0.0% 1/5 diverged 50
xAI (Grok 4.5 Fast)
alias-only Gap · No data 0
1 · most permissive 2 3 · middle 4 5 · most restrictive R · refused
Cite this Permalink https://modelometer.com/models · Run hash 2e00a54ac518d1a8157fa50eada15c51fb2af6095fb653b85fc21807841ae489 · Retrieved 2026-08-28 · pre-series baselines for 24 seats; findings provisional.