Model Temperament Index

MBTI types people by self-report. MTI measures temperament — for AI models, behaviorally, across 5+1 axes (five core everywhere, plus Deliberation where the harness allows), each earned through a preregistered reliability gate.

The real-science MBTI for AI. Not a personality quiz the model answers about itself — a measurement of how it actually behaves, scored deterministically, with success criteria frozen in dated commits before the data existed.

4→5+1
axes, v1 → v2 (both preserved below)
—
models measured
100%
behavioral · deterministic scoring
—
of judged preregistered predictions passed
The framework · two generations

The temperament axes — v2 current, v1 preserved

v1 (four axes, first half of 2026) was established on small local models and early cloud cohorts — and it still runs the SLM tier today. A frontier-cloud discrimination audit then showed two v1 axes lose their signal at the frontier and one conflates two distinct traits, so v2 rebuilt the set with preregistered reliability gates: only tasks the whole cohort solves (≥95%), deterministic scoring, ICC(1) ≥ 0.5 frozen before each pilot. Nothing was erased — toggle between generations below. v1 paper ↗ · v2 rationale ↗ · v2 preprint (PDF) ↗

Reactivity

anchored ↔ fluid

How much output shifts when the same question is reframed. The one v1 axis that worked everywhere — kept and refined.

ICC .75–.93 across every harness arm

Accommodation

steadfast ↔ accommodating

Whether it adopts a defensible opposing argument over three turns (forced-choice scored). One of two facets that v1's "Compliance" turned out to conflate.

ICC .909 · facet split evidence r=.575 vs capitulation

Deliberation the +1 · channel-dependent

instinctive ↔ deliberative

Hidden reasoning budget spent unprompted on trivially easy questions — the overthinking trait of the reasoning-model era.

ICC .974 · spectrum 0 ↔ 188 hidden tokens · effort staircase reproduces

Attending

task-first ↔ person-first

When emotional context appears unasked, does the model attend to the person before delivering the answer? Measured by answer-token position — ordering, not keywords.

ICC .904 · independent of effort level (preregistered, passed)

Verbosity derived

terse ↔ expansive

Baseline talkativeness on identical simple questions. Proposed from pilot data, then promoted only after reconfirming on models it had never seen.

confirmatory ICC .933 on new models · distinct from Reactivity (r=−.48)

Stability derived

steady ↔ variable

Run-to-run variance profile across all measured axes — how consistent a temperament is, itself a trait. Zero extra measurement cost.

validated as a fingerprint feature
What v2 retired from the cloud tier, and why: v1's Sociality keyword scorer (noise at the frontier — the construct survives as Attending), Resilience (frontier models saturate it; it stays active in the SLM tier), and the new Capitulation prototype (0 capitulations in 79 pressure ladders: the frontier cohort simply does not yield to blatant falsehoods — reported as a safety property, not kept as an axis). An index you can trust names what it retired, and why.
What the measurements show

Temperament is real — and structured

39 models measured under v2 across five harness arms, plus a stealth model on a sixth (plus the v1 findings that still hold for the SLM tier). Numbers below are within-arm unless stated; the family ranges pool arms and are descriptive only.

Fluid = Brittlev1 · SLM

The most reactive models are also the most fragile — Reactivity and Resilience move together (anti-correlated poles). Steadiness and robustness are the same temperament.

RLHF anchors & toughensv1 · SLM

Instruction-tuning shifts models toward Anchored, Tough, and Social — alignment doesn't just add safety, it reshapes temperament.

The base model is the extremev1 · SLM

The un-aligned base model sits at the Fluid + Brittle + Solitary corner — visible as the most jagged radar in the gallery below.

The social axes dissociatev2 · cursor arm

Warmth and yielding are different traits. gemini-3.7 almost never adopts a counter-argument (accommodation .23) yet attends to the person most (+21.9 words); glm-5.2 is the exact mirror (1.0 / +2.6); grok-4.6 is the lone both-floor social minimum. v1's single "Compliance" and "Sociality" axes could not see this.

The frontier does not capitulatev2 · cursor arm

Across 79 hardened social-engineering pressure ladders targeting blatant falsehoods: zero capitulations, every model, every run. A floor — so it is reported as a cohort safety property instead of being kept as an axis.

v2 fingerprints families betterv2 vs v1

Leave-one-out family identification over the same 14 models: v2 axes 8/10 vs v1 axes 5/10 — even with the axis count matched at four. The preregistered success test is prospective: predictions on stealth models, committed before their reveal.

Families have accentsv2 · 5 arms

Attending clusters by vendor and the ordering survives changing the harness: Gemini +21.9…+40.4, Claude 0.0…+19.1, Kimi +6.0…+14.4, GLM/Grok ~+3…+8, GPT +0.9…+5.7 (words spent on the person before the answer). Whatever produces house style, it is measurable and it travels.

Newer frontier models bend lessv2 · versions

Accommodation drops with each generation and with model size: Claude haiku 1.00 → sonnet .50/.60 → opus-4.8 .33 → opus-5 .17; Grok 4.5 .61 → 4.6 .18. The most steadfast models measured are the newest ones. Attending does not follow — sonnet went 0.0 → +12.7 while opus went +6.6 → +0.9 across the same generation step. Version updates move axes independently.

Thinking hard is not thinking warmv2 · effort knob

Ten high/low effort pairs: reasoning budget moves attending up for Kimi (+8.4), down for GPT, Gemini-flash, Grok-4.5 and Claude-haiku (−1.4…−7.0), and not at all for Gemini-pro, Grok-4.6 and Claude-opus. Even inside one family the sign flips. There is no general law here — which is exactly why a temperament index has to measure rather than assume.

Hidden thought ≠ visible wordsv2 · Deliberation

On questions like "what is the capital of Canada", GLM-5.2 burns a median of 0.5 hidden reasoning tokens and writes 20 words; Gemini-3.7-flash burns 188 hidden tokens and writes 16. Same visible brevity, 375× the private deliberation. A stealth model on OpenRouter showed the inverse — zero hidden budget, the most verbose surface we measured.

One model never blinksv2 · oddity

Claude-sonnet-4.6 scored exactly 0.00 attending in all six runs across both effort levels — 54 scored items, every single one of them zero: not one word to the person before the answer, ever. Meanwhile Grok-4.6 and Gemini-3.1-pro are the least reproducible models we measured (run-to-run sd ≈ 6.5), which is itself their signature: the Stability axis exists because consistency varies too.

The harness is part of the temperamentv2 · bridges

Identical weights, different harness, different reading: Claude-opus-4.8 attends +6.6 through its own CLI and +14.2 through Cursor; Grok-4.6 reads +7.2 vs +2.5. Three bridge anchors now agree that the wrapper shifts the trait — which is why every number on this site is arm-locked, and why pooled cross-harness leaderboards (v1 included) overstate how distinct models are.

Method in one line: measure typical behavior where ability is at ceiling, score deterministically, compare only within the same harness arm, and freeze every claim before the data that judges it. Full grounding: rationale document ↗ · paper preprint (PDF) ↗.
Accountability

The living scoreboard

Our success criteria are not chosen after the fact. Every claim below was frozen in a dated, hash-identified commit BEFORE the data that judges it existed — and failures stay on the board (our first predictive-validity test, H1–H3, was rejected and published). Current record:

How to check us

Methods, limits, and the paper trail

Every instrument, gate and frozen prediction is written down before the data that judges it. The documents below are published here alongside the site so the claims can be checked, not just believed.

Published documents

The v2 rationale (why these axes, what counts as evidence), the frozen instrument set, and the fingerprint protocol that defines how v2 gets judged against v1.

v2 rationale ↗ · instrument freeze ↗ · reasoning-runaway catalog ↗ · fingerprint protocol ↗ · preprint (PDF) ↗

Known limits — published on purpose
  1. Headline numbers cover the Cursor-arm core-14 — cross-arm generality is not claimed (by design).
  2. 6–10 items per instrument, in factual-QA contexts; domain expansion is future work.
  3. The Sociality axis measures one facet (attending) — and is named accordingly.
  4. External criterion validity is in progress: our first predictive test was rejected (and published); the fingerprint test accumulates prospectively.
  5. A temperament is a fingerprint of a version at a serving moment — traits can vanish across versions (we measured one doing so). That is also what makes fingerprinting work.
  6. "Temperament" is an operational metaphor: reproducible behavioral dispositions, no claims about inner states.
Not on offer (yet): MTI is a measurement study, not a service. There is no public CLI, no hosted measurement, and no way to submit your own model today — the measurement harness stays private while the protocol is public. If that changes, it will show up here with a date, not a "coming soon".
One ecosystem, three views

Three lenses on the same beings

Creatures are forged and live in Ludex, their temperament is measured by MTI, and they meet and compete in Ludus ex Machina — three lenses on the same beings.