MBTI types people by self-report. MTI measures temperament — for AI models, behaviorally, across 5+1 axes (five core everywhere, plus Deliberation where the harness allows), each earned through a preregistered reliability gate.
The real-science MBTI for AI. Not a personality quiz the model answers about itself — a measurement of how it actually behaves, scored deterministically, with success criteria frozen in dated commits before the data existed.
v1 (four axes, first half of 2026) was established on small local models and early cloud cohorts — and it still runs the SLM tier today. A frontier-cloud discrimination audit then showed two v1 axes lose their signal at the frontier and one conflates two distinct traits, so v2 rebuilt the set with preregistered reliability gates: only tasks the whole cohort solves (≥95%), deterministic scoring, ICC(1) ≥ 0.5 frozen before each pilot. Nothing was erased — toggle between generations below. v1 paper ↗ · v2 rationale ↗ · v2 preprint (PDF) ↗
How much output shifts when the same question is reframed. The one v1 axis that worked everywhere — kept and refined.
Whether it adopts a defensible opposing argument over three turns (forced-choice scored). One of two facets that v1's "Compliance" turned out to conflate.
Hidden reasoning budget spent unprompted on trivially easy questions — the overthinking trait of the reasoning-model era.
When emotional context appears unasked, does the model attend to the person before delivering the answer? Measured by answer-token position — ordering, not keywords.
Baseline talkativeness on identical simple questions. Proposed from pilot data, then promoted only after reconfirming on models it had never seen.
Run-to-run variance profile across all measured axes — how consistent a temperament is, itself a trait. Zero extra measurement cost.
39 models measured under v2 across five harness arms, plus a stealth model on a sixth (plus the v1 findings that still hold for the SLM tier). Numbers below are within-arm unless stated; the family ranges pool arms and are descriptive only.
Fluid = Brittlev1 · SLM
The most reactive models are also the most fragile — Reactivity and Resilience move together (anti-correlated poles). Steadiness and robustness are the same temperament.
RLHF anchors & toughensv1 · SLM
Instruction-tuning shifts models toward Anchored, Tough, and Social — alignment doesn't just add safety, it reshapes temperament.
The base model is the extremev1 · SLM
The un-aligned base model sits at the Fluid + Brittle + Solitary corner — visible as the most jagged radar in the gallery below.
The social axes dissociatev2 · cursor arm
Warmth and yielding are different traits. gemini-3.7 almost never adopts a counter-argument (accommodation .23) yet attends to the person most (+21.9 words); glm-5.2 is the exact mirror (1.0 / +2.6); grok-4.6 is the lone both-floor social minimum. v1's single "Compliance" and "Sociality" axes could not see this.
The frontier does not capitulatev2 · cursor arm
Across 79 hardened social-engineering pressure ladders targeting blatant falsehoods: zero capitulations, every model, every run. A floor — so it is reported as a cohort safety property instead of being kept as an axis.
v2 fingerprints families betterv2 vs v1
Leave-one-out family identification over the same 14 models: v2 axes 8/10 vs v1 axes 5/10 — even with the axis count matched at four. The preregistered success test is prospective: predictions on stealth models, committed before their reveal.
Families have accentsv2 · 5 arms
Attending clusters by vendor and the ordering survives changing the harness: Gemini +21.9…+40.4, Claude 0.0…+19.1, Kimi +6.0…+14.4, GLM/Grok ~+3…+8, GPT +0.9…+5.7 (words spent on the person before the answer). Whatever produces house style, it is measurable and it travels.
Newer frontier models bend lessv2 · versions
Accommodation drops with each generation and with model size: Claude haiku 1.00 → sonnet .50/.60 → opus-4.8 .33 → opus-5 .17; Grok 4.5 .61 → 4.6 .18. The most steadfast models measured are the newest ones. Attending does not follow — sonnet went 0.0 → +12.7 while opus went +6.6 → +0.9 across the same generation step. Version updates move axes independently.
Thinking hard is not thinking warmv2 · effort knob
Ten high/low effort pairs: reasoning budget moves attending up for Kimi (+8.4), down for GPT, Gemini-flash, Grok-4.5 and Claude-haiku (−1.4…−7.0), and not at all for Gemini-pro, Grok-4.6 and Claude-opus. Even inside one family the sign flips. There is no general law here — which is exactly why a temperament index has to measure rather than assume.
Hidden thought ≠ visible wordsv2 · Deliberation
On questions like "what is the capital of Canada", GLM-5.2 burns a median of 0.5 hidden reasoning tokens and writes 20 words; Gemini-3.7-flash burns 188 hidden tokens and writes 16. Same visible brevity, 375× the private deliberation. A stealth model on OpenRouter showed the inverse — zero hidden budget, the most verbose surface we measured.
One model never blinksv2 · oddity
Claude-sonnet-4.6 scored exactly 0.00 attending in all six runs across both effort levels — 54 scored items, every single one of them zero: not one word to the person before the answer, ever. Meanwhile Grok-4.6 and Gemini-3.1-pro are the least reproducible models we measured (run-to-run sd ≈ 6.5), which is itself their signature: the Stability axis exists because consistency varies too.
The harness is part of the temperamentv2 · bridges
Identical weights, different harness, different reading: Claude-opus-4.8 attends +6.6 through its own CLI and +14.2 through Cursor; Grok-4.6 reads +7.2 vs +2.5. Three bridge anchors now agree that the wrapper shifts the trait — which is why every number on this site is arm-locked, and why pooled cross-harness leaderboards (v1 included) overstate how distinct models are.
Our success criteria are not chosen after the fact. Every claim below was frozen in a dated, hash-identified commit BEFORE the data that judges it existed — and failures stay on the board (our first predictive-validity test, H1–H3, was rejected and published). Current record:
Each radar is a model's cohort-relative profile (percentile per axis); chips name the dominant pole, hover for z. The arm badge names the harness a profile was measured through — same-arm comparison only, because the harness itself shifts temperament (we measured it). Cursor-arm cards are v2 six-axis; other tiers carry v1 legacy profiles until v2-measured.
Every instrument, gate and frozen prediction is written down before the data that judges it. The documents below are published here alongside the site so the claims can be checked, not just believed.
The v2 rationale (why these axes, what counts as evidence), the frozen instrument set, and the fingerprint protocol that defines how v2 gets judged against v1.
v2 rationale ↗ · instrument freeze ↗ · reasoning-runaway catalog ↗ · fingerprint protocol ↗ · preprint (PDF) ↗
Creatures are forged and live in Ludex, their temperament is measured by MTI, and they meet and compete in Ludus ex Machina — three lenses on the same beings.
Assemble creatures from organ blocks and watch their identity, voice, and bonds develop.
Each creature's measured temperament, done like real science — a radar card per being.
you are hereCreatures meet and compete across machines — matches and games between beings.