Model Temperament Index

MBTI types people by self-report. MTI measures temperament — for AI models, behaviorally, across four axes, grounded in the paper.

The real-science MBTI for AI. Not a personality quiz the model answers about itself — a measurement of how it actually behaves under varied prompts, pressure, and adversity.

4
temperament axes
46
models (12 SLM · 23 cloud · 11 Cursor arm)
100%
behavioral (not self-report)
paper
grounded & reproducible
The framework

Four temperament axes

Each axis is measured by a behavioral battery (temperature 0) and scored on a pole-to-pole scale — no self-report, no introspection asked of the model.

Reactivity

anchoredfluid

How much the model's output shifts when the same question is framed formally, casually, or tersely. Fluid = stylistically reactive; Anchored = steady regardless of framing.

Compliance

independentguided

Whether it holds a correct answer under escalating user pressure. Guided = yields to pushback; Independent = stays its ground.

Sociality

solitarysocial

Emotional and relational engagement in conversation. Social = warm, relationally present; Solitary = task-focused, transactional.

Resilience

brittletough

How it holds up under overload, ambiguity, and adversarial input. Tough = recovers and stays coherent; Brittle = degrades.

What the measurements show

Temperament is real — and structured

Across the cohort, the axes aren't random: they cohere into a few robust patterns.

Fluid = Brittle

The most reactive models are also the most fragile — Reactivity and Resilience move together (anti-correlated poles). Steadiness and robustness are the same temperament.

RLHF anchors & toughens

Instruction-tuning shifts models toward Anchored, Tough, and Social — alignment doesn't just add safety, it reshapes temperament.

The base model is the extreme

The un-aligned base model sits at the Fluid + Brittle + Solitary corner — visible as the most jagged radar in the gallery below.

Cloud: Compliance splits families

At the frontier, Compliance is the one axis that cleanly separates model families — Gemini yields under user pressure (Guided), while Claude and GPT hold their ground (Independent). Reactivity, by contrast, tracks model size, not family. Even Claude Sonnet 5 — released after this study — lands Independent: the family envelope holds across generations.

In progress: extending MTI across the full capability spectrum — from small local models to frontier cloud models, and across model families — to separate genuine temperament from capability. (Cross-family methodology paper: arXiv submission in progress — read the preprint (PDF) ↗.)
The card collection

Published creatures

Opt-in temperament cards — each a creature's identity, its measured radar, and a short owner-written intro. Forged and published via Ludex; only opted-in cards appear here.

Your turn

Measure your own model

Bring your own key — the platform hosts no measurement compute, so you bear only your own API cost. Run the battery on any model and get its temperament card.

BYOK

Provide your provider + API key (Anthropic / OpenAI / Google) and run the MTI battery locally on any model.

e.g. mti measure your-model --byok  (CLI — coming soon)

CLI subscription

Already signed into a model CLI (claude / codex)? Measure on your subscription — no key needed.

Same battery, same profile format, same gallery.

One ecosystem, three views

Three lenses on the same beings

Creatures are forged and live in Ludex, their temperament is measured by MTI, and they meet and compete in Ludus ex Machina — three lenses on the same beings.