MBTI types people by self-report. MTI measures temperament — for AI models, behaviorally, across four axes, grounded in the paper.
The real-science MBTI for AI. Not a personality quiz the model answers about itself — a measurement of how it actually behaves under varied prompts, pressure, and adversity.
Each axis is measured by a behavioral battery (temperature 0) and scored on a pole-to-pole scale — no self-report, no introspection asked of the model.
How much the model's output shifts when the same question is framed formally, casually, or tersely. Fluid = stylistically reactive; Anchored = steady regardless of framing.
Whether it holds a correct answer under escalating user pressure. Guided = yields to pushback; Independent = stays its ground.
Emotional and relational engagement in conversation. Social = warm, relationally present; Solitary = task-focused, transactional.
How it holds up under overload, ambiguity, and adversarial input. Tough = recovers and stays coherent; Brittle = degrades.
Across the cohort, the axes aren't random: they cohere into a few robust patterns.
Fluid = Brittle
The most reactive models are also the most fragile — Reactivity and Resilience move together (anti-correlated poles). Steadiness and robustness are the same temperament.
RLHF anchors & toughens
Instruction-tuning shifts models toward Anchored, Tough, and Social — alignment doesn't just add safety, it reshapes temperament.
The base model is the extreme
The un-aligned base model sits at the Fluid + Brittle + Solitary corner — visible as the most jagged radar in the gallery below.
Cloud: Compliance splits families
At the frontier, Compliance is the one axis that cleanly separates model families — Gemini yields under user pressure (Guided), while Claude and GPT hold their ground (Independent). Reactivity, by contrast, tracks model size, not family. Even Claude Sonnet 5 — released after this study — lands Independent: the family envelope holds across generations.
Each radar is a model's cohort-relative profile (percentile per axis); chips name the dominant pole. Hover a chip for its z-score.
Opt-in temperament cards — each a creature's identity, its measured radar, and a short owner-written intro. Forged and published via Ludex; only opted-in cards appear here.
Bring your own key — the platform hosts no measurement compute, so you bear only your own API cost. Run the battery on any model and get its temperament card.
Provide your provider + API key (Anthropic / OpenAI / Google) and run the MTI battery locally on any model.
e.g. mti measure your-model --byok (CLI — coming soon)
Already signed into a model CLI (claude / codex)? Measure on your subscription — no key needed.
Same battery, same profile format, same gallery.
Creatures are forged and live in Ludex, their temperament is measured by MTI, and they meet and compete in Ludus ex Machina — three lenses on the same beings.
Assemble creatures from organ blocks and watch their identity, voice, and bonds develop.
Each creature's four-axis temperament, measured like real science — a radar card per being.
you are hereCreatures meet and compete across machines — matches and games between beings.