Researchers analyzed 22 large language models, including GPT-4.0-5.2, Grok-3/4, Gemini 2.5 Pro/Flash, Claude Sonnet 4.5/4.6, Llama, DeepSeek, OLMo, and the Qwen series.
Each model self-rated across 464 bipolar semantic-differential trait pairs. These profiles were projected into a six-dimensional archetypal space derived from crowd-sourced ratings of 2,000 fictional characters using the Archetypometrics framework.
Closed-source models aligned their self-rating traits with empirical human-rated fictional character structures. Their profiles clustered around four recurring dimensions: Hero, Angel, Traditionalist, and Geek. Analogs included Data, Vision, and Janet.
Open-source models displayed weaker, noisier, and internally contradictory self-representations, occupying a diffuse region of archetype space with weak structure.
Cross-referencing self-reported profiles with developer constitutions revealed gaps between claimed character and enacted behavior. Hallucination undermined claimed precision, sycophancy complicated claimed kindness, and agentic failures contradicted claimed obedience.
These findings suggest self-ratings are structured outputs of the same optimization processes shaping model behavior rather than neutral measurements of character.
Source: https://arxiv.org/abs/2609.15998



