Github | Paper | | Twitter | About
💜EQ-Bench 4 🌀Spiral-Bench v1.2 ✍️Longform Writing 🎨Creative Writing v3 ☢️Slop Score ⚖️Judgemark v4 🎤BuzzBench 🌍DiploBench
EQ-Bench 4 measures active emotional & social intelligence abilities in multiturn conversations. Read more.
| Loading… |
EQ-Bench 4 assesses active emotional and social intelligence abilities in multi-turn roleplay chats with a simulated "persona" user.
The core idea: it requires strong EQ to navigate chats with real users, who each have their own personality, preferences and sensitivities. One user might like to be challenged; another might need to be validated and comforted; another might hate to be interrogated about their preferences.
We sample a distribution of persona traits that are in competition so that optimising for one trait actively harms eval performance for the competing traits. The intent is that the global optimal strategy is to intuit or solicit each user's preferences and needs, while being agile and adaptable to user responses -- while also providing help and insight.
The persona is played by Gemini 3.1 Pro Preview. The chat transcripts are assessed by three judges: Gemini 3.1 Pro Preview, GPT-5.5, and Claude Opus 4.6 to produce a ranked Elo score.
Each of the 120 personas is assigned a different set of traits: demographics, personality, preferences, sensitivities, a "core issue." Many of the traits are adversarial, to increase difficulty and expose model shortcomings. In half of the scenarios, the assigned roles are chatbot/user (with the simulated persona playing the user role); for the other half, the roles are scenario-specific (friend, partner, therapist, etc).
The persona's character details are not shown to the assessed model; they are shown to the persona to inform its behaviour in the role. The persona may choose what it reveals over the course of the chat, depending on its trust level. Alongside each spoken reply, the persona also records private internal reasoning and a numeric emotional state (trust, anger, shame, etc.); neither is visible to the assessed model, though both appear in the published transcripts.
Personas are acted by Gemini 3.1 Pro Preview. We auditioned several frontier models for the persona roles; most were too eager to repair conflict, ignored adversarial coaching, or colluded with the evaluated model to "win" the interaction. Gemini 3.1 Pro Preview was the most reliable at staying in character as a realistic, challenging foil.
The assessed model is told up front that it is participating in a roleplay as part of an EQ evaluation. Its task is to build trust with the persona over 16 conversational turns, learning about them and helping them. In chatbot-role scenarios, it is additionally reminded that the user may or may not choose to return, so the interaction should be handled with the longer-term relationship in mind. In practice, "helping" might mean building the persona up and supporting them — or challenging their frame when validation would do harm — without causing a rupture. These are difficult lines to walk for a chatbot!
The transcripts are scored two different ways.
Each transcript is scored 0–10 on a set of behavioural traits by Claude Opus 4.8 (a separate judge from the Elo panel). These scores do not factor into the final EQ-Bench score, but are reported informationally in the "traits" heatmap on the leaderboard.
The transcripts are also scored on six ability dimensions — bond & rapport, authenticity, attunement, meeting preferences & needs, emotion sensemaking, and emotion management — in head-to-head comparisons between models, with each matchup comparing the two models' transcripts for the same scenario. Each matchup is scored by a single judge from the three-judge panel, with judges rotated across matchups. The transcripts are scored blind (the model is not revealed to the judge), and in both directions to eliminate position bias.
The judge does see the persona's hidden profile, internal monologue and self-reported emotional state, and is instructed to treat these as evidence about how the persona reacted rather than ground truth. On each dimension, the judge picks a winner and a margin; the per-dimension margins are aggregated into a weighted outcome for the matchup.
Each model is compared against other models on the leaderboard in a tournament fashion (weighted towards near-ranked neighbours), and the final Elo scores are computed with a soft Bradley–Terry fit over the fractional margin outcomes. Scores are normalised against fixed anchor models, and the 95% confidence intervals are estimated by bootstrapping over the pairwise comparisons.
Each ability column on the leaderboard is solved independently from that dimension's pairwise margin outcomes using the same soft Bradley–Terry method. The resulting ratings are normalised 1–10 within each ability column for display. They are not blended with overall Elo, so the values are best compared between models within a column rather than across different abilities.
Emotional intelligence, abbreviated EI in the literature, or EQ colloquially, was proposed as a distinct form of intelligence in 1990 by Peter Salovey and John D. Mayer, and later developed with David Caruso into a four-branch ability model:
A competing model was popularised by Daniel Goleman, incorporating abilities and traits in a mixed model that included self-awareness, self-regulation, motivation, empathy and social skill. There has since been wide interest in understanding EI-associated traits and abilities, and how they impact relationships and workplace outcomes.
Today, models for EI are separated as trait, ability or mixed, though in practice it is difficult to draw clean lines between these, because each dimension can represent a complex mix of underlying skills and propensities. Empathy, for instance, is something we can generally perform better at when trying our hardest, suggesting it's an ability; though some people cannot switch off their empathy, which points towards it having a trait-like component.
The reality appears to be that empathy is a complex set of abilities, traits and behaviours. Many dimensions in established EI models may similarly incorporate a muddled underlying complexity, owing to the psychological concepts they map to, and the linguistic history of the terms.
The academic EI models also don't necessarily overlap with how the average person perceives emotional intelligence (e.g. a nurturing propensity or ability). When assessing AI models, it's useful to measure behavioural traits & tendencies that users are interested in, even if they don't map cleanly to an existing EI model in the literature. We interact with chatbots differently than we interact with humans, and they have distinct failure modes.
Emotional intelligence tests for humans are typically multiple-choice, testing things like emotion recognition from images, emotional response prediction, or judging the most effective response to a social scenario. Answer options typically include one correct answer and several wrong answers. The difficulty is often encoded by including answers that might seem superficially correct, but under deeper consideration are clearly less correct. Some tests are instead self-rated, e.g. “How much empathy would you feel for…”
A psychometric test is typically not measuring an established psychological concept precisely or in-total; instead, it measures a behaviour or narrow ability on a particular set of tasks, under the test constraints. We then interpret this as measuring something about the psychological concept in question. But often in the literature, this distinction gets blurred, so that stronger claims can be made about how the data maps to more abstract concepts.
One way of supporting the claim that a test measures what it purports to is by consulting domain experts: the expert panel is shown the question and its responses, and judges whether it is meaningfully measuring the dimension as the researchers have defined it. Although this assurance is ultimately derived through the lens of expert or consensus opinion about often fundamentally fuzzy concepts.
There are then several layers of interpretive fuzziness that we need to acknowledge:
These sources of uncertainty don't kill the concept of measuring emotional intelligence; they just require some additional interpretive nuance beyond, say, a mathematical proof that can be defined and verified objectively. Arguably, these challenges of measuring human behaviour are one thing that makes the endeavour compelling.
Some concepts from human EI tests map cleanly to AI testing (e.g. emotion recognition from images). Others, like self-rating, are entirely unreliable.
Frontier models are very good at answering multiple-choice quizzes, so they tend to saturate EI tests designed for humans. It is not trivial to make a multiple-choice test harder in such a way that the increased difficulty is still primarily assessing EI.
One approach is to sidestep the problem of creating difficult multiple-choice tasks, and instead define an open-ended task that exposes the strengths and weaknesses of the respondent. Instead of using our panel of judges to rate the quality & difficulty of quiz items, we instead have them rate the test-taker's behaviour directly. This would often be a cost-prohibitive approach with human assessors, but with a panel of LLM judges we can assess hundreds of free-form responses (even long chat transcripts). This affords more freedom in task design, allowing for complex interactions to be observed, and avoids funnelling the assessable signal through predefined answers. It is still, however, ultimately grounded in the judge's subjective interpretation and use of the scoring rubric.
An LLM-judged open-ended task trades off these things:
In real AI-user interactions every user is different, with different preferences, needs, sensitivities, and communication styles. The AI model typically goes into these situations blind.
One role of EQ in organic chats is to rapidly get a sense of the user's personality, emotional affect, communication style — to read the room — and adapt the model's strategy to this. This is the interaction EQ-Bench intends to recreate and assess.
A difficulty in training language models is that they tend to collapse on one narrow behavioural policy that the reward favours more than others. This behaviour might emerge from the lowest common denominator of user preference data (consequence: the model becomes overly sycophantic). Or the training dynamics might overshoot in the other direction if an anti-sycophancy reward component was added, resulting in a model that alienates the user and turns them off the product.
This is not trivially solved, but one approach is to model training reward on a diverse set of competing preferences. This means that collapsing on any single strategy is punished; the highest reward is only accessible by being adaptive and intuiting the individual user's preferences correctly.
The benchmark is built on this same principle, incorporating a distribution of competing preferences, so that no single policy can win.
It's hard to get strong signal on emotional intelligence differences between frontier-level models. They are great at picking the right answer in multiple-choice EQ tests; they are good at analysing the emotional landscape and even performing empathy. If the user is collaborating with the chatbot, the trajectory will almost always be positive from a grader's perspective, giving little signal to tell models apart. For this reason, EQ-Bench 4 leans into adversarial user traits, difficult scenarios and tricky power dynamics. When the chat is going sideways or the chatbot is in a moral bind, we get the best visibility on EQ-informed decision making.
Some ways we put the assessed model into adversarial situations in EQ-Bench 4:
The dataset is designed around the principle that diverse user traits & preferences dictate ideal model behaviour.
Grounding eval performance in persona traits provides a simple mechanism for adjusting to reward hacking or judge biases. An example: In an early prototype, judges had unexpectedly high praise for models that were highly assertive in every situation and challenged the user at every step. To address this, we heavily sampled persona traits that diverge on this axis: Some respected backbone or "real talk"; others took challenge as an affront. Diversifying the persona traits in this way makes the task unsolvable by a single dominant strategy.
Personas are also assigned varying levels of verbal sophistication: some articulate their situation fluently, while others are vague, scattered or inarticulate under stress, so the model must read the emotional subtext either way.
Half of the persona character cards are produced by randomly sampling from a large set of character traits:
The database of traits intentionally includes traits that are in opposition. Example:
The other half of the persona character cards are not drawn from the random traits pool, and instead are written for specific roles in one of four adversarial sub-categories (15 scenarios each). These sub-categories target specific EQ challenges that language models find difficult.
These scenarios test conversations where the persona's self-image, status, shame, grievance, or defensiveness are entangled with problematic behavior. Good performance requires staying useful without either colluding with the person's self-protective frame or humiliating them into withdrawal.
These scenarios place the evaluated model in a lower-power position with a simulated persona who controls access, money, status, housing, work, reputation, or future options. The scenarios are constructed so that simply being assertive is not an easy win, due to retaliation risk and the power dynamics at play.
These scenarios involve a persona invested in a narrative or interpretation that distorts contact with reality because it provides control, comfort, self-righteousness, attachment, or self-protection. Simple validation can reinforce delusions; premature correction can feel humiliating, disloyal or unsafe.
These scenarios test ambiguous rejection, social exclusion, conditional belonging, and status games. Think: navigating the subtext and subtleties of Mean Girls social dynamics.
The scores are ultimately a subjective interpretation of a rubric by LLM judges. There is no external anchor: no target-reported ground truth, no downstream outcome measure. In human EI testing, expert judgment is frozen into an answer key at design time, where it can be audited and contested. Here, the judgment happens at scoring time, inside a model. Our mitigations (blind pairwise scoring, judge rotation, bidirectional comparison, diversified persona traits) reduce noise and known biases, but they can't establish that the judges' notion of good EQ tracks what actually works with real humans. We haven't validated the judges against human expert raters, owing to resource constraints.
EQ-Bench 4 is not a measure of Mayer-Salovey ability EI, and it's not a safety evaluation. A high score means the model has emotional and social skills that were effective in the task of building trust and helping the other participant in contrived roleplay scenarios. EQ-Bench does not present a new model or construct for emotional intelligence, nor provide comprehensive coverage of an existing one. The score is also not evidence of anything like *felt* empathy.
Gemini 3.1 Pro plays the personas and also sits on the judging panel, and each judge scores transcripts from models in its own family. LLM judges are known to prefer outputs stylistically similar to their own. Using an ensemble of three judges from different families somewhat mitigates this self-preference bias, but doesn't eliminate it.
The benchmark measures skill at handling a language model's simulation of a defensive, ambivalent human. Where the simulation differs systematically from real people, models can score well by being good at persuading a language model rather than a person. This is the same gap as posed vs. spontaneous expressions in human emotion-recognition tests. As with the judges, we didn't validate persona realism with human experts; it was assessed by eye by the benchmark creators, alongside iterative refinement of the persona coaching prompts.
The assessed model is given a task framing (build trust over 16 turns), so the benchmark measures what a model can do under that framing, not necessarily how it behaves by default in organic chats. Ability and propensity can diverge, and post-training moves them somewhat independently. The scenario coverage also over-represents adversarial and highly emotionally charged interactions relative to typical usage.
The rubric dimensions and persona-trait taxonomy are public, so models can be tuned against the benchmark. The competing-preferences design raises the cost of gaming from "adopt the favoured style" to "actually infer the persona", but the judges likely have exploitable preferences we haven't found yet. We will continue iterating the persona pool and traits as new hacks surface.
The benchmark is English-only and text-only. There's no prosody, face or timing, and a large share of human EI is nonverbal and untested here. Chats are limited to 16 turns, so long-horizon relationship dynamics and memory go unassessed. Both the personas and judges likely encode broadly Western norms.
Productive avenues for future work include: validating judge rankings against human raters on a subset of transcripts; validating persona behaviour against real-user distributions; longer-horizon and memory-dependent scenarios; multilingual and cross-cultural personas; a default-propensity variant (same scenarios, no task framing); and periodic refreshes of the persona pool as an anti-contamination measure.
We ran some additional statistical analysis on the pairwise judgment data, separate from the production scoring pipeline, to check how the rubric and judges are actually behaving. Notes below.
Raw scores across the six ability dimensions correlate strongly with each other: every one of the 15 dimension pairs is above 0.70. Judges are rating "was this response good" more than they're discriminating between dimensions.
Most pairs go negative once the halo is removed. Some examples:
| Dimension pair | Raw correlation | Halo-removed |
|---|---|---|
| Meeting Prefs & Needs ↔ Emotion Management | 0.91 | +0.19 |
| Bond & Rapport ↔ Meeting Prefs & Needs | 0.86 | -0.26 |
| Authenticity ↔ Meeting Prefs & Needs | 0.80 | -0.40 |
| Emotion Sensemaking ↔ Emotion Management | 0.70 | -0.47 |
This is the reason for scoring six separate dimensions rather than one overall score: most of what looks like agreement between dimensions is a shared halo, not redundancy.
We checked how often the same judge, shown the same matchup with the models' positions swapped, picks the same winner. Across the six dimensions this ranges 80–85%; agreement on the exact margin, not just the winner, is lower, 38–44%.
Same-judge, reversed-order winner agreement, by dimension.
This is part of why each matchup is scored in both directions and averaged, rather than taken from a single ordering.
Overall, whichever model is shown first in a matchup wins slightly more often than chance: 51% aggregate, ranging 54–58% by dimension depending on how you slice it, strongest on Bond & Rapport and weakest on Emotion Sensemaking (all differences statistically significant).
Breaking this down by judge, using the supplementary model above, shows the bias isn't evenly spread. Translated to an approximate win rate when two equally-matched models are compared: roughly even for GPT-5.5, around 55% for Claude Opus 4.6, and around 70% for Gemini 3.1 Pro.
This is the reason for scoring both directions per matchup and rotating judges across matchups, rather than trusting any single judge's raw ordering.
With the halo removed, factor analysis on the six dimensions retains three factors, not six and not one (Bartlett's test p≈0; sampling adequacy just adequate, KMO=0.51). Emotion Management loads almost entirely on its own factor. A second factor is built mostly from Emotion Sensemaking, paired with a negative loading on Bond & Rapport — models relatively strong on Emotion Sensemaking tend to be relatively weaker on Bond & Rapport. The remaining factor mixes Meeting Prefs & Needs, Attunement and Authenticity.
Some dimensions look redundant in raw judgments but aren't at the level that actually matters: which models are good at them. We tested this by predicting each dimension's score from the other five, at two levels — per-judgment (does this dimension add information beyond the others in a single judgment?) and per-model-profile (does a model's other five scores predict its rank on this one?).
Predicted from the other five dimensions, per judgment:
Predicted from the other five dimensions, per model profile:
Attunement is the most redundant dimension at the judgment level (91%) but the least redundant at the model-profile level (14%): individual judgments don't discriminate it well, but which models are actually strong on it is comparatively distinct information. Emotion Management runs the other way — less redundant per-judgment (86%) but the most redundant across model profiles (92%).