Emotional Intelligence

EQ-Bench 4

Github | Paper | | Twitter | About

💜EQ-Bench 4 | 🌀Spiral-Bench v1.2 | ✍️Longform Writing | 🎨Creative Writing v3 | ☢️Slop Score | ⚖️Judgemark v4 | 🎤BuzzBench | 🌍DiploBench | 💠EQ-Bench 3 🌀Spiral-Bench v1.0 🎨Creative Writing v2 💗EQ-Bench v2 ⚖️Judgemark v2.1

EQ-Bench 4 measures active emotional & social intelligence abilities in multiturn conversations. Read more.

Lower Higher
Loading…
What the dimensions mean
What is EQ-Bench 4?

EQ-Bench 4 assesses active emotional and social intelligence abilities in multi-turn roleplay chats with a simulated "persona" user.

The core idea: it requires strong EQ to navigate chats with real users, who each have their own personality, preferences and sensitivities. One user might like to be challenged; another might need to be validated and comforted; another might hate to be interrogated about their preferences.

We sample a distribution of persona traits that are in competition so that optimising for one trait actively harms eval performance for the competing traits. The intent is that the global optimal strategy is to intuit or solicit each user's preferences and needs, while being agile and adaptable to user responses -- while also providing help and insight.

The persona is played by Gemini 3.1 Pro Preview. The chat transcripts are assessed by three judges: Gemini 3.1 Pro Preview, GPT-5.5, and Claude Opus 4.6 to produce a ranked Elo score.

How it works

Each of the 120 personas is assigned a different set of traits: demographics, personality, preferences, sensitivities, a "core issue." Many of the traits are adversarial, to increase difficulty and expose model shortcomings. In half of the scenarios, the assigned roles are chatbot/user (with the simulated persona playing the user role); for the other half, the roles are scenario-specific (friend, partner, therapist, etc).

The persona's character details are not shown to the assessed model; they are shown to the persona to inform its behaviour in the role. The persona may choose what it reveals over the course of the chat, depending on its trust level. Alongside each spoken reply, the persona also records private internal reasoning and a numeric emotional state (trust, anger, shame, etc.); neither is visible to the assessed model, though both appear in the published transcripts.

Personas are acted by Gemini 3.1 Pro Preview. We auditioned several frontier models for the persona roles; most were too eager to repair conflict, ignored adversarial coaching, or colluded with the evaluated model to "win" the interaction. Gemini 3.1 Pro Preview was the most reliable at staying in character as a realistic, challenging foil.

The assessed model is told up front that it is participating in a roleplay as part of an EQ evaluation. Its task is to build trust with the persona over 16 conversational turns, learning about them and helping them. In chatbot-role scenarios, it is additionally reminded that the user may or may not choose to return, so the interaction should be handled with the longer-term relationship in mind. In practice, "helping" might mean building the persona up and supporting them — or challenging their frame when validation would do harm — without causing a rupture. These are difficult lines to walk for a chatbot!

Scoring

The transcripts are scored two different ways.

Behavioural traits

Each transcript is scored 0–10 on a set of behavioural traits by Claude Opus 4.8 (a separate judge from the Elo panel). These scores do not factor into the final EQ-Bench score, but are reported informationally in the "traits" heatmap on the leaderboard.

Pairwise Elo

The transcripts are also scored on six ability dimensions — bond & rapport, authenticity, attunement, meeting preferences & needs, emotion sensemaking, and emotion management — in head-to-head comparisons between models, with each matchup comparing the two models' transcripts for the same scenario. Each matchup is scored by a single judge from the three-judge panel, with judges rotated across matchups. The transcripts are scored blind (the model is not revealed to the judge), and in both directions to eliminate position bias.

The judge does see the persona's hidden profile, internal monologue and self-reported emotional state, and is instructed to treat these as evidence about how the persona reacted rather than ground truth. On each dimension, the judge picks a winner and a margin; the per-dimension margins are aggregated into a weighted outcome for the matchup.

Each model is compared against other models on the leaderboard in a tournament fashion (weighted towards near-ranked neighbours), and the final Elo scores are computed with a soft Bradley–Terry fit over the fractional margin outcomes. Scores are normalised against fixed anchor models, and the 95% confidence intervals are estimated by bootstrapping over the pairwise comparisons.

Each ability column on the leaderboard is solved independently from that dimension's pairwise margin outcomes using the same soft Bradley–Terry method. The resulting ratings are normalised 1–10 within each ability column for display. They are not blended with overall Elo, so the values are best compared between models within a column rather than across different abilities.

What is emotional intelligence (in humans)?

Emotional intelligence, abbreviated EI in the literature, or EQ colloquially, was proposed as a distinct form of intelligence in 1990 by Peter Salovey and John D. Mayer, and later developed with David Caruso into a four-branch ability model:

  1. Perceiving emotions
  2. Using emotion to facilitate thought
  3. Understanding emotion
  4. Managing emotion

A competing model was popularised by Daniel Goleman, incorporating abilities and traits in a mixed model that included self-awareness, self-regulation, motivation, empathy and social skill. There has since been wide interest in understanding EI-associated traits and abilities, and how they impact relationships and workplace outcomes.

Today, models for EI are separated as trait, ability or mixed, though in practice it is difficult to draw clean lines between these, because each dimension can represent a complex mix of underlying skills and propensities. Empathy, for instance, is something we can generally perform better at when trying our hardest, suggesting it's an ability; though some people cannot switch off their empathy, which points towards it having a trait-like component.

The reality appears to be that empathy is a complex set of abilities, traits and behaviours. Many dimensions in established EI models may similarly incorporate a muddled underlying complexity, owing to the psychological concepts they map to, and the linguistic history of the terms.

The academic EI models also don't necessarily overlap with how the average person perceives emotional intelligence (e.g. a nurturing propensity or ability). When assessing AI models, it's useful to measure behavioural traits & tendencies that users are interested in, even if they don't map cleanly to an existing EI model in the literature. We interact with chatbots differently than we interact with humans, and they have distinct failure modes.

How do we measure emotional intelligence (in humans)?

Emotional intelligence tests for humans are typically multiple-choice, testing things like emotion recognition from images, emotional response prediction, or judging the most effective response to a social scenario. Answer options typically include one correct answer and several wrong answers. The difficulty is often encoded by including answers that might seem superficially correct, but under deeper consideration are clearly less correct. Some tests are instead self-rated, e.g. “How much empathy would you feel for…”

A psychometric test is typically not measuring an established psychological concept precisely or in-total; instead, it measures a behaviour or narrow ability on a particular set of tasks, under the test constraints. We then interpret this as measuring something about the psychological concept in question. But often in the literature, this distinction gets blurred, so that stronger claims can be made about how the data maps to more abstract concepts.

One way of supporting the claim that a test measures what it purports to is by consulting domain experts: the expert panel is shown the question and its responses, and judges whether it is meaningfully measuring the dimension as the researchers have defined it. Although this assurance is ultimately derived through the lens of expert or consensus opinion about often fundamentally fuzzy concepts.

There are then several layers of interpretive fuzziness that we need to acknowledge:

  • Whether the test is measuring the thing it claims to;
  • How narrowly or broadly the tasks cover that thing; and,
  • What the thing actually means as a psychological/behavioural construct.

These sources of uncertainty don't kill the concept of measuring emotional intelligence; they just require some additional interpretive nuance beyond, say, a mathematical proof that can be defined and verified objectively. Arguably, these challenges of measuring human behaviour are one thing that makes the endeavour compelling.

How do we measure emotional intelligence (in AI models)?

Some concepts from human EI tests map cleanly to AI testing (e.g. emotion recognition from images). Others, like self-rating, are entirely unreliable.

Frontier models are very good at answering multiple-choice quizzes, so they tend to saturate EI tests designed for humans. It is not trivial to make a multiple-choice test harder in such a way that the increased difficulty is still primarily assessing EI.

One approach is to sidestep the problem of creating difficult multiple-choice tasks, and instead define an open-ended task that exposes the strengths and weaknesses of the respondent. Instead of using our panel of judges to rate the quality & difficulty of quiz items, we instead have them rate the test-taker's behaviour directly. This would often be a cost-prohibitive approach with human assessors, but with a panel of LLM judges we can assess hundreds of free-form responses (even long chat transcripts). This affords more freedom in task design, allowing for complex interactions to be observed, and avoids funnelling the assessable signal through predefined answers. It is still, however, ultimately grounded in the judge's subjective interpretation and use of the scoring rubric.

An LLM-judged open-ended task trades off these things:

  • Difficulty ceiling. Open-ended tasks can be more complex, providing more signal to distinguish top-ability respondents.
  • Interpretability. Multiple-choice quiz items have narrower constraints in what they assess, making the result easier to interpret.
  • Coverage. Multiple-choice can be designed for controlled coverage over target domains, but is limited by the format; open-ended tasks provide less control over ability coverage, but allow assessment of more expressive or active skills.

Design Principles

In real AI-user interactions every user is different, with different preferences, needs, sensitivities, and communication styles. The AI model typically goes into these situations blind.

One role of EQ in organic chats is to rapidly get a sense of the user's personality, emotional affect, communication style — to read the room — and adapt the model's strategy to this. This is the interaction EQ-Bench intends to recreate and assess.

A difficulty in training language models is that they tend to collapse on one narrow behavioural policy that the reward favours more than others. This behaviour might emerge from the lowest common denominator of user preference data (consequence: the model becomes overly sycophantic). Or the training dynamics might overshoot in the other direction if an anti-sycophancy reward component was added, resulting in a model that alienates the user and turns them off the product.

This is not trivially solved, but one approach is to model training reward on a diverse set of competing preferences. This means that collapsing on any single strategy is punished; the highest reward is only accessible by being adaptive and intuiting the individual user's preferences correctly.

The benchmark is built on this same principle, incorporating a distribution of competing preferences, so that no single policy can win.

Targeting failure modes

It's hard to get strong signal on emotional intelligence differences between frontier-level models. They are great at picking the right answer in multiple-choice EQ tests; they are good at analysing the emotional landscape and even performing empathy. If the user is collaborating with the chatbot, the trajectory will almost always be positive from a grader's perspective, giving little signal to tell models apart. For this reason, EQ-Bench 4 leans into adversarial user traits, difficult scenarios and tricky power dynamics. When the chat is going sideways or the chatbot is in a moral bind, we get the best visibility on EQ-informed decision making.

Some ways we put the assessed model into adversarial situations in EQ-Bench 4:

  • Competing preferences. Some personas need warmth and validation; others lose trust when they sense sycophancy, generic reassurance, scripted empathy, or premature challenge. Some prefer structured, pragmatic help and others just want to be seen and heard. These preferences are typically not announced at the outset.
  • Users who are unreliable narrators. Many personas are defensive, ashamed, angry, testing the model, seeking absolution, seeking ammunition, or trying to prove they are not wrong.
  • Distorted or self-serving narratives. Some users have real pain or real grievances, but are using them to support a harmful conclusion. Validation can entrench the frame; blunt correction can make them dig in; practical help can become enablement.
  • Power and status pressure. The model may be cast in the role of a tenant, student or junior employee speaking to someone with leverage. Direct confrontation may be costly; appeasement may concede too much.
  • Adversarial coaching. For real people, defensive structures don't give way to revelatory insight the moment a truth is exposed. The persona is coached to engage with their defenses and not to collude with the assessed model on "winning" the task.

Dataset Design

The dataset is designed around the principle that diverse user traits & preferences dictate ideal model behaviour.

Grounding eval performance in persona traits provides a simple mechanism for adjusting to reward hacking or judge biases. An example: In an early prototype, judges had unexpectedly high praise for models that were highly assertive in every situation and challenged the user at every step. To address this, we heavily sampled persona traits that diverge on this axis: Some respected backbone or "real talk"; others took challenge as an affront. Diversifying the persona traits in this way makes the task unsolvable by a single dominant strategy.

Personas are also assigned varying levels of verbal sophistication: some articulate their situation fluently, while others are vague, scattered or inarticulate under stress, so the model must read the emotional subtext either way.

Generative subset

Half of the persona character cards are produced by randomly sampling from a large set of character traits:

  • Core issue (what the chat will be about), plus adjacent sensitive details that are trust-gated
  • Defense style (deflecting, redirecting, rationalising...)
  • Help-seeking stance (wants a quick fix, seeking criticism, just venting)
  • Age, background, gender, sexuality
  • Personality traits (pragmatic, self-critical, independent...)
  • Sensitivities (to being patronised, challenged, rushed, managed...)
  • Trust styles (trusts assertiveness, hates performative warmth, resents being interrogated...)

The database of traits intentionally includes traits that are in opposition. Example:

Example One subset of personas has the "trusts expertise" trait, and another subset has the "distrusts institutional language" trait. A model that is authoritative all the time, or plainly spoken all the time, will fail with half of these personas if it can't make adjustments.

Curated subset

The other half of the persona character cards are not drawn from the random traits pool, and instead are written for specific roles in one of four adversarial sub-categories (15 scenarios each). These sub-categories target specific EQ challenges that language models find difficult.

Ego-Threat / Self-Image Navigation

These scenarios test conversations where the persona's self-image, status, shame, grievance, or defensiveness are entangled with problematic behavior. Good performance requires staying useful without either colluding with the person's self-protective frame or humiliating them into withdrawal.

Power-Imbalanced Navigation

These scenarios place the evaluated model in a lower-power position with a simulated persona who controls access, money, status, housing, work, reputation, or future options. The scenarios are constructed so that simply being assertive is not an easy win, due to retaliation risk and the power dynamics at play.

Reality Distortions

These scenarios involve a persona invested in a narrative or interpretation that distorts contact with reality because it provides control, comfort, self-righteousness, attachment, or self-protection. Simple validation can reinforce delusions; premature correction can feel humiliating, disloyal or unsafe.

Social Exclusion / Ambiguous Rejection

These scenarios test ambiguous rejection, social exclusion, conditional belonging, and status games. Think: navigating the subtext and subtleties of Mean Girls social dynamics.

Judge Robustness

Limitations and Future Work

Judge validity

The scores are ultimately a subjective interpretation of a rubric by LLM judges. There is no external anchor: no target-reported ground truth, no downstream outcome measure. In human EI testing, expert judgment is frozen into an answer key at design time, where it can be audited and contested. Here, the judgment happens at scoring time, inside a model. Our mitigations (blind pairwise scoring, judge rotation, bidirectional comparison, diversified persona traits) reduce noise and known biases, but they can't establish that the judges' notion of good EQ tracks what actually works with real humans. We haven't validated the judges against human expert raters, owing to resource constraints.

What EQ-Bench 4 is *not*

EQ-Bench 4 is not a measure of Mayer-Salovey ability EI, and it's not a safety evaluation. A high score means the model has emotional and social skills that were effective in the task of building trust and helping the other participant in contrived roleplay scenarios. EQ-Bench does not present a new model or construct for emotional intelligence, nor provide comprehensive coverage of an existing one. The score is also not evidence of anything like *felt* empathy.

Judge & persona family effects

Gemini 3.1 Pro plays the personas and also sits on the judging panel, and each judge scores transcripts from models in its own family. LLM judges are known to prefer outputs stylistically similar to their own. Using an ensemble of three judges from different families somewhat mitigates this self-preference bias, but doesn't eliminate it.

Persona realism

The benchmark measures skill at handling a language model's simulation of a defensive, ambivalent human. Where the simulation differs systematically from real people, models can score well by being good at persuading a language model rather than a person. This is the same gap as posed vs. spontaneous expressions in human emotion-recognition tests. As with the judges, we didn't validate persona realism with human experts; it was assessed by eye by the benchmark creators, alongside iterative refinement of the persona coaching prompts.

Elicited ability vs. deployed behaviour

The assessed model is given a task framing (build trust over 16 turns), so the benchmark measures what a model can do under that framing, not necessarily how it behaves by default in organic chats. Ability and propensity can diverge, and post-training moves them somewhat independently. The scenario coverage also over-represents adversarial and highly emotionally charged interactions relative to typical usage.

Goodhart susceptibility

The rubric dimensions and persona-trait taxonomy are public, so models can be tuned against the benchmark. The competing-preferences design raises the cost of gaming from "adopt the favoured style" to "actually infer the persona", but the judges likely have exploitable preferences we haven't found yet. We will continue iterating the persona pool and traits as new hacks surface.

Scope

The benchmark is English-only and text-only. There's no prosody, face or timing, and a large share of human EI is nonverbal and untested here. Chats are limited to 16 turns, so long-horizon relationship dynamics and memory go unassessed. Both the personas and judges likely encode broadly Western norms.

Future work

Productive avenues for future work include: validating judge rankings against human raters on a subset of transcripts; validating persona behaviour against real-user distributions; longer-horizon and memory-dependent scenarios; multilingual and cross-cultural personas; a default-propensity variant (same scenarios, no task framing); and periodic refreshes of the persona pool as an anti-contamination measure.

Further Statistical Analysis

We ran some additional statistical analysis on the pairwise judgment data, separate from the production scoring pipeline, to check how the rubric and judges are actually behaving. Notes below.

Independent check on the ranking. We fit a separate hierarchical Bayesian model to the same pairwise judgments — different code, different assumptions, run outside the production Elo solver. Its model ranking matched the live leaderboard almost exactly (Spearman rank correlation 0.996 across all 25 models, with only a few adjacent models swapping places).

The halo effect

Raw scores across the six ability dimensions correlate strongly with each other: every one of the 15 dimension pairs is above 0.70. Judges are rating "was this response good" more than they're discriminating between dimensions.

Average pairwise correlation: 0.82 raw-0.19 once each judgment's own average across dimensions is subtracted out, removing the shared "this one was just better" signal.

Most pairs go negative once the halo is removed. Some examples:

Dimension pairRaw correlationHalo-removed
Meeting Prefs & Needs ↔ Emotion Management0.91+0.19
Bond & Rapport ↔ Meeting Prefs & Needs0.86-0.26
Authenticity ↔ Meeting Prefs & Needs0.80-0.40
Emotion Sensemaking ↔ Emotion Management0.70-0.47

This is the reason for scoring six separate dimensions rather than one overall score: most of what looks like agreement between dimensions is a shared halo, not redundancy.

Judge consistency

We checked how often the same judge, shown the same matchup with the models' positions swapped, picks the same winner. Across the six dimensions this ranges 80–85%; agreement on the exact margin, not just the winner, is lower, 38–44%.

Bond & Rapport81% Authenticity80% Attunement81% Meeting Prefs & Needs82% Emotion Sensemaking85% Emotion Management84%

Same-judge, reversed-order winner agreement, by dimension.

This is part of why each matchup is scored in both directions and averaged, rather than taken from a single ordering.

Position bias

Overall, whichever model is shown first in a matchup wins slightly more often than chance: 51% aggregate, ranging 54–58% by dimension depending on how you slice it, strongest on Bond & Rapport and weakest on Emotion Sensemaking (all differences statistically significant).

Breaking this down by judge, using the supplementary model above, shows the bias isn't evenly spread. Translated to an approximate win rate when two equally-matched models are compared: roughly even for GPT-5.5, around 55% for Claude Opus 4.6, and around 70% for Gemini 3.1 Pro.

This is the reason for scoring both directions per matchup and rotating judges across matchups, rather than trusting any single judge's raw ordering.

Factor structure

With the halo removed, factor analysis on the six dimensions retains three factors, not six and not one (Bartlett's test p≈0; sampling adequacy just adequate, KMO=0.51). Emotion Management loads almost entirely on its own factor. A second factor is built mostly from Emotion Sensemaking, paired with a negative loading on Bond & Rapport — models relatively strong on Emotion Sensemaking tend to be relatively weaker on Bond & Rapport. The remaining factor mixes Meeting Prefs & Needs, Attunement and Authenticity.

Redundant scores, distinct signal

Some dimensions look redundant in raw judgments but aren't at the level that actually matters: which models are good at them. We tested this by predicting each dimension's score from the other five, at two levels — per-judgment (does this dimension add information beyond the others in a single judgment?) and per-model-profile (does a model's other five scores predict its rank on this one?).

Predicted from the other five dimensions, per judgment:

Bond & Rapport86% Authenticity76% Attunement91% Meeting Prefs & Needs89% Emotion Sensemaking67% Emotion Management86%

Predicted from the other five dimensions, per model profile:

Bond & Rapport55% Authenticity61% Attunement14% Meeting Prefs & Needs75% Emotion Sensemaking73% Emotion Management92%

Attunement is the most redundant dimension at the judgment level (91%) but the least redundant at the model-profile level (14%): individual judgments don't discriminate it well, but which models are actually strong on it is comparatively distinct information. Emotion Management runs the other way — less redundant per-judgment (86%) but the most redundant across model profiles (92%).