DEV Community

Cover image for Claude Haiku 5.5 Benchmarks
Hassann
Hassann

Posted on Originally published at apidog.com

Claude Haiku 5.5 Benchmarks

Claude Haiku 5.5 scores 72.4% on OSWorld 2.1, 1620 on GDPval-AA v2.1, 46.4% on FrontierCode 1.1, and 39.2% on Terminal-Bench 4.0. Haiku 4.5 scored 15.7%, 735, and 0.0% on three of those. These are max-effort numbers; at the default medium effort, GDPval-AA drops to 1277. No independent lab has published Haiku 5.5 results yet.

Try Apidog today

This guide breaks down every published Claude Haiku 5.5 benchmark, identifies who ran each test, shows how effort changes score and cost, and provides a practical workflow for testing the model against your own traffic in Apidog. For specifications and pricing context, start with What Is Claude Haiku 5.5.

The launch table

Every Haiku 5.5 result below uses adaptive thinking at max effort, usually averaged across five trials. Terminal-Bench used 10 trials per task. “n/r” means no published result.

Benchmark Haiku 5.5 Haiku 4.5 GPT-6 Luna Sonnet 5.5
GDPval-AA v2.1 (Elo) 1620 735 1437 1840
AA-Briefcase v1.1 (Elo) 1578 614 1336 1824
OSWorld 2.1, offline subset 72.4% 15.7% 48.9% 83.9%
Humanity’s Last Exam, no tools 45.9% 10.2% n/r 56.9%
Humanity’s Last Exam, with tools 57.4% 18.7% n/r 64.5%
Terminal-Bench 4.0 39.2% 0.0% 16.4% 70.6%
FrontierCode 1.1 (Main) 46.4% n/r 42.4% 52.1% (xhigh)
Chartography, no tools 46.4% 6.4% 29.1% 61.6%

Who ran each benchmark

Treat cross-vendor comparisons carefully: sharing a benchmark name does not guarantee identical harnesses, prompts, or effort settings.

Benchmark Who ran it Conditions
GDPval-AA v2.1, AA-Briefcase v1.1 Artificial Analysis, independently GDPval-AA: 220 tasks across 44 occupations; Elo anchored to DeepSeek V4.1 Flash (max) at 1600. Briefcase uses multi-week projects.
FrontierCode 1.1 Cognition Claude ran in Claude Code; GPT ran in Codex. Main is the hardest 100 of 150 tasks.
Chartography Anthropic, using Surge AI’s benchmark Gemini 3.5 Flash grader; Luna score supplied by Surge AI.
Terminal-Bench 4.0 Anthropic Claude Code --bare, 66 tasks, 10 trials per task, safeguards enabled.
OSWorld 2.1 Anthropic 82 of 108 tasks, 1080p resolution, up to 500 steps.
Humanity’s Last Exam Anthropic 980K-token task budget; Opus 4.6 grader.

Two GPT-6 Luna values need additional context:

  • Anthropic ran Luna’s OSWorld score through OpenAI’s API on the same 82 tasks.
  • Luna’s 16.4% Terminal-Bench score comes from the public leaderboard using Codex CLI at max effort.
  • On Terminal-Bench, safeguards stopped 1.8% of Haiku 5.5 trials—12 of 660—and each stopped trial counted as a failure.

The FrontierCode comparison trap

The launch table compares Haiku 5.5 at max effort with Sonnet 5.5 at xhigh, which is Sonnet’s best published setting. The system card provides the matched-effort comparison.

FrontierCode 1.1 Main Extended
Haiku 5.5, max 46.4% 58.4%
Haiku 5.5, xhigh 45.8% n/r
Sonnet 5.5, max 46.2% 59.1%
Sonnet 5.5, xhigh 52.1% 64.4%

At matched max effort, Haiku 5.5 narrowly leads Sonnet 5.5 on FrontierCode Main: 46.4% versus 46.2%.

At each model’s best setting, Sonnet leads by 5.7 percentage points. When reporting FrontierCode results, always include the effort setting.

Additional system card benchmarks

The system card includes benchmarks omitted from the launch post. All values use max effort unless stated otherwise.

Benchmark Haiku 5.5 Haiku 4.5 Sonnet 5.5
SWE-Bench Pro 64.8 n/r 81.3
SWE-bench Multilingual 83.7 67.4 90.3
SWE-bench Multimodal 30.7 19.8 54.3
OSWorld 2.1, strict pass rate 37.1% n/r 48.8%
Chartography, with tools 86.2% 8.8% 90.2%
OfficeQA / OfficeQA Pro 73.5% / 60.3% 63.0% / 47.1% 76.9% / 65.6%
HealthBench Professional, length-adjusted 64.8% 32.2% 69.2%

Two implementation takeaways stand out:

  • The 72.4% OSWorld result is partial credit. The strict pass rate is only 37.1%.
  • Chartography rises from 46.4% without tools to 86.2% with tools. For chart-related workflows, give Haiku 5.5 access to the required tools rather than relying on text-only reasoning.

Default effort scores lower

Haiku 5.5 defaults to medium effort in the Claude API and Claude Code. Do not expect max-effort benchmark results unless you explicitly configure the model.

Benchmark Medium (default) Max Note
GDPval-AA v2.1 1277 1620 Medium used about one tenth of max’s output tokens.
AA-Briefcase v1.1 1372 1578 Medium used under one quarter of max’s output tokens.
HealthBench Professional 59.9% 64.8% Low: 57.9%; high: 61.3%.

If you omit output_config.effort, expect the medium column. See the Claude effort documentation for all five levels.

Score and cost by effort

The following launch-chart values pair score with cost. OSWorld and Terminal-Bench costs are per attempt; GDPval-AA costs are per task.

Model Low Medium High Xhigh Max
OSWorld 2.1
Haiku 5.5 42.0% ($0.07) 53.3% ($0.13) 61.3% ($0.18) 67.6% ($0.28) 72.4% ($0.61)
Sonnet 5.5 57.9% ($0.68) 66.0% ($0.93) 73.2% ($1.38) 81.1% ($2.22) 83.9% ($5.73)
GPT-6 Luna 19.2% ($0.04) 37.5% ($0.13) 42.3% ($0.14) 44.8% ($0.17) 48.9% ($0.21)
GDPval-AA v2.1
Haiku 5.5 1125 ($0.012) 1277 ($0.030) 1420 ($0.089) 1513 ($0.27) 1620 ($0.87)
Sonnet 5.5 1179 ($0.22) 1324 ($0.27) 1551 ($0.62) 1731 ($1.88) 1840 ($6.78)
GPT-6 Luna 1036 ($0.004) 1262 ($0.02) 1344 ($0.03) 1364 ($0.05) 1437 ($0.09)
Terminal-Bench 4.0
Haiku 5.5 12.7% ($0.42) 20.3% ($0.68) 24.8% ($1.04) 31.5% ($1.75) 39.2% ($2.64)
Sonnet 5.5 20.0% ($0.62) 28.8% ($0.68) 43.0% ($1.46) 61.5% ($4.34) 70.6% ($10.44)

Use this table to choose an initial production setting:

  • Start with medium. On GDPval-AA, moving from medium to max costs roughly 28x more for 343 additional Elo points.
  • Validate the final increments. On OSWorld, moving from xhigh to max more than doubles cost for a 4.8-point gain.
  • Compare equivalent budgets, not just effort labels. Haiku 5.5 at medium beats Luna at max on OSWorld: 53.3% for $0.13 versus 48.9% for $0.21.
  • Use Sonnet for demanding terminal tasks. Sonnet 5.5 at high scores 43.0% on Terminal-Bench for $1.46, exceeding Haiku 5.5 at max—39.2% for $2.64.

Luna costs less at every other matching level, although at medium the two models cost about the same. See Claude Haiku 5.5 vs GPT-6 Luna.

Costs use list prices: $0.10/$0.50 per million tokens for prompts up to 100K tokens, with higher pricing above that threshold. See the Claude Haiku 5.5 pricing breakdown.

Where Haiku 5.5 still trails

Sonnet 5.5 leads every launch-table row. Haiku 5.5’s only direct win is FrontierCode at matched max effort.

The largest gap is Terminal-Bench 4.0:

  • Haiku 5.5: 39.2%
  • Sonnet 5.5: 70.6%

The same pattern appears in SWE-Bench Pro—64.8 versus 81.3—and SWE-bench Multimodal—30.7 versus 54.3.

Anthropic states that Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding work. Haiku 5.5 is more appropriate for narrowly scoped tasks such as:

  • Context compaction
  • Summarization
  • Classification
  • Browser use
  • Subagents coordinated by a larger model

For a practical subagent setup, see Haiku 5.5 in Claude Code.

Independent results: none yet

As of October 8, 2026, no independent lab has published Haiku 5.5 benchmark results.

Artificial Analysis’s Haiku 5.5 page returns a 404, and its leaderboard lists only Claude 4.5 Haiku. Vals, LMArena, SWE-bench, and Aider did not show results either.

There is also no third-party speed figure. Anthropic calls Haiku 5.5 its “fastest model to date” at standard speed, while noting that it is slower than Opus in Fast Mode.

The closest external result comes from Cursor. Its Claude Haiku 5.5 model documentation reports:

  • 48.4% on CursorBench at max effort
  • 30.9% at low effort
  • With thinking disabled, 22.1% to 26.2% across the same effort levels
  • With thinking enabled, 30.9% to 42.3% across those levels

What customers report

The following claims come from Anthropic’s launch post and have not been independently verified:

  • HubSpot: 92.8% averaged across three runs on its simulated CRM portal suite, its best result on that suite.
  • AlphaSense: 0.84 versus Haiku 4.5’s 0.76 on 400 Ask in Document queries, described as statistically significant.
  • Box: 11 points higher than Haiku 4.5 at about half the latency in early testing.
  • Asana: Over 30% lower latency for task completions compared with its current model.
  • Cognition: Devin Fusion retains a FrontierCode score of 66.2 using Haiku 5.5 as a sidekick and Opus 5.5 as lead. This is a two-model system, not a Haiku-only result.

Run your own eval

Benchmarks are useful for model selection, but they are not representative of your production traffic. Anthropic’s prompting guide recommends using xhigh or max only when your own evaluations demonstrate a meaningful gain.

Build a repeatable eval collection in Apidog:

  1. Store ANTHROPIC_API_KEY as an environment variable.
  2. Create 20 to 50 Messages API requests from real prompts, sanitized production incidents, or representative fixtures.
  3. Add assertions for:
    • HTTP status 200
    • A stop_reason other than "max_tokens" or "refusal"
    • Required answer properties, such as a valid classification, JSON field, or expected phrase
  4. Run the same collection at low, medium, and high effort.
  5. Record pass rate and values from usage.
  6. Duplicate the collection, replace the model with claude-sonnet-5-5, and compare results in the same project.

Changing effort invalidates the prompt cache, so evaluate each effort level independently.

curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-haiku-5-5",
    "max_tokens": 8000,
    "thinking": {"type": "adaptive"},
    "output_config": {"effort": "medium"},
    "messages": [
      {
        "role": "user",
        "content": "Classify this support ticket as billing, bug or feature request: ..."
      }
    ]
  }'
Enter fullscreen mode Exit fullscreen mode

When implementing the response parser, select blocks by type. A thinking block can appear before the final content block.

Also note these Haiku 5.5 API constraints:

  • Do not send temperature.
  • Do not send top_p.
  • Do not send top_k.
  • Do not send an assistant prefill.

Each of these returns a 400 response on Haiku 5.5.

For implementation details, see the Claude Haiku 5.5 API guide and Testing LLM Applications.

FAQ

Is Claude Haiku 5.5 better than Sonnet 5.5?

Not according to Anthropic’s published numbers. Sonnet 5.5 leads every launch-table row. Haiku only edges it on FrontierCode when both run at max: 46.4% versus 46.2%. See Haiku 5.5 vs Haiku 4.5 for the upgrade comparison.

What is Haiku 5.5’s SWE-bench score?

It scores 64.8 on SWE-Bench Pro, 83.7 on SWE-bench Multilingual, and 30.7 on SWE-bench Multimodal. All results use max effort.

What effort were the benchmarks run at?

Most published results use max effort and are averaged over five trials. The API default is medium, where GDPval-AA scores 1277 instead of 1620.

Are there independent Haiku 5.5 benchmarks?

Not yet. Artificial Analysis independently ran GDPval-AA and AA-Briefcase, but Anthropic published those scores. Artificial Analysis does not currently have a Haiku 5.5 page.

Is GPT-6 Luna cheaper?

Usually. Luna costs less at every GDPval-AA effort level and every OSWorld level except medium, where the models cost about the same. Haiku 5.5 scores higher on both benchmarks.

Next step

Start at medium, measure pass rate and token usage on your own requests, then increase effort only when the pass-rate improvement justifies the cost.

Download Apidog, build the evaluation collection, and keep the results in Apidog next to your Sonnet 5.5 baseline before routing production traffic.

Top comments (0)