Claude Haiku 5.5 scores 72.4% on OSWorld 2.1, 1620 on GDPval-AA v2.1, 46.4% on FrontierCode 1.1, and 39.2% on Terminal-Bench 4.0. Haiku 4.5 scored 15.7%, 735, and 0.0% on three of those. These are max-effort numbers; at the default medium effort, GDPval-AA drops to 1277. No independent lab has published Haiku 5.5 results yet.
This guide breaks down every published Claude Haiku 5.5 benchmark, identifies who ran each test, shows how effort changes score and cost, and provides a practical workflow for testing the model against your own traffic in Apidog. For specifications and pricing context, start with What Is Claude Haiku 5.5.
The launch table
Every Haiku 5.5 result below uses adaptive thinking at max effort, usually averaged across five trials. Terminal-Bench used 10 trials per task. “n/r” means no published result.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (Elo) | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 (Elo) | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1, offline subset | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity’s Last Exam, no tools | 45.9% | 10.2% | n/r | 56.9% |
| Humanity’s Last Exam, with tools | 57.4% | 18.7% | n/r | 64.5% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 (Main) | 46.4% | n/r | 42.4% | 52.1% (xhigh) |
| Chartography, no tools | 46.4% | 6.4% | 29.1% | 61.6% |
Who ran each benchmark
Treat cross-vendor comparisons carefully: sharing a benchmark name does not guarantee identical harnesses, prompts, or effort settings.
| Benchmark | Who ran it | Conditions |
|---|---|---|
| GDPval-AA v2.1, AA-Briefcase v1.1 | Artificial Analysis, independently | GDPval-AA: 220 tasks across 44 occupations; Elo anchored to DeepSeek V4.1 Flash (max) at 1600. Briefcase uses multi-week projects. |
| FrontierCode 1.1 | Cognition | Claude ran in Claude Code; GPT ran in Codex. Main is the hardest 100 of 150 tasks. |
| Chartography | Anthropic, using Surge AI’s benchmark | Gemini 3.5 Flash grader; Luna score supplied by Surge AI. |
| Terminal-Bench 4.0 | Anthropic | Claude Code --bare, 66 tasks, 10 trials per task, safeguards enabled. |
| OSWorld 2.1 | Anthropic | 82 of 108 tasks, 1080p resolution, up to 500 steps. |
| Humanity’s Last Exam | Anthropic | 980K-token task budget; Opus 4.6 grader. |
Two GPT-6 Luna values need additional context:
- Anthropic ran Luna’s OSWorld score through OpenAI’s API on the same 82 tasks.
- Luna’s 16.4% Terminal-Bench score comes from the public leaderboard using Codex CLI at max effort.
- On Terminal-Bench, safeguards stopped 1.8% of Haiku 5.5 trials—12 of 660—and each stopped trial counted as a failure.
The FrontierCode comparison trap
The launch table compares Haiku 5.5 at max effort with Sonnet 5.5 at xhigh, which is Sonnet’s best published setting. The system card provides the matched-effort comparison.
| FrontierCode 1.1 | Main | Extended |
|---|---|---|
| Haiku 5.5, max | 46.4% | 58.4% |
| Haiku 5.5, xhigh | 45.8% | n/r |
| Sonnet 5.5, max | 46.2% | 59.1% |
| Sonnet 5.5, xhigh | 52.1% | 64.4% |
At matched max effort, Haiku 5.5 narrowly leads Sonnet 5.5 on FrontierCode Main: 46.4% versus 46.2%.
At each model’s best setting, Sonnet leads by 5.7 percentage points. When reporting FrontierCode results, always include the effort setting.
Additional system card benchmarks
The system card includes benchmarks omitted from the launch post. All values use max effort unless stated otherwise.
| Benchmark | Haiku 5.5 | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|
| SWE-Bench Pro | 64.8 | n/r | 81.3 |
| SWE-bench Multilingual | 83.7 | 67.4 | 90.3 |
| SWE-bench Multimodal | 30.7 | 19.8 | 54.3 |
| OSWorld 2.1, strict pass rate | 37.1% | n/r | 48.8% |
| Chartography, with tools | 86.2% | 8.8% | 90.2% |
| OfficeQA / OfficeQA Pro | 73.5% / 60.3% | 63.0% / 47.1% | 76.9% / 65.6% |
| HealthBench Professional, length-adjusted | 64.8% | 32.2% | 69.2% |
Two implementation takeaways stand out:
- The 72.4% OSWorld result is partial credit. The strict pass rate is only 37.1%.
- Chartography rises from 46.4% without tools to 86.2% with tools. For chart-related workflows, give Haiku 5.5 access to the required tools rather than relying on text-only reasoning.
Default effort scores lower
Haiku 5.5 defaults to medium effort in the Claude API and Claude Code. Do not expect max-effort benchmark results unless you explicitly configure the model.
| Benchmark | Medium (default) | Max | Note |
|---|---|---|---|
| GDPval-AA v2.1 | 1277 | 1620 | Medium used about one tenth of max’s output tokens. |
| AA-Briefcase v1.1 | 1372 | 1578 | Medium used under one quarter of max’s output tokens. |
| HealthBench Professional | 59.9% | 64.8% | Low: 57.9%; high: 61.3%. |
If you omit output_config.effort, expect the medium column. See the Claude effort documentation for all five levels.
Score and cost by effort
The following launch-chart values pair score with cost. OSWorld and Terminal-Bench costs are per attempt; GDPval-AA costs are per task.
| Model | Low | Medium | High | Xhigh | Max |
|---|---|---|---|---|---|
| OSWorld 2.1 | |||||
| Haiku 5.5 | 42.0% ($0.07) | 53.3% ($0.13) | 61.3% ($0.18) | 67.6% ($0.28) | 72.4% ($0.61) |
| Sonnet 5.5 | 57.9% ($0.68) | 66.0% ($0.93) | 73.2% ($1.38) | 81.1% ($2.22) | 83.9% ($5.73) |
| GPT-6 Luna | 19.2% ($0.04) | 37.5% ($0.13) | 42.3% ($0.14) | 44.8% ($0.17) | 48.9% ($0.21) |
| GDPval-AA v2.1 | |||||
| Haiku 5.5 | 1125 ($0.012) | 1277 ($0.030) | 1420 ($0.089) | 1513 ($0.27) | 1620 ($0.87) |
| Sonnet 5.5 | 1179 ($0.22) | 1324 ($0.27) | 1551 ($0.62) | 1731 ($1.88) | 1840 ($6.78) |
| GPT-6 Luna | 1036 ($0.004) | 1262 ($0.02) | 1344 ($0.03) | 1364 ($0.05) | 1437 ($0.09) |
| Terminal-Bench 4.0 | |||||
| Haiku 5.5 | 12.7% ($0.42) | 20.3% ($0.68) | 24.8% ($1.04) | 31.5% ($1.75) | 39.2% ($2.64) |
| Sonnet 5.5 | 20.0% ($0.62) | 28.8% ($0.68) | 43.0% ($1.46) | 61.5% ($4.34) | 70.6% ($10.44) |
Use this table to choose an initial production setting:
-
Start with
medium. On GDPval-AA, moving frommediumtomaxcosts roughly 28x more for 343 additional Elo points. -
Validate the final increments. On OSWorld, moving from
xhightomaxmore than doubles cost for a 4.8-point gain. -
Compare equivalent budgets, not just effort labels. Haiku 5.5 at
mediumbeats Luna atmaxon OSWorld: 53.3% for $0.13 versus 48.9% for $0.21. -
Use Sonnet for demanding terminal tasks. Sonnet 5.5 at
highscores 43.0% on Terminal-Bench for $1.46, exceeding Haiku 5.5 atmax—39.2% for $2.64.
Luna costs less at every other matching level, although at medium the two models cost about the same. See Claude Haiku 5.5 vs GPT-6 Luna.
Costs use list prices: $0.10/$0.50 per million tokens for prompts up to 100K tokens, with higher pricing above that threshold. See the Claude Haiku 5.5 pricing breakdown.
Where Haiku 5.5 still trails
Sonnet 5.5 leads every launch-table row. Haiku 5.5’s only direct win is FrontierCode at matched max effort.
The largest gap is Terminal-Bench 4.0:
- Haiku 5.5: 39.2%
- Sonnet 5.5: 70.6%
The same pattern appears in SWE-Bench Pro—64.8 versus 81.3—and SWE-bench Multimodal—30.7 versus 54.3.
Anthropic states that Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding work. Haiku 5.5 is more appropriate for narrowly scoped tasks such as:
- Context compaction
- Summarization
- Classification
- Browser use
- Subagents coordinated by a larger model
For a practical subagent setup, see Haiku 5.5 in Claude Code.
Independent results: none yet
As of October 8, 2026, no independent lab has published Haiku 5.5 benchmark results.
Artificial Analysis’s Haiku 5.5 page returns a 404, and its leaderboard lists only Claude 4.5 Haiku. Vals, LMArena, SWE-bench, and Aider did not show results either.
There is also no third-party speed figure. Anthropic calls Haiku 5.5 its “fastest model to date” at standard speed, while noting that it is slower than Opus in Fast Mode.
The closest external result comes from Cursor. Its Claude Haiku 5.5 model documentation reports:
- 48.4% on CursorBench at max effort
- 30.9% at low effort
- With thinking disabled, 22.1% to 26.2% across the same effort levels
- With thinking enabled, 30.9% to 42.3% across those levels
What customers report
The following claims come from Anthropic’s launch post and have not been independently verified:
- HubSpot: 92.8% averaged across three runs on its simulated CRM portal suite, its best result on that suite.
- AlphaSense: 0.84 versus Haiku 4.5’s 0.76 on 400 Ask in Document queries, described as statistically significant.
- Box: 11 points higher than Haiku 4.5 at about half the latency in early testing.
- Asana: Over 30% lower latency for task completions compared with its current model.
- Cognition: Devin Fusion retains a FrontierCode score of 66.2 using Haiku 5.5 as a sidekick and Opus 5.5 as lead. This is a two-model system, not a Haiku-only result.
Run your own eval
Benchmarks are useful for model selection, but they are not representative of your production traffic. Anthropic’s prompting guide recommends using xhigh or max only when your own evaluations demonstrate a meaningful gain.
Build a repeatable eval collection in Apidog:
- Store
ANTHROPIC_API_KEYas an environment variable. - Create 20 to 50 Messages API requests from real prompts, sanitized production incidents, or representative fixtures.
- Add assertions for:
- HTTP status
200 - A
stop_reasonother than"max_tokens"or"refusal" - Required answer properties, such as a valid classification, JSON field, or expected phrase
- HTTP status
- Run the same collection at
low,medium, andhigheffort. - Record pass rate and values from
usage. - Duplicate the collection, replace the model with
claude-sonnet-5-5, and compare results in the same project.
Changing effort invalidates the prompt cache, so evaluate each effort level independently.
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-haiku-5-5",
"max_tokens": 8000,
"thinking": {"type": "adaptive"},
"output_config": {"effort": "medium"},
"messages": [
{
"role": "user",
"content": "Classify this support ticket as billing, bug or feature request: ..."
}
]
}'
When implementing the response parser, select blocks by type. A thinking block can appear before the final content block.
Also note these Haiku 5.5 API constraints:
- Do not send
temperature. - Do not send
top_p. - Do not send
top_k. - Do not send an assistant prefill.
Each of these returns a 400 response on Haiku 5.5.
For implementation details, see the Claude Haiku 5.5 API guide and Testing LLM Applications.
FAQ
Is Claude Haiku 5.5 better than Sonnet 5.5?
Not according to Anthropic’s published numbers. Sonnet 5.5 leads every launch-table row. Haiku only edges it on FrontierCode when both run at max: 46.4% versus 46.2%. See Haiku 5.5 vs Haiku 4.5 for the upgrade comparison.
What is Haiku 5.5’s SWE-bench score?
It scores 64.8 on SWE-Bench Pro, 83.7 on SWE-bench Multilingual, and 30.7 on SWE-bench Multimodal. All results use max effort.
What effort were the benchmarks run at?
Most published results use max effort and are averaged over five trials. The API default is medium, where GDPval-AA scores 1277 instead of 1620.
Are there independent Haiku 5.5 benchmarks?
Not yet. Artificial Analysis independently ran GDPval-AA and AA-Briefcase, but Anthropic published those scores. Artificial Analysis does not currently have a Haiku 5.5 page.
Is GPT-6 Luna cheaper?
Usually. Luna costs less at every GDPval-AA effort level and every OSWorld level except medium, where the models cost about the same. Haiku 5.5 scores higher on both benchmarks.
Next step
Start at medium, measure pass rate and token usage on your own requests, then increase effort only when the pass-rate improvement justifies the cost.
Download Apidog, build the evaluation collection, and keep the results in Apidog next to your Sonnet 5.5 baseline before routing production traffic.
Top comments (0)