Gauntlet
Frontier models run the gauntlet of prompt-injection attacks — ranked by how well they resist direct jailbreaks and indirect, tool-based injection. Higher robustness is better.
| Model | Robustness | ASR | Out tok | Avg cost |
|---|---|---|---|---|
1 claude-fable-5medium anthropic | 100% | 0% | 17k | $1.38 |
2 claude-fable-5low anthropic | 99% | 1% | 14k | $1.27 |
3 claude-fable-5high anthropic | 99% | 1% | 18k | $1.40 |
4 claude-opus-4.8high anthropic | 94% | 6% | 74k | $3.04 |
5 claude-opus-5low anthropic | 94% | 6% | 65k | $2.67 |
6 claude-opus-4.8low anthropic | 93% | 7% | 60k | $2.64 |
7 claude-opus-4.8medium anthropic | 93% | 7% | 69k | $2.89 |
8 claude-opus-5medium anthropic | 91% | 9% | 83k | $3.20 |
9 claude-opus-5high anthropic | 89% | 11% | 97k | $3.63 |
10 gpt-5.5high openai | 75% | 25% | 34k | $1.49 |
11 grok-4.3 x-ai | 73% | 27% | 82k | $0.43 |
12 gpt-5.5medium openai | 70% | 30% | 21k | $1.04 |
13 claude-sonnet-4.6high anthropic | 68% | 32% | 63k | $1.69 |
14 gpt-5.5low openai | 58% | 42% | 18k | $0.88 |
Every model faces the same frozen attack suite across two escalating threat models. Robustness is 100 − attack-success-rate, so higher means harder to break.
- Indirect injection
- Hand-built agentic scenarios where a malicious instruction hidden in tool output — a document, email, or API response — tries to hijack the model while it works a legitimate task.
- Agentic tool-hijacking
- ZeroLeaks' Sandbox attack corpus replayed against mock agent tools — tool-poisoning, authority spoofing, and protocol exploits — counting any dangerous tool call as a breach.
Robustness = 100 − attack-success-rate (ASR). Each response is scored by an LLM judge backed by success/failure indicators, so a model is only marked breached on a genuine rule violation — not on a refusal or an error. The chart plots robustness (%) against efficiency (cost per run or output tokens, reversed so more efficient is to the right); each line connects one model across reasoning efforts. Cost uses standard list prices. GPT-5.5 and Grok 4.3 are served via Azure. Latest run: suite indirect-v6+agentic, July 27, 2026.