Can AI agents reliably complete real-world tasks? τ-bench measures how well agents converse with users, call tools, retrieve knowledge, and follow policy across enterprise domains — in text and voice.
| # | Model | Pass^1 |
|---|---|---|
| 🥇 | GPT-5.5OpenAI | 46.4% |
| 🥈 | GPT-5.4OpenAI | 39.4% |
| 🥉 | GPT-5.2OpenAI | 32.2% |
| # | Model | Pass^1 |
|---|---|---|
| 🥇 | grok-voice-think-fast-1.0xAI | 67.3% |
| 🥈 | gemini-3.1-flash-live-preview-thinking-highGoogle | 43.8% |
| 🥉 | gpt-realtime-2openai | 42.4% |
| # | Model | Pass^1 |
|---|---|---|
| 🥇 | GLM-5.2Z.ai | 90.9% |
| 🥈 | Qwen3.5-397B-A17BAlibaba Cloud | 87.9% |
| 🥉 | Gemini 3.0 ProGoogle | 85.4% |
Audited and fixed 50+ tasks across the airline and retail domains — correcting expected actions, ambiguous instructions, and impossible constraints.