NEWτ³-bench is here: τ-knowledge evaluates agents on knowledge-intensive tasks, and τ-voice benchmarks real-time voice agents.

How τ-bench has evolved

τ-voiceτ³-benchMarch 2026Paper →Blog →

Real-time voice: full-duplex conversations with interruptions, accents, and background noise — the same task rigor, now in speech.

Voice mode:🛍️ Retail✈️ Airline📱 Telecom
τ-knowledgeτ³-benchMarch 2026Paper →Blog →

Knowledge-intensive tasks: agents retrieve and reason over a realistic knowledge base of ~700 documents, combining retrieval with policy application.

Adds domain:🏦 Banking
Task audit & fixesτ³-benchFebruary 2026Blog →

Audited and fixed 50+ tasks across the airline and retail domains — correcting expected actions, ambiguous instructions, and impossible constraints.

Updates:🛍️ Retail✈️ Airline
τ²-benchJune 2025Paper →Blog →

Dual control: the user can now act on the world too. Agents must guide users through steps only the user can perform, not just act on their behalf.

Adds domain:📱 Telecom
τ-benchJune 2024Paper →Blog →

The original benchmark for tool-agent-user interaction. Agents converse with a simulated user and call tools, scored against verifiable database outcomes with the passk reliability metric.

Domains:🛍️ Retail✈️ Airline