Τ³

Τ³-Bench is out – can agents handle complex docs and live calls?

Hacker News

Τ³-Bench is out – can agents handle complex docs and live calls?

τ-Bench is an open benchmark for evaluating AI agents on grounded, multi-turn customer service tasks with verifiable outcomes. It's been great to see the community adopt it since launch — this is now the third iteration. With τ³-Bench, we're extending it to two new settings: knowledge-intensive retrieval and full-duplex voice. τ-Knowledge: agents must navigate ~700 interconnected policy documents to complete multi-step tasks. Best frontier model (GPT-5.2, high reasoning) hits ~25%. The surprising part: even when you hand the model the exact documents it needs, performance only reaches ~40%. We found that the bottleneck isn't retrieval — it's reasoning over complex, interlinked policies and executing the right actions in the right order. τ-Voice: same grounded tasks, but over live full-duplex voice with realistic audio — accents, background noise, interruptions, compressed phone lines. Voice agents score 31–51% in clean audio conditions and 26–38% in realistic ones. A consistent failure pattern across providers (OpenAI, Gemini, xAI): agent mishears a name or email during authentication, and everything downstream fails. We also incorporated 75+ task fixes to the original airline, retail, and telecom domains — many based on community audits and PRs (including contributions from Amazon and Anthropic). We believe a benchmark is only as good as its maintenance, and we're grateful for the community's help improving it. Code and leaderboard are open — we'd welcome community submissions and feedback. Blog post (papers, code, leaderboard): https://sierra.ai/blog/bench-advancing-agent-benchmarking-to...

Share card

Actual performance

12points
1comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, model · Missing: mac, macos, cursor
98%98% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: including, gemini · Missing: supports, reddit linkedin, podcasting
89%89% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, io · Missing: https docs, excited, just released
57%57% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: calls · Missing: plus, platform, intuitive
43%43% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
39%39% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
24%24% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: audio · Missing: web3, chat, crypto
2%2% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Te
Text handle.it and it calls to make appts for you63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Text handle.it and it calls to make appts for you

Hacker News4
We
We manufacture complex movements for smartwatches51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

We manufacture complex movements for smartwatches

Hacker News1
MeetingFriends
MeetingFriends23%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

We handle the when. You handle the fun.

Indie Hackers
Fr
FrustratedMonkey – A tool to handle your frustration64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

FrustratedMonkey – A tool to handle your frustration

Hacker News5
Ha
Harris docs on TweedJS homepage36%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Harris docs on TweedJS homepage

Hacker News1
Ho
How to handle API downtime with 2 LOC66%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

How to handle API downtime with 2 LOC

Hacker News1
Hi
Hijax: Intercept Ajax Calls59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Hijax: Intercept Ajax Calls

Hacker News1
Du
DumbAss – be too stupid for complex tools57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

DumbAss – be too stupid for complex tools

Hacker News19
Reminders API
Reminders API43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A simple API to handle complex reminders.

Indie Hackers2apis
Ga
Game of Cubes Complex Behaviour48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Game of Cubes Complex Behaviour

Hacker News1