Vo

Voice bots with 500ms response times

Hacker News

Voice bots with 500ms response times

Last year when GPT-4 was released I started making lots of little voice + LLM experiments. Voice interfaces are fun; there are several interesting new problem spaces to explore. I'm convinced that voice is going to be a bigger and bigger part of how we all interact with generative AI. But one thing that's hard, today, is building voice bots that respond as quickly as humans do in conversation. A 500ms voice-to-voice response time is just barely possible with today's AI models. You can get down to 500ms if you: host transcription, LLM inference, and voice generation all together in one place; are careful about how you route and pipeline all the data; and the gods of both wifi and vram caching smile on you. Here's a demo of a 500ms-capable voice bot, plus a container you can deploy to run it yourself on an A10/A100/H100 if you want to: https://fastvoiceagent.cerebrium.ai/ We've been collecting lots of metrics. Here are typical numbers (in milliseconds) for all the easily measurable parts of the voice-to-voice response cycle. macOS mic input 40 opus encoding 30 network stack and transit 10 packet handling 2 jitter buffer 40 opus decoding 30 transcription and endpointing 200 llm ttfb 100 sentence aggregation 100 tts ttfb 80 opus encoding 30 packet handling 2 network stack and transit 10 jitter buffer 40 opus decoding 30 macOS speaker output 15 ---------------------------------- total ms 759 Everything in AI is changing all the time. LLMs with native audio input and output capabilities will likely make it easier to build fast-responding voice bots soon. But for the moment, I think this is the fastest possible approach/tech stack.

Share card

Actual performance

315points
99comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, macos, agent · Missing: agents, cursor, claude
88%88% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: started · Missing: supports, reddit linkedin, podcasting
87%87% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
47%47% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: pipe, io · Missing: https docs, excited, just released
38%38% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: plus, host, interface · Missing: platform, intuitive, reviews
36%36% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
23%23% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: audio · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Ro
Roq – Demonstrating low μs response times for algorithmic trading54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Roq – Demonstrating low μs response times for algorithmic trading

Hacker News2
Av
Average support response times for 60+ bitcoin exchanges41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Average support response times for 60+ bitcoin exchanges

Hacker News2
Th
The Probability Times39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Probability Times

Hacker News20
Ho
How many more times will you see your mother?40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

How many more times will you see your mother?

Hacker News1
I
I built a platform to your improve efficiency and response times42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built a platform to your improve efficiency and response times

Hacker News2
Wh
Where's the Latency? Decompose API Response Times into 4 Parts63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Where's the Latency? Decompose API Response Times into 4 Parts

Hacker News1
Vo
Voicegain RTC Callback API for IVR and voice bots released58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Voicegain RTC Callback API for IVR and voice bots released

Hacker News1
Sh
Shortwave.ai Email Bots – Send an email. Get a helpful response53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Shortwave.ai Email Bots – Send an email. Get a helpful response

Hacker News4
InnovaBot
InnovaBot39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

BOTS

Product Hunt+8
Co
Code in Response to “The Trouble with Symlinks.”53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Code in Response to “The Trouble with Symlinks.”

Hacker News5