Op

Open-source turn detection model for voice AI

Hacker News

Open-source turn detection model for voice AI

Hey HN, it’s Russ - cofounder of LiveKit. An open source stack for building realtime AI applications. We’re sharing our first homegrown AI model for turn detection. Here’s a live demo: https://cerebras.vercel.app/ Voice AI has come a long way in the last year. We now have end-to-end systems that can generate a response to user input in 300-500ms — human level speeds! As latency reduces, a common problem that surfaces is the LLM responds too quickly. Any time there’s a short pause in a user’s speech, it ends up interrupting them. This is largely due to how voice AI applications perform “turn detection” — that is, figuring out when the user has finished speaking and when the model can run inference and respond. Pretty much everyone uses a signal processing technique called voice activity detection (VAD). In a nutshell, it figures out when the audio signal switches from speech to silence and then triggers an end of turn once a configurable amount of silence has transpired. One obvious delta between VAD and how humans do turn detection is we also consider the content of speech (i.e. what someone says). These past few months, we’ve been working on an open weights, content-aware turn detection model for voice AI applications. It was fine-tuned from SmolLM v2 on text, runs on CPU (currently takes 50ms for inference), and uses speech transcriptions as input to predict when a user has completed a thought (also called an “utterance”). Since it was trained on text, notably it works well for pipeline-based architectures (i.e. STT ⇒ LLM ⇒ TTS). We use this model together with VAD to make better predictions about whether a user is done speaking. Here’s some demos -- - Podcast interview: https://youtu.be/EYDrSSEP0h0 - Ordering food: https://youtu.be/fcr8Y-3c4E0 - Providing shipping address: https://youtu.be/2pQWvd6xozw - Customer support: https://youtu.be/YoSRg3ORKtQ In our testing we’ve found: - 85% reduction in unintentional interruptions - 3% false positives (where the user is done speaking, but the model thinks they aren’t) In practice, we still have work to do. We currently delay inference if the model predicts a < 15% chance the user is done speaking. This threshold misses a bunch of middle-of-the-pack probabilities. Next steps are improving the model accuracy, tuning performance, and expanding to support more languages (only supports English rn). Separately, we’re starting to explore an audio-based model that considers not just what someone says but how they say it, which can be used with natively multimodal models like GPT-4o that directly process and generate audio. Code here: https://github.com/livekit/agents/tree/main/livekit-plugins/... Let us know what you think!

Share card

Actual performance

8points
1comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, model · Missing: mac, macos, cursor
99%99% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: supports, para · Missing: reddit linkedin, podcasting, created
87%87% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: open source, ide, pipe · Missing: https docs, excited, just released
78%78% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: month, way, para · Missing: mobile apps, ios, personal
52%52% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
29%29% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
17%17% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: audio · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Oc
Octo.ai, Open source analytics hypervisor54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Octo.ai, Open source analytics hypervisor

Hacker News67
Comp AI
Comp AI77%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The open source Vanta & Drata alternative

Product Hunt+612Open Source
Op
Open-Source SDXL UI – Invoke AI 3.0.162%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-Source SDXL UI – Invoke AI 3.0.1

Hacker News1
Ec
EchoKit – open-source voice AI agent framework58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

EchoKit – open-source voice AI agent framework

Hacker News3
SurfaceBrief AI
SurfaceBrief AI43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open Source Intelligence Automation with AI

Indie Hackers1$2/moai
Vo
Voxos.ai – An Open-Source Desktop Voice Assistant56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Voxos.ai – An Open-Source Desktop Voice Assistant

Hacker News123
In
Invoke AI 3.1 – open-source Workflows and Canvas for SDXL61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Invoke AI 3.1 – open-source Workflows and Canvas for SDXL

Hacker News1
ShareAI
ShareAI67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Think Uber, but for AI Open Source Models

Indie Hackers1ai
se
secinsights.ai – An open-source full-stack app using LlamaIndex53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

secinsights.ai – An open-source full-stack app using LlamaIndex

Hacker News7
Ta
TalkForm AI – open source conversational forms64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

TalkForm AI – open source conversational forms

Hacker News5