Na

Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

Hacker News

Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

Hey HN, Toby from Nari Labs here. We've been working on making OSS speech models super-fast. Last year, we built Dia, the first OSS text-to-speech model capable of doing natural dialogue. Since then, so many more great speech models have been released to the public. But the market is still dominated by closed source models. We think that's an inference problem. Existing systems such as vLLM / SGLang are not well suited for multimodal inference. To prove this, we built an inference engine specialized for Qwen3-TTS and open-sourced it ( https://github.com/nari-labs/nari-qwen3-tts ). Running at sub-50 ms latency at 10 RPS, this showed open models can be run much faster and cheaper. Since then, we've been working hard to bring cheap, fast, and high quality serving to all. And we've even beat closed models at their game! Measured on the highly cited Coval (YC S24) voice AI benchmarks, our Qwen3-TTS endpoint not just is #2 in latency, but #1 in accuracy (WER) compared to 11Labs, Cartesia etc. while being the cheapest endpoint. Our Qwen3-ASR endpoint has the lowest latency and #2 accuracy, just 0.1% away from #1. It is the second cheapest model on the list. It took a lot of clever inference engineering to make these models quick, perform well while keeping costs low. Interestingly, Alibaba's official endpoints seem to perform worse in terms of accuracy and latency compared to ours. But nonetheless, much love to the Qwen team for OSS-ing these amazing speech models. We want to continue to push prices down to make speech technology a commodity - so that every app can have great TTS and STT without worrying about unit costs. We're also working on other parts of audio such as diarization - as well as video and world model inference. More to come!

Share card

Actual performance

38points
10comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, models, open · Missing: mac, agents, macos
94%94% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
89%89% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: exist, existing, ide · Missing: https docs, excited, just released
80%80% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: video, way · Missing: mobile apps, ios, personal
51%51% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoStrong fit for a featured deal · Missing: plus, platform, intuitive
50%50% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
21%21% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: audio · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Vu
VulcanSQL – Serve high-concurrency, low-latency API from OLAP79%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

VulcanSQL – Serve high-concurrency, low-latency API from OLAP

Hacker News8
In
Inworld TTS – high-quality, affordable, and low-latency TTS70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Inworld TTS – high-quality, affordable, and low-latency TTS

Hacker News24
aptselect
aptselect43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Map LLM Latency, Cost, and Accuracy.

Indie Hackerscommitment-full-time
Lo
Low-latency jamming over the internet74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Low-latency jamming over the internet

Hacker News236
No
NoraSector – low-latency WebRTC police/fire/EMS scanner (Seattle)75%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

NoraSector – low-latency WebRTC police/fire/EMS scanner (Seattle)

Hacker News3
st
stationary_vector: A Parallelizable, Low-Latency C++ Vector71%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

stationary_vector: A Parallelizable, Low-Latency C++ Vector

Hacker News2
Gl
Global Low Latency APIs68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Global Low Latency APIs

Hacker News2
Sw
Swift-Qwen3.8-27B, -58.3% thinking, x1.95 speed, accuracy of xhigh63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Swift-Qwen3.8-27B, -58.3% thinking, x1.95 speed, accuracy of xhigh

Hacker News29
Ki
Kitten TTS Based Low-Latency Streaming Voice Assistant on CPU66%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Kitten TTS Based Low-Latency Streaming Voice Assistant on CPU

Hacker News3
EasyTagger
EasyTagger7%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A low latency API for assigning high-quality topics to text

Indie Hackers2ai