In

Inworld TTS – high-quality, affordable, and low-latency TTS

Hacker News

Inworld TTS – high-quality, affordable, and low-latency TTS

Hi HN, Igor here, one of the engineers behind this project. High-quality voice APIs are usually either expensive, slow, or both. Cheaper and faster solutions very often lack realism. We decided to build Inworld TTS to bridge this gap. We just released two multilingual models. Our small model, named TTS-1, is on par with SOTA models quality-wise given objective metrics WER/SIM/DNSMOS. A larger model, TTS-1-Max, is even better. It can produce more nuanced speech and has ~3.5% better WER across all 11 supported languages averaged. Both models also support markup tags (e.g. prepend "[happy]" to the text to make the generation more enthusiastic, etc). The models are built with LLaMA 1B and 8B being the SpeechLM backbones for TTS-1 and TTS-1-Max respectively. We up-trained both models on a mixture of text and audio, then finetuned on text-audio pairs and polished final checkpoints with GRPO on a small high-quality dataset. Our Speech Lab team (4 MLEs) started to work on collecting audio data around late February and exploring different audio codec architectures. We got inspired by the simplicity of the single vector quantization Xcodec2 neural audio codec architecture used and decided to use a similar idea. We started the training early April. Once codec was ready, SpeechLMs’ training took another month and a half. We finished mid-June, all - using 32 H100 GPUs. To make models real-time ready during serving, we collaborated with Modular to migrate from vanilla vLLM solution to Mojo- written MAX server. Our bet of keeping serving architecture as simple as possible played out well: both models turned out to be really fast. TTS-1, which can be accessed via streaming API, has ~500ms p90 latency for returning the first ~2 seconds of audio. The pricing is simple, pay $5/1M characters. A larger model’s API access will be opened soon. We’ll share more details about serving performance optimizations made in the coming weeks. We are also about to release all the training, modeling, and benchmarking code on GitHub to be transparent about how we made it. This repo is very flexible and can easily be adjusted to train an arbitrary neural net, but we’ll release the code with the focus on speech modeling. By the way, we’ve used PyTorch Lightning as the framework for multi-node/multi-GPU training as it proved to be very easy-to-use and reliable. -- Check the TTS out at https://inworld.ai/tts Happy to answer any questions you have!

Share card

Actual performance

24points
15comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: started · Missing: supports, reddit linkedin, podcasting
95%95% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, models, single · Missing: mac, agents, macos
88%88% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: just released, llama, ide · Missing: https docs, excited, exist
71%71% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoStrong fit for a featured deal · Strong signals: soon · Missing: plus, platform, intuitive
59%59% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: month, way · Missing: mobile apps, ios, personal
49%49% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: training · Missing: arr, mrr, revenue
27%27% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: audio, collaborate · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Bath Remodeling San Jose
Bath Remodeling San Jose49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Affordable & High-Quality Bathroom Remodeling in San Jose

Indie Hackerscommitment-full-time
EasyTagger
EasyTagger7%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A low latency API for assigning high-quality topics to text

Indie Hackers2ai
Vu
VulcanSQL – Serve high-concurrency, low-latency API from OLAP79%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

VulcanSQL – Serve high-concurrency, low-latency API from OLAP

Hacker News8
Na
Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost81%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

Hacker News38
The Church Co
The Church Co66%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

High Quality, low cost websites for churches & nonprofits.

Indie Hackers189saas
Dr
Dreamweaver – high quality t-shirts made with AI42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Dreamweaver – high quality t-shirts made with AI

Hacker News5
Ba
Backprop GPU Cloud for affordable and high quality VMs73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Backprop GPU Cloud for affordable and high quality VMs

Hacker News3
Painting Services Qatar
Painting Services Qatar51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Affordable and High Quality Painting Services in Qatar

Indie Hackerscommitment-full-time
fananas for seo and backlinks
fananas for seo and backlinks69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fananas specializes in high-quality backlinks, enhancing onl

Indie Hackers1$1/moadvertising
Hi
High-Quality NSFW Video App55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

High-Quality NSFW Video App

Hacker News1