Kitten TTS Based Low-Latency Streaming Voice Assistant on CPU
Kitten TTS Based Low-Latency Streaming Voice Assistant on CPU
We asked Neo AI to build a small voice assistant pipeline that runs with low latency on CPU instead of requiring a GPU. The goal was to see how responsive a LLM → speech system can be on normal laptops or edge devices. It includes: - Voice Activity Detection - CPU-friendly LLM + TTS streaming - Async pipeline to reduce latency Modular LLM backend Useful for local assistants, robotics prototypes, privacy-first setups, or benchmarking STT/LLM/TTS latency. We’ve been experimenting with similar CPU-first pipelines inside NEO workflows for on-device agents, and this repo is a minimal standalone version. Would love suggestions on lightweight STT/TTS models or latency tricks people have used on CPU.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
Audio Repeater Pro – Low latency audio streaming tool
Low-Latency Streaming for Gamers & Creators
Bambuser – Ultra low latency live streaming SDK
Inworld TTS – high-quality, affordable, and low-latency TTS
Low-latency jamming over the internet
NoraSector – low-latency WebRTC police/fire/EMS scanner (Seattle)
Lenma – A Voice Assistant for macOS
LeagueOfLegends Voice Assistant and Context
Analog audio to Google Cast via low-latency streaming
stationary_vector: A Parallelizable, Low-Latency C++ Vector