Op

Open-source simulation testing infra for voice agents

Hacker News

Open-source simulation testing infra for voice agents

Hey HN, we’re Nischal & Naman. We’re brothers, and together we’re building an open-source platform for simulation based testing of voice agents (try it out in 5 mins - https://docs.egma.ai/docs/get-started/quickstart , 2 min demo video - https://youtu.be/wgDWEe5UAUY ) Platforms that help you do simulation testing already exist. But they all charge a heavy premium on top of inference costs. We believe if the industry truly wants to scale simulation testing of voice agents, we need to stop charging a premium on inference and start providing infrastructure to scale simulations. The project is built in a way that allows you to bring your own STT, LLM, and TTS provider keys while also providing a way for us to provide inference directly within the platform at a 0% markup so you don’t have to wrestle with multiple keys/ rate limits. A bit on why’re we’re building this - We started working on voice agents in early 2024 and since then, have worked on numerous voice ai systems - like screenless voice-powered hardware for kids[1], AI receptionist deployed in healthcare practices like med spas & therapy clinics and personal accountability coaches. Some of these were full-fledged startups others were just side projects. But we kept on encountering similar issues all throughout. One of the most frustrating parts was calling the agent again and again, reciting the same script to test its behavior. We also kept on encountering new issues all the time in production that we couldn't have simulated pre launch. This frustration led us to a bigger question: how can developers trust the voice agents they’re shipping? We believe building that trust takes two things - - First, you need a way to test the major scenarios your agent will face in production before you release it, without having to make every call yourself. But you can’t test for everything. Real world is too messy to predict in advance. - Which brings us to the second: you need a way to find issues once your agent is in production, whether it’s handling a few dozen conversations or millions. Detect drift in known behaviors AND surface unknown-unknown ways in which your agent is going wrong. We think it’s a really hard problem to solve. And having felt it firsthand, we’re deeply motivated to take a shot at it. Today we’re launching the simulation testing side of the platform. We feel it’s mature enough that real teams can depend on it. For eval design, we took inspiration from anthropic’s evals design[2] and extended it to voice systems. For technical & business model design we borrowed ideas from Langfuse[3]. Our stack is postgres, clickhouse & minio. Its easy to self-host[4] & the code has a permissive MIT license. We also have managed cloud version. We’ve written more about our [testing]( https://docs.egma.ai/docs/core-philosophies/testing-philosop... ) and [monitoring]( https://docs.egma.ai/docs/core-philosophies/monitoring-philo... ) philosophy in the docs. We’d love feedback from the HN community and people building voice agents - how are you testing today, what’s working and what’s frustrating? We’ll be in the comments. Thanks! [1] https://x.com/theBhulawat/status/1966200231705595932?s=20 [2] https://www.anthropic.com/engineering/demystifying-evals-for... [3] https://github.com/langfuse/langfuse [4] https://docs.egma.ai/self-hosting/get-started

Share card

Actual performance

16points
6comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, model · Missing: mac, macos, cursor
96%96% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: started, ios · Missing: supports, reddit linkedin, podcasting
93%93% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: exist, ide, clickhouse · Missing: https docs, excited, just released
81%81% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: ios, personal, video · Missing: mobile apps, entrepreneurs, apps
57%57% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform, host · Missing: plus, intuitive, reviews
35%35% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
21%21% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: real world · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Fi
Fixa – an open source Python package for testing voice agents56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fixa – an open source Python package for testing voice agents

Hacker News15
An
An open-source ELO benchmark for voice agents52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An open-source ELO benchmark for voice agents

Hacker News8
Egma AI
Egma AI20%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OSS Stack for simulation testing & monitoring voice agents

Indie Hackerscommitment-full-time
Op
Open source balloon simulation with Three.js72%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open source balloon simulation with Three.js

Hacker News307
Vocera
Vocera82%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Launch voice agents faster with simulation & monitoring

Product Hunt+583Analytics
Mu
Mutatr – an open source A/B testing agent42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Mutatr – an open source A/B testing agent

Hacker News3
Op
Open-source A/B testing platform61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source A/B testing platform

Hacker News11
Au
Autumn – Open-source infra over Stripe60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Autumn – Open-source infra over Stripe

Hacker News141
Ma
MarinaBox: Open-Source Sandbox Infra for AI Agents78%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

MarinaBox: Open-Source Sandbox Infra for AI Agents

Hacker News6
open source testing tool
open source testing tool23%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Best Open Source Testing Tool

Indie Hackers