E2

E2E Testing for Chatbots

Hacker News

E2E Testing for Chatbots

Hi HN, Tired of shipping chatbot features based on gut feelings and "it seems to work" manual testing? We built SigmaEval, an open-source Python library that brings statistical rigor to testing conversational AI. Instead of simple pass/fail checks, SigmaEval uses an AI User Simulator and an AI Judge to let you make data-driven statements like: "We're confident that at least 90% of user issues will be resolved with a quality score of 8/10 or higher." This allows you to set and enforce objective quality bars for your AI's behavior, response latency, and more, directly within your existing Pytest/Unittest suites. It's built on top of LiteLLM to support 100+ LLM providers and is licensed under Apache 2.0. We just launched and would love to get your feedback. GitHub: https://github.com/Itura-AI/SigmaEval Docs: https://docs.sigmaeval.com/

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: user, open · Missing: mac, agents, macos
73%73% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
66%66% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: exist, existing, ide · Missing: https docs, excited, just released
59%59% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
48%48% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
34%34% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
17%17% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Ma
Manner – A/B testing and personalisation platform for chatbots41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Manner – A/B testing and personalisation platform for chatbots

Hacker News2
Bo
BotEngine – chatbots for LiveChat46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

BotEngine – chatbots for LiveChat

Hacker News17
swiftr
swiftr23%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Chatbots for bloggers

Indie Hackerscommitment-side-project
inchat
inchat22%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Chatbots for qualitative data gathering

Indie Hackers
Co
ContextCheck – Open-source tool for testing LLMs, RAGs and Chatbots54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ContextCheck – Open-source tool for testing LLMs, RAGs and Chatbots

Hacker News4
Wonop
Wonop72%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Chatbots for developers with deadlines

Indie Hackers3ai
Li
LinkedIn for Chatbots34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LinkedIn for Chatbots

Hacker News42
Presbot
Presbot50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Linkedin for Chatbots

Indie Hackers3ai
Ch
ChefSpec – RSpec testing for Chef Cookbooks33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ChefSpec – RSpec testing for Chef Cookbooks

Hacker News3
A/
A/B testing baked into RequireJS33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A/B testing baked into RequireJS

Hacker News1