Re

Reference-free evaluation of LLM-powered chatbots

Hacker News

Reference-free evaluation of LLM-powered chatbots

Hey HN! This an interactive demo with a *somewhat* helpful AI assistant. The goal is to demonstrate a good way to reference-free evaluate interactions between humans and AI assistants. Reference-free means that you do not provide a correct answer to a query. The used metric in this context is the goal success ratio, which measures how many queries a user needs to send to reach their goal. In the near future, there will be a guide on how to reference-free evaluate any LLM app (chat, RAG, summarization, etc.). Try it out and please share any feedback!

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: user, context · Missing: mac, agents, macos
66%66% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Hacker NewsMay not resonate with HN audience · Strong signals: lua, ide, io · Missing: https docs, excited, just released
49%49% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: way · Missing: mobile apps, ios, personal
47%47% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Indie HackersIH features products with proven revenue · Missing: supports, reddit linkedin, podcasting
30%30% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
18%18% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: active · Missing: arr, mrr, revenue
12%12% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
2%2% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Do
Dokimos – LLM evaluation framework for Java44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Dokimos – LLM evaluation framework for Java

Hacker News1
Go
Go-Mnemonic – Reference Implementation of a BIP-39 Go36%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Go-Mnemonic – Reference Implementation of a BIP-39 Go

Hacker News3
Ru
Rues an Expression Evaluation Sidecar56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Rues an Expression Evaluation Sidecar

Hacker News1
Bo
BotEngine – chatbots for LiveChat46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

BotEngine – chatbots for LiveChat

Hacker News17
swiftr
swiftr23%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Chatbots for bloggers

Indie Hackerscommitment-side-project
Fa
Faster LLM evaluation with Bayesian optimization50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Faster LLM evaluation with Bayesian optimization

Hacker News131
Greenman
Greenman39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Greenman is a British witchcraft reference app

Indie Hackers1ai
DataTau
DataTau58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

DataTau is the reference newsboard for Data Scientists

Indie Hackers2ai
Un
Unity3D unassigned reference warnings at compile time50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Unity3D unassigned reference warnings at compile time

Hacker News2
LL
LLM-Powered Sysadmin57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM-Powered Sysadmin

Hacker News1