Reference-free evaluation of LLM-powered chatbots
Reference-free evaluation of LLM-powered chatbots
Hey HN! This an interactive demo with a *somewhat* helpful AI assistant. The goal is to demonstrate a good way to reference-free evaluate interactions between humans and AI assistants. Reference-free means that you do not provide a correct answer to a query. The used metric in this context is the goal success ratio, which measures how many queries a user needs to send to reach their goal. In the near future, there will be a guide on how to reference-free evaluate any LLM app (chat, RAG, summarization, etc.). Try it out and please share any feedback!
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
Dokimos – LLM evaluation framework for Java
Go-Mnemonic – Reference Implementation of a BIP-39 Go
Rues an Expression Evaluation Sidecar
BotEngine – chatbots for LiveChat
Chatbots for bloggers
Faster LLM evaluation with Bayesian optimization
Greenman is a British witchcraft reference app
DataTau is the reference newsboard for Data Scientists
Unity3D unassigned reference warnings at compile time
LLM-Powered Sysadmin