Be

Beval – Simple evaluations for your AI product

Hacker News

Beval – Simple evaluations for your AI product

I have been working on a web app called Beval - Simple evaluations for your AI product. In my day to day as a Product Manager working in a team that ships AI products, I often found myself wanting to do 'quick and dirty' LLM-based evaluation on conversation transcripts and traces. I didn't need anything fancy, just 'did the agent answer the question', 'did the agent cover the 5 things it needed to' - that type of thing. I found myself blocked by 'Gemini in Google Sheets', it was too slow and cumbersome, and it didn't handle eval changes well - particularly when trying to associate evals with ground truth. And because I was exploring or working on new and experimental features, it wasn't helpful to try and set up something more robust with the team. To fix the problem I eventually learned to call the OpenAI API in Python, but I really felt that I wanted a 'product' to help me and potentially help others who need answers fast - outside of building infrastructure and pipelines. So over the last few weeks I built: https://beval.space It has: - LLM-as-judge evals: boolean checks (yes/no), scores (1-5), categories, and freeform comments - Reusable eval definitions you can run across different datasets - Ground truth labelling so you can compare eval versions against human judgments - Per-trace reasoning so you can see why the judge scored something the way it did - An example dataset so you can try it without having your own traces ready One of our early users described it as 'quick n dirty evals when you don't want to touch a shit load of infra.' I'm trying to figure out if that's a common need or just a niche thing. Free during beta. Would love HN's take — what's missing, and would you actually use something like this?

Share card

Actual performance

2points
1comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agent, google, user · Missing: mac, agents, macos
87%87% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: gemini · Missing: supports, reddit linkedin, podcasting
79%79% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, pipe · Missing: https docs, excited, just released
51%51% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: google, answers, users · Missing: mobile apps, ios, personal
46%46% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: users · Missing: plus, platform, intuitive
26%26% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
21%21% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Sortlist
Sortlist50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI Product Organiser for Shoppers

Indie Hackerscommitment-side-project
ShopTag.ai
ShopTag.ai72%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI product tagging and categorization

Product Hunt+8
Yo
Your AI Product Manager38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Your AI Product Manager

Hacker News11
Pr
Prototyping with AI as a Product Manager38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Prototyping with AI as a Product Manager

Hacker News5
Iteraite
Iteraite38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Your AI Product Team

Indie Hackerscommitment-full-time
Mazora
Mazora29%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI Product Context Wizard

Indie Hackers1ai
BrainGrid
BrainGrid35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The AI Product Planner

Indie Hackerscommitment-full-time
MEHMARK
MEHMARK36%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Find the boredom tax in your AI product.

Indie Hackerscommitment-full-time
Polymet
Polymet65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI product designer

Product Hunt+622Design Tools
Aiproductphotography
Aiproductphotography60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI Product Photography

Indie Hackerscommitment-side-project