Op

Open Evaluation

Hacker News

Open Evaluation

Hey HN, this will likely interest you if you're a) into dense data visualization or b) trying to figure out how to measure the quality of a retrieval-augmented generation (RAG) system. There's an OSS tool called Open RAG Eval that analyzes RAG-based query-and-answer sets to generate a dense set of metrics in an "evaluation report". This report is in CSV format and the data is basically impossible for a human to read because there's so much of it. I built Open Evaluation to enable folks to load in a report and visualize the evaluation metrics in a more human-readable way. The challenge was the sheer amount of information to visualize. I went with a collapsible table with sticky headers to presenting the info, so you can compare metrics across reports and questions. I also tried to make everything clickable, so if you want to understand the meaning behind a metric you can just click it to open up an info panel to learn more about it. The site has built-in sample evaluation reports, so you can try it out without needing to generate your own reports. If you give it a shot please share your feedback. I'd love to find ways to make this more usable. Full disclosure: I did this for work and my coworkers also made Open RAG Eval.

Share card

Actual performance

3points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: visual, open · Missing: mac, agents, macos
86%86% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
69%69% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, io · Missing: https docs, excited, just released
55%55% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
39%39% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: visualize, way · Missing: mobile apps, ios, personal
35%35% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
18%18% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Ru
Rues an Expression Evaluation Sidecar56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Rues an Expression Evaluation Sidecar

Hacker News1
Op
Opik, an open source LLM evaluation framework79%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Opik, an open source LLM evaluation framework

Hacker News86
Vi
Visualizing arithmetic and logical expression evaluation in Swift61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Visualizing arithmetic and logical expression evaluation in Swift

Hacker News3
An
An Empirical Evaluation of Linear Probing Algorithms37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An Empirical Evaluation of Linear Probing Algorithms

Hacker News14
Co
Convert VHDL to Verilog using GHDL (+ first evaluation)51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Convert VHDL to Verilog using GHDL (+ first evaluation)

Hacker News2
Op
Open-Source RAG Evaluation Toolkit64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-Source RAG Evaluation Toolkit

Hacker News6
Flapico
Flapico57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Prompt versioning, testing, and evaluation

Product Hunt+149Developer Tools
Intellirate
Intellirate27%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The AI intelligence evaluation company

Indie Hackerscommitment-full-time
Sp
Spellout – Simple expressions (set) evaluation microservice55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Spellout – Simple expressions (set) evaluation microservice

Hacker News1
Sp
Spellout – Simple expressions (set) evaluation microservice55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Spellout – Simple expressions (set) evaluation microservice

Hacker News2