Gr

Grading Notes for LLM-as-Judge

Hacker News

Grading Notes for LLM-as-Judge

I've created a very simple package for the idea introduced in the Databricks blog post ( https://www.databricks.com/blog/enhancing-llm-as-a-judge-wit... ). It proved to be quite useful for the use-cases I've worked on since with grading notes you can leave small details on around domain concepts that the LLMs make mistakes on rather than have a full answer which consumes a lot more time labeling time. I'd like to learn more if such an approach or similar has been useful for others too.

Share card

Actual performance

2points
3comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: notes · Missing: mac, agents, macos
68%68% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: ide · Missing: https docs, excited, just released
66%66% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Indie HackersFits the IH revenue-focused audience · Strong signals: created, mistakes · Missing: supports, reddit linkedin, podcasting
53%53% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRFits verified-revenue profile · Missing: mobile apps, ios, personal
52%52% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
32%32% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
13%13% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: introduce · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Ri
RiteTag – Hashtag grading system46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

RiteTag – Hashtag grading system

Hacker News4
Hoverr
Hoverr46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Crowdsourced Pokemon grading game

Indie Hackerscommitment-full-time
Vi
Vinyl record and sleeve grading tool56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Vinyl record and sleeve grading tool

Hacker News63
Fr
Free Prompt Grading Tool30%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Free Prompt Grading Tool

Hacker News1
Pr
Programming Assignment Grading System63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Programming Assignment Grading System

Hacker News6
GradeProAI
GradeProAI25%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI grading assistant for Canvas & Brightspace

Indie Hackerscommitment-side-project
Re
Redis-LLM – Redis module integrates LLM with Redis45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Redis-LLM – Redis module integrates LLM with Redis

Hacker News2
Li
LitLLM the Spiciest LLM Wrapper34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LitLLM the Spiciest LLM Wrapper

Hacker News1
LL
LLM Reasonsers46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Reasonsers

Hacker News2
LLM Hotkey
LLM Hotkey39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.
TrustMRROther