I

I built a small audit layer for LLM-as-judge decisions

Hacker News

I built a small audit layer for LLM-as-judge decisions

I made this while checking model graded answer and helped me to check the odd cases by hand. Not sure if it’s useful to anyone else. TL;DR: it breaks an LLM judge run into claims->evidence->verdicts and flags when a verdict is not supported by the evidence, so i can check it manually.

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model · Missing: mac, agents, macos
62%62% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
39%39% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
34%34% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Indie HackersIH features products with proven revenue · Missing: supports, reddit linkedin, podcasting
29%29% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: ide, io · Missing: https docs, excited, just released
20%20% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
19%19% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
2%2% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

je
jevals – replacing LLM judges with typed Jev decisions59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

jevals – replacing LLM judges with typed Jev decisions

Hacker News39
Psyloom
Psyloom24%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Psychological Layer for LLM

Indie Hackerscommitment-full-time
DecisionBear
DecisionBear28%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Decisions made easy | Audit trail for your decisions

Indie Hackers1productivity
I
I built a small OSS kernel for replaying and diffing AI decisions48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built a small OSS kernel for replaying and diffing AI decisions

Hacker News1
Gl
Glimpse the multiverse of outcomes before decisions using LLM37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Glimpse the multiverse of outcomes before decisions using LLM

Hacker News1
Bu
Built a small tool to hack my beliefs22%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Built a small tool to hack my beliefs

Hacker News1
A
A reliability layer that prevents LLM downtime and unpredictable cost39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A reliability layer that prevents LLM downtime and unpredictable cost

Hacker News1
Au
Audit your ISP with a speedtest cronjob40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Audit your ISP with a speedtest cronjob

Hacker News12
Re
Reliability layer to prevent LLM hallucinations31%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Reliability layer to prevent LLM hallucinations

Hacker News3
sm
small lisp interpreter in Haskell79%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

small lisp interpreter in Haskell

Hacker News25