I built a small audit layer for LLM-as-judge decisions
I built a small audit layer for LLM-as-judge decisions
I made this while checking model graded answer and helped me to check the odd cases by hand. Not sure if it’s useful to anyone else. TL;DR: it breaks an LLM judge run into claims->evidence->verdicts and flags when a verdict is not supported by the evidence, so i can check it manually.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
jevals – replacing LLM judges with typed Jev decisions
Psychological Layer for LLM
Decisions made easy | Audit trail for your decisions
I built a small OSS kernel for replaying and diffing AI decisions
Glimpse the multiverse of outcomes before decisions using LLM
Built a small tool to hack my beliefs
A reliability layer that prevents LLM downtime and unpredictable cost
Audit your ISP with a speedtest cronjob
Reliability layer to prevent LLM hallucinations
small lisp interpreter in Haskell