Fa

Faster LLM evaluation with Bayesian optimization

Hacker News

Faster LLM evaluation with Bayesian optimization

Recently I've been working on making LLM evaluations fast by using bayesian optimization to select a sensible subset. Bayesian optimization is used because it’s good for exploration / exploitation of expensive black box (paraphrase, LLM). I would love to hear your thoughts and suggestions on this!

Share card

Actual performance

131points
43comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: using · Missing: mac, agents, macos
73%73% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Hacker NewsStrong engagement from HN community · Strong signals: lua, io · Missing: https docs, excited, just released
53%53% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
46%46% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: para · Missing: mobile apps, ios, personal
44%44% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Indie HackersIH features products with proven revenue · Strong signals: para · Missing: supports, reddit linkedin, podcasting
43%43% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
14%14% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
2%2% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Do
Dokimos – LLM evaluation framework for Java44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Dokimos – LLM evaluation framework for Java

Hacker News1
Ru
Rues an Expression Evaluation Sidecar56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Rues an Expression Evaluation Sidecar

Hacker News1
Op
Open Evaluation53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open Evaluation

Hacker News3
Wh
WhitestormJS r11: modularity, optimization for webpack and more!45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

WhitestormJS r11: modularity, optimization for webpack and more!

Hacker News1
Ge
Geopt – GEneric OPTimization by Genetically Evolved OPeration Trees42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Geopt – GEneric OPTimization by Genetically Evolved OPeration Trees

Hacker News1
Ta
Tail Recursion Optimization for the JVM55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tail Recursion Optimization for the JVM

Hacker News107
Li
Lizard Optimization59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Lizard Optimization

Hacker News1
Op
Opik, an open source LLM evaluation framework79%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Opik, an open source LLM evaluation framework

Hacker News86
Op
Opensource resume evaluation LLM agents45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Opensource resume evaluation LLM agents

Hacker News1
Joule
Joule44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI Optimization

Indie Hackers1ai