Ll

Llmfao – Human-Ranked LLM Leaderboard with Sixty Models

Hacker News

Llmfao – Human-Ranked LLM Leaderboard with Sixty Models

In September 2023, I noticed a tweet [1] on difficulties with LLM evaluation, which resonated with me a lot. A bit later, I spotted a nice LLMonitor Benchmarks dataset [2] with a small set of prompts and a large set of model completions. I decided to make my attempt without running a comprehensive suite of hundreds of benchmarks: https://dustalov.github.io/llmfao/ I also wrote a detailed post describing the methodology and analysis: https://evalovernite.substack.com/p/llmfao-human-ranking [1]: https://twitter.com/_jasonwei/status/1707104739346043143 [2]: https://benchmarks.llmonitor.com/ Unfortunately, I did my analysis before the Mistral AI model was released, but published it after the model was released. I’d be happy to add it to the comparison if I had their completions.

Share card

Actual performance

2points
2comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, models · Missing: mac, agents, macos
78%78% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
75%75% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: lua, ide, io · Missing: https docs, excited, just released
46%46% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
43%43% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
36%36% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
15%15% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
1%1% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

De
Debategle – ranked 1v1 debates judged by an LLM30%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Debategle – ranked 1v1 debates judged by an LLM

Hacker News3
Fi
Five Thousand Novels, Ranked by Vividness43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Five Thousand Novels, Ranked by Vividness

Hacker News54
Th
The Instavest Leaderboard40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Instavest Leaderboard

Hacker News6
Ranked.Fun
Ranked.Fun51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Online leaderboard with your friends for your favorite games

Indie Hackers2$20/mogames
Ap
Apps ranked by valuation51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Apps ranked by valuation

Hacker News41
Ranked in Google
Ranked in Google59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Ranked in Google

Indie Hackers1$1/moadvertising
Fr
FreePoll – Ranked Voting for Teams53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

FreePoll – Ranked Voting for Teams

Hacker News6
Ev
Every "Launch HN" Ranked39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Every "Launch HN" Ranked

Hacker News1
ScoreLeader
ScoreLeader36%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Scoreboard & Leaderboard App

Product Hunt+6
Ce
Cerno – CAPTCHA that targets LLM reasoning, not human biology49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cerno – CAPTCHA that targets LLM reasoning, not human biology

Hacker News12