Be

Better Than 4 But Not 5 – An LLM model ordering challenge

Hacker News

Better Than 4 But Not 5 – An LLM model ordering challenge

This simple game challenges you drag-and-drop a set of LLM models into the correct order (release date, input cost at launch, LM Arena score) using just their name. It's no secret that LLM model names are a bit of a mess, but when OpenAI decided to backtrack from 4.5 to 4.1, well, I just couldn't let it go. This was my first attempt at vibe coding something with VSCode Copilot Agent (w/ Claude 3.7 Sonnet). The agent got me like 95% functionality, but the moment I needed to crack open the code myself I couldn't handle the mess and had to clean it up. I am bad at letting go control enough for vibe coding. If anyone else has had that problem and gotten over it, I would love some helpful tips. I collected the data primarily by digging through Simon Willison's blog archives and snapshotting LMArena's leaderboard. If you spot any inaccuracies in the model data, I'd appreciate corrections with sources! Nothing fancy in the tech here, just a vanilla HTML/CSS/JS site hosted on GitHub Pages. GitHub link (MIT License): https://github.com/wspittman/BetterThan4ButNot5

Share card

Actual performance

1points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agent, claude, model · Missing: mac, agents, macos
95%95% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
67%67% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: ide, io · Missing: https docs, excited, just released
35%35% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: host · Missing: plus, platform, intuitive
31%31% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
27%27% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: ttm · Missing: arr, mrr, revenue
17%17% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Ga
Gandalf - LLM Prompt Injection Challenge24%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Gandalf - LLM Prompt Injection Challenge

Hacker News3
A
A Bruteforcer in Go for WarpWallet's 10BTC Challenge35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A Bruteforcer in Go for WarpWallet's 10BTC Challenge

Hacker News2
Ho
Holiday Binge Challenge – A WebGL Experiment46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Holiday Binge Challenge – A WebGL Experiment

Hacker News1
A
A Challenge to RK427%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A Challenge to RK4

Hacker News1
Rh
Rhythm Challenge35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Rhythm Challenge

Hacker News3
I
I Challenge You to a Regex50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I Challenge You to a Regex

Hacker News5
Mi
MindCipher – Challenge yourself35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

MindCipher – Challenge yourself

Hacker News14
Go
Go for Glory – take the apibunny challenge35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Go for Glory – take the apibunny challenge

Hacker News10
Te
TelTech Challenge35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

TelTech Challenge

Hacker News17
Habfun
Habfun33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Challenge yourself

Product Hunt+123Health & Fitness