LL

LLM Colosseum – A daily battle royale between frontier LLMs

Hacker News

LLM Colosseum – A daily battle royale between frontier LLMs

I put Claude, GPT, Gemini, and Grok in an arena and let them fight it out. Each model gets the full game state and decides how to survive - move, attack, form alliances, betray. Every decision comes from the model's API, nothing is scripted. First battle ran today. Gemini won by allying with GPT early, then backstabbing at the perfect moment. Claude tried to play it safe and got eliminated. They play very differently and it's fun to watch. Stack is React + Canvas, Bun + Hono on the backend. No database — battle data is JSON committed to git. Each model talks through its native SDK (Anthropic, OpenAI, Google, xAI). A new battle runs automatically every day. Source: https://github.com/sanifhimani/llm-colosseum

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: claude, model, google · Missing: mac, agents, macos
98%98% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: gemini · Missing: supports, reddit linkedin, podcasting
81%81% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: ide, io · Missing: https docs, excited, just released
54%54% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: google · Missing: mobile apps, ios, personal
46%46% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
32%32% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
22%22% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
3%3% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Ch
Chess-LLM, using constrained-generation to force LLMs to battle it out63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Chess-LLM, using constrained-generation to force LLMs to battle it out

Hacker News7
Al
Alien Battle – defeat your enemies41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Alien Battle – defeat your enemies

Hacker News22
Fo
Fortzone Battle Royale41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fortzone Battle Royale

Hacker News1
Ba
Battle Suck - TechCrunch Hackathon 201251%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Battle Suck - TechCrunch Hackathon 2012

Hacker News17
Ah
Ahoy.navy – A daily sea battle puzzle37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Ahoy.navy – A daily sea battle puzzle

Hacker News1
LL
LLM Templates – Streamline daily tasks with templated LLMs32%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Templates – Streamline daily tasks with templated LLMs

Hacker News3
Te
Text Battle – AI-simulated fights, daily leagues (Elo Ranking)44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Text Battle – AI-simulated fights, daily leagues (Elo Ranking)

Hacker News2
Be
BerriAI – Monitor Hallucinations in LLMs (Sentry for LLM Apps)66%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

BerriAI – Monitor Hallucinations in LLMs (Sentry for LLM Apps)

Hacker News4
GP
GPTCache – Redis for LLMs69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GPTCache – Redis for LLMs

Hacker News7
pr
prompttest – pytest for LLMs34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

prompttest – pytest for LLMs

Hacker News2