I

I built an open-source benchmark that evaluates LLMs through gameplay

Hacker News

I built an open-source benchmark that evaluates LLMs through gameplay

I built an open-source framework that evaluates LLM models using competitive gameplay. So far we have 3 games - a debate contest where LLMs try to persuade each other of different positions, a poetry slam where they judge each others' creativity, and a simple strategy game of cooperation and defection based on the prisoner's dilemma. The idea is that by pitting models against one another and evaluating their relative strengths we can scale the benchmark with model capability improvements. Some interesting results have emerged. DeepSeek R1 seems to be the most persuasive model - it's ranked #1 in debate slam and often sweeps the votes (as one example, in a debate against ChatGPT-4.5 it convinced all of the judges both for and against genetic engineering). DeepSeek R1 is also the current poetry slam champion, by quite a lot. Its poems are also often unanimous favorites. I'm not sure if this constitutes "creativity" per se or more like a different flavor of persuasion, but either way it seems impressive. I've read some of its poems and find them to be beautiful. Grok-2, meanwhile, is the current champion in prisoner's dilemma. It seems to be able to find the optimal time to defect in order to optimize its score (it is the first defector in 90% of its games). This is, to my knowledge, the only open-source benchmark of its kind. I think the open part is important, because it means the methodology and results are verifiable and reproducible. It also means (I hope) that others can jump in to contribute, either by adding new games, coming up with new ways to analyze and visualize the results, or by providing feedback. This has a lot of room to grow. I'm open to any and all critiques and feedback. And if you'd like to contribute please visit the project on github: https://github.com/jmogielnicki/llmshowdown Cheers, John

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, new, models · Missing: mac, agents, macos
95%95% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
86%86% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, io · Missing: https docs, excited, just released
50%50% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: visualize, way · Missing: mobile apps, ios, personal
35%35% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
35%35% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
13%13% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

A
A benchmark classifer for an open source weed sprayer63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A benchmark classifer for an open source weed sprayer

Hacker News2
Po
Poozle – open-source Plaid for LLMs65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Poozle – open-source Plaid for LLMs

Hacker News132
Le
Leaping – Open-source debugging with LLMs69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Leaping – Open-source debugging with LLMs

Hacker News15
An
An open source benchmark for prompt-injection detectors41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An open source benchmark for prompt-injection detectors

Hacker News2
Go
Golem – A beautiful open-source UI for LLMs, built on Nuxt 379%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Golem – A beautiful open-source UI for LLMs, built on Nuxt 3

Hacker News4
La
Launching ChessArena – open-source Chess Benchmark to evaluate LLMs48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Launching ChessArena – open-source Chess Benchmark to evaluate LLMs

Hacker News2
I
I built an open-source, privacy-first tool to map contexts for LLMs73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built an open-source, privacy-first tool to map contexts for LLMs

Hacker News1
An
An open-source ELO benchmark for voice agents52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An open-source ELO benchmark for voice agents

Hacker News8
Te
TensorZero – open-source data and learning flywheel for LLMs70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

TensorZero – open-source data and learning flywheel for LLMs

Hacker News49
I
I built an open source Favicon API67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built an open source Favicon API

Hacker News1