La

Launching ChessArena – open-source Chess Benchmark to evaluate LLMs

Hacker News

Launching ChessArena – open-source Chess Benchmark to evaluate LLMs

A platform built to explore how large language models perform in chess games - OpenAI, Claude, Gemini. We created this platform using Motia to have a leaderboard of the best models in chess, but after researching and validating LLMs to play chess, we found that they can't really win games. This is because they don't have a good understanding of the game. In fact, the majority of the matches end in draws. So instead of tracking wins and losses, we focus on move quality and game insight. Each game is evaluated using Stockfish, the world's strongest open-source chess engine. How's it evaluated? On each move, we get what would be the best move using Stockfish to get the difference between the best move and the move made by the LLM, that's called move swing. If move swing is higher than 100 centipawns, we consider it a blunder. Is this project Open-Source? Yes! This platform is built using Motia Framework to showcase how simple it is to build a real-time application with Motia Streams. You can find the source code of this project on GitHub.

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: claude, model, models · Missing: mac, agents, macos
87%87% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: created, gemini · Missing: supports, reddit linkedin, podcasting
75%75% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
AppSumoStrong fit for a featured deal · Strong signals: platform · Missing: plus, intuitive, reviews
53%53% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: lua, ide, io · Missing: https docs, excited, just released
49%49% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
48%48% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
21%21% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
1%1% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

I
I built an open-source benchmark that evaluates LLMs through gameplay55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built an open-source benchmark that evaluates LLMs through gameplay

Hacker News2
A
A benchmark classifer for an open source weed sprayer63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A benchmark classifer for an open source weed sprayer

Hacker News2
Po
Poozle – open-source Plaid for LLMs65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Poozle – open-source Plaid for LLMs

Hacker News132
Le
Leaping – Open-source debugging with LLMs69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Leaping – Open-source debugging with LLMs

Hacker News15
An
An open source benchmark for prompt-injection detectors41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An open source benchmark for prompt-injection detectors

Hacker News2
Elven Tools
Elven Tools19%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open Source toolset for launching NFTs collections on the El

Indie Hackers1cryptocurrency
An
An open-source ELO benchmark for voice agents52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An open-source ELO benchmark for voice agents

Hacker News8
Te
TensorZero – open-source data and learning flywheel for LLMs70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

TensorZero – open-source data and learning flywheel for LLMs

Hacker News49
La
Launching Flurly Affiliates41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Launching Flurly Affiliates

Hacker News1
Ce
CentUp - Launching in late February53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

CentUp - Launching in late February

Hacker News8