Launching ChessArena – open-source Chess Benchmark to evaluate LLMs
Launching ChessArena – open-source Chess Benchmark to evaluate LLMs
A platform built to explore how large language models perform in chess games - OpenAI, Claude, Gemini. We created this platform using Motia to have a leaderboard of the best models in chess, but after researching and validating LLMs to play chess, we found that they can't really win games. This is because they don't have a good understanding of the game. In fact, the majority of the matches end in draws. So instead of tracking wins and losses, we focus on move quality and game insight. Each game is evaluated using Stockfish, the world's strongest open-source chess engine. How's it evaluated? On each move, we get what would be the best move using Stockfish to get the difference between the best move and the move made by the LLM, that's called move swing. If move swing is higher than 100 centipawns, we consider it a blunder. Is this project Open-Source? Yes! This platform is built using Motia Framework to showcase how simple it is to build a real-time application with Motia Streams. You can find the source code of this project on GitHub.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
I built an open-source benchmark that evaluates LLMs through gameplay
A benchmark classifer for an open source weed sprayer
Poozle – open-source Plaid for LLMs
Leaping – Open-source debugging with LLMs
An open source benchmark for prompt-injection detectors
Open Source toolset for launching NFTs collections on the El
An open-source ELO benchmark for voice agents
TensorZero – open-source data and learning flywheel for LLMs
Launching Flurly Affiliates
CentUp - Launching in late February