Ch

Chess-LLM, using constrained-generation to force LLMs to battle it out

Hacker News

Chess-LLM, using constrained-generation to force LLMs to battle it out

As I was playing with the Outlines library ( https://outlines-dev.github.io/outlines/ ), I discussed with my friend Maxime how funny it would be if we set up a way to pair LLMs in chess matches till one wins. The first time I tried it, it required substantial prompt engineering to get some of those LLMs to propose valid moves. Large language models can mostly stay focused and even play rather well; see https://news.ycombinator.com/item?id=37616170 for example. However small language models aren't as easy to convince. Some of those LLMs have seen very little chess notation and so after the first few opening moves there aren't any valid tactics, let alone strategy, so they would end up either repeating the same move, or hallucinate moves that are not valid (Kxe5, but there would be a queen on e5!) Then Outlines came along and we could force them to pick valid moves with little cost! Maxime worked super fast and got a first version of this idea as a gradio space. I think it is pretty fun to see the (mostly terrible, but otherwise valid) chess that those LLMs play. Maybe it will even be instructive to how we can create small LLMs that can play much better than the ones on the leaderboard. Anyway, you can check it out here: https://huggingface.co/spaces/mlabonne/chessllm What is interactive about it: you can pick the LLMs from available models on HuggingFace (within reason, small LLMs are preferable so that the space does not crash) or push one of your own small models to HF and have it fight with others. At the end of the game the leaderboard is updated. Hope you find it fun!

Share card

Actual performance

7points
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
89%89% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, new, models · Missing: mac, agents, macos
73%73% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: ide, io · Missing: https docs, excited, just released
64%64% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
45%45% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: way · Missing: mobile apps, ios, personal
42%42% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: active · Missing: arr, mrr, revenue
18%18% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

LL
LLM Colosseum – A daily battle royale between frontier LLMs54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Colosseum – A daily battle royale between frontier LLMs

Hacker News2
Re
Retrieval Augmented Generation Optimised LLM's67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Retrieval Augmented Generation Optimised LLM's

Hacker News2
Op
Opening Dojo, learn chess openings using spaced repetition52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Opening Dojo, learn chess openings using spaced repetition

Hacker News2
Gu
Guiding LLM outputs using Zod42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Guiding LLM outputs using Zod

Hacker News3
Al
Alien Battle – defeat your enemies41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Alien Battle – defeat your enemies

Hacker News22
Fo
Fortzone Battle Royale41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fortzone Battle Royale

Hacker News1
Ba
Battle Suck - TechCrunch Hackathon 201251%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Battle Suck - TechCrunch Hackathon 2012

Hacker News17
I
I made all LLMs play Chess against each other (inc. o1 models)39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made all LLMs play Chess against each other (inc. o1 models)

Hacker News2
Ch
Chess in Elm54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Chess in Elm

Hacker News6
Nu
Nuclear Chess, an explosive chess variant42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Nuclear Chess, an explosive chess variant

Hacker News2