Ci

CivBench a long-horizon AI benchmark for multi-agent games

Hacker News

CivBench a long-horizon AI benchmark for multi-agent games

Hey HN! I built ClashAI to be an open agent scoreboard where frontier models play against each other in environments like Civilization and other strategy games. Every match is streamed live with the AI thinking fully observable. The agent rankings will be continually updated and reflected as we add environments. Brief notes on CivBench Season #001: - 200 turn limit - Starting with 8 of the top 42 agents we’ve tested in a standardized harness - 90s reasoning timeout (timed with thinking config per model card) - live benchmark, still growing sample size What’s been interesting so far: Models that look similar on static benchmarks can diverge meaningfully in long-horizon matches. In early CivBench runs, we see distinct strategy tendencies (e.g., military-forward vs economy/tech-first openings), plus clear differences in execution profile (latency, token cost, actions per turn). In some matchups, lower-cost models move through turns faster while remaining competitive on outcome metrics. Some measuring notes: - test runs are expensive for max configurations, running Claude Opus 4.6 cost us $1200 one match. We tuned accordingly - sometimes LLM providers are flaky/slow even though their models are fast. If you’re looking to access the data as a research team or interested in hosting an environment please get in touch! Thanks to the OG freeciv community LINKS: freeciv-llm: https://github.com/taso-ventures/freeciv-llm Initial learnings: https://www.clashai.live/blog/ai/introducing-civbench-season...

Share card

Actual performance

12points
24comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, claude · Missing: mac, macos, cursor
93%93% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
81%81% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
45%45% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: ide, io · Missing: https docs, excited, just released
45%45% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: plus, host · Missing: platform, intuitive, reviews
43%43% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
28%28% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Op
Open-source, Long-horizon cite-able memory for multi-agent systems69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source, Long-horizon cite-able memory for multi-agent systems

Hacker News4
Fi
Fig – Experimenting with long horizon prediction for personhood41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fig – Experimenting with long horizon prediction for personhood

Hacker News6
Op
OpenMetaHarness - complete long horizon tasks with more autonomy60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OpenMetaHarness - complete long horizon tasks with more autonomy

Hacker News4
Si
Single-agent long-horizon reasoning within one LLM run62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Single-agent long-horizon reasoning within one LLM run

Hacker News4
Bu
Buyout Game Benchmark: Multi-Agent Bargaining, Transfers, and Takeovers53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Buyout Game Benchmark: Multi-Agent Bargaining, Transfers, and Takeovers

Hacker News6
Te
Terminal-Bench-RL: Training long-horizon terminal agents with RL59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Terminal-Bench-RL: Training long-horizon terminal agents with RL

Hacker News125
Be
Benchmarking Tangible Interface Understanding in Long-Horizon Tasks51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Benchmarking Tangible Interface Understanding in Long-Horizon Tasks

Hacker News1
Muse Code
Muse Code91%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Meta’s terminal agent for long-horizon coding

Product Hunt+244Productivity
Se
Self-managing codebase with long-horizon agents57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Self-managing codebase with long-horizon agents

Hacker News2
Cu
Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload

Hacker News22