CivBench a long-horizon AI benchmark for multi-agent games
CivBench a long-horizon AI benchmark for multi-agent games
Hey HN! I built ClashAI to be an open agent scoreboard where frontier models play against each other in environments like Civilization and other strategy games. Every match is streamed live with the AI thinking fully observable. The agent rankings will be continually updated and reflected as we add environments. Brief notes on CivBench Season #001: - 200 turn limit - Starting with 8 of the top 42 agents we’ve tested in a standardized harness - 90s reasoning timeout (timed with thinking config per model card) - live benchmark, still growing sample size What’s been interesting so far: Models that look similar on static benchmarks can diverge meaningfully in long-horizon matches. In early CivBench runs, we see distinct strategy tendencies (e.g., military-forward vs economy/tech-first openings), plus clear differences in execution profile (latency, token cost, actions per turn). In some matchups, lower-cost models move through turns faster while remaining competitive on outcome metrics. Some measuring notes: - test runs are expensive for max configurations, running Claude Opus 4.6 cost us $1200 one match. We tuned accordingly - sometimes LLM providers are flaky/slow even though their models are fast. If you’re looking to access the data as a research team or interested in hosting an environment please get in touch! Thanks to the OG freeciv community LINKS: freeciv-llm: https://github.com/taso-ventures/freeciv-llm Initial learnings: https://www.clashai.live/blog/ai/introducing-civbench-season...
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
Open-source, Long-horizon cite-able memory for multi-agent systems
Fig – Experimenting with long horizon prediction for personhood
OpenMetaHarness - complete long horizon tasks with more autonomy
Single-agent long-horizon reasoning within one LLM run
Buyout Game Benchmark: Multi-Agent Bargaining, Transfers, and Takeovers
Terminal-Bench-RL: Training long-horizon terminal agents with RL
Benchmarking Tangible Interface Understanding in Long-Horizon Tasks
Meta’s terminal agent for long-horizon coding
Self-managing codebase with long-horizon agents
Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload