Ag

Agent Runner – open-source agent harness to benchmark real coding

Hacker News

Agent Runner – open-source agent harness to benchmark real coding

Hey HN! We built Agent Runner, a model-agnostic, open-source agent harness that executes the same prompt against two anonymized coding agents in parallel sandboxes. Each agent can make tool calls, edit multiple files, and self-correct through iterative reasoning. You pick the better result - this becomes the ground truth for the leaderboard. Why we built it Traditional benchmarks often fall short for modern agentic systems: they rely on static tasks and only measure final outputs. But real coding agents modify multiple files across a repo, answer to user re-prompts, use tool calls, and recover from partial failures What Agent Runner does You ask it to build anything Agent Runner kicks off two generations from different sandboxed LLM providers (OpenAI, Anthropic, Google, xAI, Mistral, Kimi, and more) Anonymized models make tool calls, multi-file edits, and cater to reprompts You pick your favorite - this preference powers the benchmark Because different providers handle tool calls, prompts, and execution semantics differently, we worked with each provider to ensure configurations reflect intended behavior. These provider-specific setups remain private, but Agent Runner itself is open-source. How to try it Kick off Agent Runner at https://www.designarena.ai/agentarena Repo at https://github.com/Design-Arena/agent-runner Use it as a CLI tool: https://pypi.org/project/agent-runner/ pip install agent-runner agentrunner run “create a nextjs replica of Discord” We hope this provides a provider-agnostic, framework-agnostic, realistic benchmark for state-of-the-art coding agents. Video demo: https://youtu.be/rdtiuCHatjs

Share card

Actual performance

4points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, model · Missing: mac, macos, cursor
99%99% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: para · Missing: supports, reddit linkedin, podcasting
74%74% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: video, google, para · Missing: mobile apps, ios, personal
46%46% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: calls · Missing: plus, platform, intuitive
45%45% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: ide, io · Missing: https docs, excited, just released
45%45% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
17%17% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

A
A benchmark classifer for an open source weed sprayer63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A benchmark classifer for an open source weed sprayer

Hacker News2
zo
zot – Yet another coding agent harness57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

zot – Yet another coding agent harness

Hacker News7
Zo
Zot – Yet another coding agent harness53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Zot – Yet another coding agent harness

Hacker News107
Emdash
Emdash94%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

One app. Every coding agent. Open-source.

Product Hunt+413Productivity
deepsec
deepsec83%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source coding security harness

Product Hunt+256Open Source
Op
Open-source web embeddable code runner79%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source web embeddable code runner

Hacker News31
An
An open source benchmark for prompt-injection detectors41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An open source benchmark for prompt-injection detectors

Hacker News2
HAR
HAR91%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open Source harness for multi-agent coding workflows

Product Hunt+112Open Source
CrabTalk
CrabTalk95%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The agent daemon that hides nothing. 5MB. Open Source

Product Hunt+192Developer Tools
Op
Open-source agent harness with team collaboration58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source agent harness with team collaboration

Hacker News1