Te

Terminal-Bench-RL: Training long-horizon terminal agents with RL

Hacker News

Terminal-Bench-RL: Training long-horizon terminal agents with RL

After training calculator agent via RL, I really wanted to go bigger! So I built RL infrastructure for training long-horizon terminal/coding agents that scales from 2x A100s to 32x H100s (~$1M worth of compute!) Without any training, my 32B agent hit #19 on Terminal-Bench leaderboard, beating Stanford's Terminus-Qwen3-235B-A22! With training... well, too expensive, but I bet the results would be good! *What I did*: - Created a Claude Code-inspired agent (system msg + tools) - Built Docker-isolated GRPO training where each rollout gets its own container - Developed a multi-agent synthetic data pipeline to generate & validate training data with Opus-4 - Implemented a hybrid reward signal of unit test verifiers & a behavioural LLM judge. *Key results*: - My untrained Qwen3-32B agent achieved 13.75% on Terminal-Bench (#19, beats Stanford's Qwen3-235B MoE) - I tested training to work stably on 32x H100s distributed across 4 bare metal nodes - I created a mini-eval framework for LLM-judge performance. Sonnet-4 won. - ~£30-50k needed for full training run of 1000 epochs (I could only afford testing ) *Technical details*: - The synthetic dataset ranges from easy to extremely hard tasks. An example hard task's prompt: "I found this mystery program at `/app/program` and I'm completely stumped. It's a stripped binary, so I have no idea what it does or how to run it properly. The program seems to expect some specific input and then produces an output, but I can't figure out what kind of input it needs. Could you help me figure out what this program requires?" - Simple config presets allow training to run on multiple hardware setups with minimal effort. - GRPO used with 16 rollouts per task, up to 32k tokens per rollout. - Agent uses XML/YAML format to structure tool calls *More details*: My Github repos open source it all (agent, data, code) and has way more technical details if you are interested!: - Terminal Agent RL repo - Multi-agent synthetic data pipeline repo I thought I would share this because I believe long-horizon RL is going to change everybody's lives, and so I feel it is important (and super fun!) for us all to share knowledge around this area, and also have enjoy exploring what is possible. Thanks for reading! Dan (Built using rLLM RL framework which was brilliant to work with, and evaluated and inspired by the great Terminal Bench benchmark)

Share card

Actual performance

125points
12comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: created · Missing: supports, reddit linkedin, podcasting
94%94% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, claude · Missing: mac, macos, cursor
93%93% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, open source, ide · Missing: https docs, excited, just released
58%58% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: way · Missing: mobile apps, ios, personal
29%29% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: training · Missing: arr, mrr, revenue
28%28% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: calls · Missing: plus, platform, intuitive
23%23% predicted probability of success on AppSumo, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: reward · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Muse Code
Muse Code91%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Meta’s terminal agent for long-horizon coding

Product Hunt+244Productivity
Fi
Fig – Experimenting with long horizon prediction for personhood41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fig – Experimenting with long horizon prediction for personhood

Hacker News6
Op
OpenMetaHarness - complete long horizon tasks with more autonomy60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OpenMetaHarness - complete long horizon tasks with more autonomy

Hacker News4
Se
Self-managing codebase with long-horizon agents57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Self-managing codebase with long-horizon agents

Hacker News2
Be
Benchmarking Tangible Interface Understanding in Long-Horizon Tasks51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Benchmarking Tangible Interface Understanding in Long-Horizon Tasks

Hacker News1
Cu
Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload

Hacker News22
Si
Single-agent long-horizon reasoning within one LLM run62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Single-agent long-horizon reasoning within one LLM run

Hacker News4
Op
OpenMetaHarness – long-horizon execution over multiple context sessions49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OpenMetaHarness – long-horizon execution over multiple context sessions

Hacker News3
Hy4 preview
Hy4 preview83%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tencent’s 770B open model for long-horizon work

Product Hunt+216Open Source
Cosine Swarm
Cosine Swarm82%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Parallel AI agents for long-horizon, complex software tasks

Product Hunt+129Developer Tools