Ne

New SWE-bench leaderboard compares LMs without fancy agent scaffolds

Hacker News

New SWE-bench leaderboard compares LMs without fancy agent scaffolds

Hello from the SWE-bench/SWE-agent team at Princeton/Stanford. When we created the SWE-bench benchmark in 2023 from hundreds of real-life GitHub issues/pull requests, the highest score was just a couple of percent. The tasks were so challenging for LMs, that most people didn't even want to work on them. Half a year later, SWE-agent showed that the early 2024 LMs were actually good enough to resolve up to 20% of the GitHub issues in the benchmark. This kicked off a whole wave of coding agents. Back then, developing agents was all about working around tons of silly behavior from the LMs. For example, if a command didn't work, they would try running the exact command again. If a command didn't return output, they would assume it never ran. They also couldn't get whitespace right in their edits, would get stuck into repetitive attempts and much much more. So agents got pretty complicated to work around all of that bad LM behavior. But now it's 2025, and LM companies have invested a whole lot of money to make their LMs really good at being agents. So we asked two questions: 1. What's the simplest agent we can write that still scores near SotA? 2. How do LMs compare when we evaluate them using this simple agent? Turns out, the agent can be very simple indeed! mini-swe-agent ( https://github.com/SWE-agent/mini-swe-agent ) has only 100 lines of code for the agent class (plus some 100 lines for environment etc.). It is little more than a loop that parses LM output for shell commands, executing them in a subshell, and continuing. We then took various LMs and put them to the test in a real apples-to-apples comparison without a fancy agent scaffold to prop up bad LMs. Our new leaderboard https://www.swebench.com/ shows the results. The highest score is currently 65% with Claude Sonnet 4 (which is not much less than the 70% that most fancier agents observe). o3, o4-mini, and Gemini 2.5 Pro are significantly behind, but not hopeless, achieving 50-60%. We were really surprised by these strong numbers overall: It shows that as LMs get stronger and better adapted at performing difficult, highly iterative tasks, we can take our hands off of the steering wheel, provide the minimal necessary environment, and let the LM figure out the rest. Let us know if you have any questions, our team is here on HN today :)

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, claude · Missing: mac, macos, cursor
94%94% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: created, gemini · Missing: supports, reddit linkedin, podcasting
79%79% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, io · Missing: https docs, excited, just released
64%64% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: plus · Missing: platform, intuitive, reviews
41%41% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
32%32% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
25%25% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Th
The Instavest Leaderboard40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Instavest Leaderboard

Hacker News6
FORLOGIS LMS
FORLOGIS LMS58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Learning Managment System

Indie Hackerscommitment-full-time
Tutor LMS 3.0
Tutor LMS 3.028%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

All-in-one WordPress LMS

Product Hunt+406WordPress
Manara LMS
Manara LMS26%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The All-in-One LMS for Students, Teachers & Centers.

Product Hunt+7
Cubite LMS
Cubite LMS67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Launch a branded AI LMS in minutes

Indie Hackers1$1,000/moai
ScoreLeader
ScoreLeader36%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Scoreboard & Leaderboard App

Product Hunt+6
Op
OpenCastor Agent Harness Evaluator Leaderboard47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OpenCastor Agent Harness Evaluator Leaderboard

Hacker News3
Tu
Turn any LMS into a native tablet app52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Turn any LMS into a native tablet app

Hacker News2
EzyCourse
EzyCourse73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LMS Platform, LMS Software, Community & Membership builder

Indie Hackerscommitment-full-time
Brusnika.LMS
Brusnika.LMS50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Corporate LMS that lives inside your CRM

Product Hunt+1