LL

LLMs suck at writing integration code… for now

Hacker News

LLMs suck at writing integration code… for now

Hi HN! Stefan here from superglue and today I’d like to share a new benchmark we’ve just open sourced: an Agent-API Benchmark, in which we test how well LLMs handle APIs. We gave LLMs API documentation and asked them to write code that makes actual API calls. Things like "create a Stripe customer" or "send a Slack message". We're not testing if they can use SDKs; we're testing if they can write raw HTTP requests (with proper auth, headers, body formatting) that actually work when executed against real API endpoints and can extract relevant information from that response. tl:dr: LLMs suck at writing code to use APIs. We ran 630 integration tests across 21 common APIs (Stripe, Slack, GitHub, etc.) using 6 different LLMs. Here are our key findings: - Best general LLM: 68% success rate. That's 1 in 3 API calls failing, which most would agree isn’t viable in production - Our integration layer scored a 91% success rate, showing us that just throwing bigger/better LLMs at the problem won't solve it. - Only 6 out of 21 APIs worked 100% of the time, every other API had failures. - Anthropic’s models are significantly better at building API integrations than other providers. Here is the results chart: https://superglue.ai/files/performance.png What made LLMs fail: - Lack of context (LLMs are just not great at understanding what API endpoints exist and what they do, even if you give them documentation which we did) - Multi-step workflows (chaining API calls) - Complex API design: APIs like Square, PostHog, Asana (Forcing project selection among other things trips llms over) We've open-sourced the benchmark so you can test any API and see where it ranks: https://github.com/superglue-ai/superglue/tree/main/packages... Check out the repo, consider giving it a star, or see the full ranking at https://superglue.ai/api-ranking/ . If you're building agents that need reliable API access, we'd love to hear your approach, or you can try our integration layer at superglue.ai. Next up: benchmarking MCP.

Share card

Actual performance

20points
13comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, model · Missing: mac, macos, cursor
97%97% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
93%93% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: exist, open source, ide · Missing: https docs, excited, just released
69%69% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: calls · Missing: plus, platform, intuitive
33%33% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
29%29% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
16%16% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Ah
Aha and HipChat Integration43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Aha and HipChat Integration

Hacker News2
Ah
Aha and Rally Integration43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Aha and Rally Integration

Hacker News7
Ph
PhotoSwipe Hugo Integration43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

PhotoSwipe Hugo Integration

Hacker News1
Ru
Runnaroo's Integration of SaaSHub43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Runnaroo's Integration of SaaSHub

Hacker News2
Re
Readwise – NotePlan Integration43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Readwise – NotePlan Integration

Hacker News2
Cl
Clockodo Integration for Emacs45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Clockodo Integration for Emacs

Hacker News2
Ba
Bazel.vim – Bazel Integration for Vim47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Bazel.vim – Bazel Integration for Vim

Hacker News1
Dj
Django-cypress: Django integration with Cypress41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Django-cypress: Django integration with Cypress

Hacker News1
Pa
Painless PayPal integration with Flask33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Painless PayPal integration with Flask

Hacker News1
Dr
Dropbox Integration For Educators49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Dropbox Integration For Educators

Hacker News1