Cl

Claude 3.5 Sonnet beats GPT-4o at Competitive Programming

Hacker News

Claude 3.5 Sonnet beats GPT-4o at Competitive Programming

I've designed a benchmark to evaluate the performance of different LLMs against high-quality representative competitive coding problems sourced from cses.fi. For now, I have benchmarked some of the SOTA models, including Claude 3.5 Sonnet and GPT-4o, and created initial visualizations and evaluations of the data. Suggestions are encouraged : ).

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: claude, model, models · Missing: mac, agents, macos
83%83% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: created, including · Missing: supports, reddit linkedin, podcasting
65%65% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
46%46% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: lua, io, including · Missing: https docs, excited, just released
45%45% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
41%41% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
14%14% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
5%5% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Ho
How to get started with Competitive Programming44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

How to get started with Competitive Programming

Hacker News1
Co
Competitive Programming Made Easy43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Competitive Programming Made Easy

Hacker News3
O3
O3 beats Sonnet 4 at coding (in our codebase, wrt our preferences)54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

O3 beats Sonnet 4 at coding (in our codebase, wrt our preferences)

Hacker News2
Chat100.ai
Chat100.ai43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Access ChatGPT 4o and Claude 3.5 Sonnet Free Online

Product Hunt+3
Co
Competitive Programming Helper28%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Competitive Programming Helper

Hacker News4
Ch
Chat with multiple LLMs: o1-high-effort, Sonnet 3.5, GPT-4o, and more60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Chat with multiple LLMs: o1-high-effort, Sonnet 3.5, GPT-4o, and more

Hacker News62
Fr
Free Chat with GPT-4o40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Free Chat with GPT-4o

Hacker News2
Ar
Arch-Function: 3B parameter LLM that beats GPT-4o on function calling44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Arch-Function: 3B parameter LLM that beats GPT-4o on function calling

Hacker News5
On
One API for GPT-5, Claude-Sonnet-4, DeepSeek, Gemini30%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

One API for GPT-5, Claude-Sonnet-4, DeepSeek, Gemini

Hacker News1
An
An open-source Socratic coach for competitive programming51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An open-source Socratic coach for competitive programming

Hacker News1