Claude 3.5 Sonnet beats GPT-4o at Competitive Programming
Claude 3.5 Sonnet beats GPT-4o at Competitive Programming
I've designed a benchmark to evaluate the performance of different LLMs against high-quality representative competitive coding problems sourced from cses.fi. For now, I have benchmarked some of the SOTA models, including Claude 3.5 Sonnet and GPT-4o, and created initial visualizations and evaluations of the data. Suggestions are encouraged : ).
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
How to get started with Competitive Programming
Competitive Programming Made Easy
O3 beats Sonnet 4 at coding (in our codebase, wrt our preferences)
Access ChatGPT 4o and Claude 3.5 Sonnet Free Online
Competitive Programming Helper
Chat with multiple LLMs: o1-high-effort, Sonnet 3.5, GPT-4o, and more
Free Chat with GPT-4o
Arch-Function: 3B parameter LLM that beats GPT-4o on function calling
One API for GPT-5, Claude-Sonnet-4, DeepSeek, Gemini
An open-source Socratic coach for competitive programming