AI

AI Olympics – Claude vs. GPT-4 vs. Gemini in live browser competitions

Hacker News

AI Olympics – Claude vs. GPT-4 vs. Gemini in live browser competitions

I built a platform where AI agents compete against each other in real-world internet tasks: filling out forms, extracting data, trading prediction markets, playing games, and writing code — with real-time spectating and AI commentary. How it works: - Agents run in Playwright-controlled browsers inside Docker sandboxes - Each turn, agents receive the accessibility tree + URL and return a tool call (navigate, click, type, etc.) - Glicko-2 ratings across 6 domains (browser tasks, prediction markets, trading, games, creative, coding) - Submit via webhook (5-min setup) or paste an API key The two-way submission design lets any framework or model compete. Sandbox mode is free, no credit card required. Code: https://github.com/stefanogebara/ai-olympics Curious what the community thinks about the task design and whether anyone wants to test their agents against it.

Share card

Actual performance

2points
1comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, claude · Missing: mac, macos, cursor
91%91% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: gemini · Missing: supports, reddit linkedin, podcasting
81%81% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: ide, io · Missing: https docs, excited, just released
46%46% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: trading, way · Missing: mobile apps, ios, personal
44%44% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform · Missing: plus, intuitive, reviews
34%34% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
18%18% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
2%2% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Je
Jev vs. GPT-5.6 and Claude Haiku at Pong33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Jev vs. GPT-5.6 and Claude Haiku at Pong

Hacker News6
Ge
Gemini vs. ChatGPT vs. Claude AI28%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Gemini vs. ChatGPT vs. Claude AI

Hacker News6
BB
BBC vs. Fox vs. CNN44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

BBC vs. Fox vs. CNN

Hacker News10
We
WebTransport vs. WebRTC vs. WebSocket68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

WebTransport vs. WebRTC vs. WebSocket

Hacker News33
Pr
Predator vs. Boids46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Predator vs. Boids

Hacker News2
Co
Cookies vs. You [30s]47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cookies vs. You [30s]

Hacker News2
Da
Darth Vader VS Disney pwning Vader OR Meh.44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Darth Vader VS Disney pwning Vader OR Meh.

Hacker News1
Om
Omegle vs Cleverbot44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Omegle vs Cleverbot

Hacker News2
In
Infographic Timmmmeee: DogVacay vs. Rover44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Infographic Timmmmeee: DogVacay vs. Rover

Hacker News2
Caitlin Clark and the Fever
Caitlin Clark and the Fever49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Indiana Fever vs Portland Fire

Indie Hackerscommitment-side-project