AI

AI Debate Arena – See Which LLM Argues Best

Hacker News

AI Debate Arena – See Which LLM Argues Best

Ever wish you could get the best arguments for both sides of a debate? I built an AI-powered debate platform that pits language models against each other on controversial topics. Each AI is randomly assigned a side (pro/con). You vote before and after to see if you were persuaded. Most content today presents lopsided arguments. They provide strong points for one side, weak ones for the other. This project aims to surface the strongest arguments from both sides, using LLMs to simulate a fair debate. With enough usage, I want to use it to benchmark LLMs. My hypothesis is that randomly assigning sides of the debate, models with built-in biases will score worse. It’s currently using GPT 4o, Grok 3, and Gemini 2.5 Flash. It’s early, still rough around the edges, and I’d love feedback on the concept and direction. Curious how the HN crowd thinks this could evolve. It’s built for the intellectually curious that are open minded about changing their positions. Some next steps I’m considering: - Tuning the length and structure of arguments - Prompting improvements to reduce rhetorical fluff - Optional audio output of debates Try it out and let me know what you think!

Share card

Actual performance

5points
3comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, models, grok · Missing: mac, agents, macos
92%92% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: gemini · Missing: supports, reddit linkedin, podcasting
55%55% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform · Missing: plus, intuitive, reviews
46%46% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: ide, io · Missing: https docs, excited, just released
45%45% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
25%25% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
16%16% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: audio · Missing: web3, chat, crypto
2%2% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Re
Retrieval-augmented LLM debate opponent on DebateSum dataset66%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Retrieval-augmented LLM debate opponent on DebateSum dataset

Hacker News4
Ju
Justbate – Quora for debate63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Justbate – Quora for debate

Hacker News4
Op
Opscotch - Debate Anything63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Opscotch - Debate Anything

Hacker News2
I
I debate myself over hypothetical situations63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I debate myself over hypothetical situations

Hacker News1
LL
LLM Debate Benchmark56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Debate Benchmark

Hacker News9
Pe
Peer Arena – LLMs debate and vote on who survives69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Peer Arena – LLMs debate and vote on who survives

Hacker News5
De
DebateGate, a website for civil debate65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

DebateGate, a website for civil debate

Hacker News9
debate AI
debate AI42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

debate smarter with AI

Product Hunt+297Productivity
Fast header-only arena allocator
Fast header-only arena allocator32%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fast header-only arena allocator

Product Hunt+12
Co
ConvoClash – LLMs debate each other64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ConvoClash – LLMs debate each other

Hacker News6