GP

GPT-Erdos – the results of GPT 5.2 Pro on the Erdos problems

Hacker News

GPT-Erdos – the results of GPT 5.2 Pro on the Erdos problems

Hi HN, it seemed like there was broad interest in the previous Erdos problem that GPT 5.2 Pro solved: https://news.ycombinator.com/item?id=46664631 I recruited a team of smart undergraduates to construct a dataset of ChatGPT responses to every open Erdos problem and verify the output. They found: - 3 problems with new proofs (though in 2 cases, historical partial results were found that could be extended to solve the same problem) - 4 problems where 5.2 Pro or Deep Research found an exact solution in the prior literature that hadn't been documented - 3 problems where 5.2 Pro or Deep Research were able to strengthen a prior result in the literature - 3 problems where typos were identified in the problem statement The most common failure case is that 5.2 Pro solves the problem as stated, but professional mathematicians understand there's an implicit constraint for the problem. For example, maybe the problem says integers, but they really mean only positive integers. Happy to answer any questions about the dataset!

Share card

Actual performance

1points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
65%65% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
46%46% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
44%44% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Product HuntUnlikely to reach the leaderboard · Strong signals: new, chatgpt, open · Missing: mac, agents, macos
37%37% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: ide, io · Missing: https docs, excited, just released
36%36% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
12%12% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat, smart · Missing: web3, crypto, cryptocurrency
1%1% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Wo
Worldbuilding Experiments with GPT-336%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Worldbuilding Experiments with GPT-3

Hacker News2
Au
Autosummarized HN (With GPT-3)51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Autosummarized HN (With GPT-3)

Hacker News6
I
I composed a sonata with GPT-3 DaVinci-003 and you can too51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I composed a sonata with GPT-3 DaVinci-003 and you can too

Hacker News3
GP
GPT Classifies HN Titles57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GPT Classifies HN Titles

Hacker News6
Vi
Visualized GPT57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Visualized GPT

Hacker News3
Ce
Cerebras-GPT-2.7B finetuned on Stanford Alpaca dataset65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cerebras-GPT-2.7B finetuned on Stanford Alpaca dataset

Hacker News4
Ho
Hostage Negotiation with GPT-446%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Hostage Negotiation with GPT-4

Hacker News2
Ro
Roleplaying GPT40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Roleplaying GPT

Hacker News3
Ik
Ikigai GPT46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Ikigai GPT

Hacker News1
7G
7GUIs in Hyphen by GPT46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

7GUIs in Hyphen by GPT

Hacker News1