GPT-Erdos – the results of GPT 5.2 Pro on the Erdos problems
GPT-Erdos – the results of GPT 5.2 Pro on the Erdos problems
Hi HN, it seemed like there was broad interest in the previous Erdos problem that GPT 5.2 Pro solved: https://news.ycombinator.com/item?id=46664631 I recruited a team of smart undergraduates to construct a dataset of ChatGPT responses to every open Erdos problem and verify the output. They found: - 3 problems with new proofs (though in 2 cases, historical partial results were found that could be extended to solve the same problem) - 4 problems where 5.2 Pro or Deep Research found an exact solution in the prior literature that hadn't been documented - 3 problems where 5.2 Pro or Deep Research were able to strengthen a prior result in the literature - 3 problems where typos were identified in the problem statement The most common failure case is that 5.2 Pro solves the problem as stated, but professional mathematicians understand there's an implicit constraint for the problem. For example, maybe the problem says integers, but they really mean only positive integers. Happy to answer any questions about the dataset!
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
Worldbuilding Experiments with GPT-3
Autosummarized HN (With GPT-3)
I composed a sonata with GPT-3 DaVinci-003 and you can too
GPT Classifies HN Titles
Visualized GPT
Cerebras-GPT-2.7B finetuned on Stanford Alpaca dataset
Hostage Negotiation with GPT-4
Roleplaying GPT
Ikigai GPT
7GUIs in Hyphen by GPT