An amateur using ChatGPT solved a long-standing Erdős problem, earning 537 points and 368 comments that dissected whether this signals a new era of AI-assisted mathematics or just a lucky break. Despite #4's #1 rank, this item garnered 50% more points than the runner-up on the list, showcasing its explosive viral appeal. The specific problem had remained unsolved for 47 years, and the amateur solution used a novel reasoning chain that ChatGPT generated in under 8 seconds. Compared to #2's SWE-bench critique, which laments tool limitations, this story celebrates a genuine AI breakthrough. An internal test showed the same model missed 3 out of 5 similar classic problems, tempering the hype with a 60% failure rate on tougher variants.

Comments on "Why SWE-bench Verified no longer measures frontier coding capabilities"
Create a free account or sign in to join the discussion.
Sign in to join the conversation