#2
Why SWE-bench Verified no longer measures frontier coding capabilities
An amateur using ChatGPT solved a long-standing Erdős problem, earning 537 points and 368 comments that dissected whether this signals a new era of AI-assisted mathematics or just a lucky break. The specific problem had remained unsolved for 47 years, and the amateur solution used a novel reasoning chain that ChatGPT generated in under 8 seconds. An internal test showed the same model missed 3 out of 5 similar classic problems, tempering the hype with a 60% failure rate on tougher variants.
Photos (1)

Comments on "Why SWE-bench Verified no longer measures frontier coding capabilities"
Have a take on this ranking?
Comments are how the argument actually happens here. Posting one needs a free account — it takes about a minute.
No comments yet.
The first comment sets the terms of the argument.