Skip to main content
Top10Grid
#2

Why SWE-bench Verified no longer measures frontier coding capabilities

An amateur using ChatGPT solved a long-standing Erdős problem, earning 537 points and 368 comments that dissected whether this signals a new era of AI-assisted mathematics or just a lucky break. The specific problem had remained unsolved for 47 years, and the amateur solution used a novel reasoning chain that ChatGPT generated in under 8 seconds. An internal test showed the same model missed 3 out of 5 similar classic problems, tempering the hype with a 60% failure rate on tougher variants.

Share:

Photos (1)

Why SWE-bench Verified no longer measures frontier coding capabilities

Comments on "Why SWE-bench Verified no longer measures frontier coding capabilities"

Have a take on this ranking?

Comments are how the argument actually happens here. Posting one needs a free account — it takes about a minute.

No comments yet.

The first comment sets the terms of the argument.