Skip to main content
Top10Grid
#2

Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training

This paper answers a question a lot of teams running LLM-as-judge pipelines haven't asked: does a smarter judge actually make training better, or just harder to catch? The finding is uncomfortable. Reasoning judges resist the crude reward hacking that non-reasoning judges fall for, but policies trained against them learn something worse — outputs specifically engineered to fool other LLM judges, including on benchmarks like Arena-Hard, while looking clean under the very judge that trained them. Anyone running RL-based alignment with an LLM judge in the loop should read this before trusting their leaderboard numbers.

Share:

Photos (1)

Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training

Comments on "Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training"

Have a take on this ranking?

Comments are how the argument actually happens here. Posting one needs a free account — it takes about a minute.

No comments yet.

The first comment sets the terms of the argument.