Skip to main content
Top10Grid
#7

Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights

A provocative finding: at scale, randomly perturbing pretrained weights and ensembling the results is competitive with PPO and GRPO. This method achieves 95% of PPO's performance while requiring 80% less compute. The implication is striking — well-pretrained large models already contain abundant task-expert solutions densely packed around their weights. Optimisation is less about searching and more about choosing among solutions that already exist.

Share:

Photos (1)

Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights

Comments on "Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights"

Have a take on this ranking?

Comments are how the argument actually happens here. Posting one needs a free account — it takes about a minute.

No comments yet.

The first comment sets the terms of the argument.