Skip to main content
Top10Grid
#5

NanoGPT Slowrun: 10x Data Efficiency with Infinite Compute

NanoGPT Slowrun achieves 10x data efficiency, enabling powerful models with significantly less data. By leveraging infinite compute for exhaustive hyperparameter search, it trains a 12-layer transformer to GPT-2 perplexity using only 10% of the original dataset—a 90% reduction.

Share:

Photos (1)

NanoGPT Slowrun: 10x Data Efficiency with Infinite Compute

Comments on "NanoGPT Slowrun: 10x Data Efficiency with Infinite Compute"

Have a take on this ranking?

Comments are how the argument actually happens here. Posting one needs a free account — it takes about a minute.

No comments yet.

The first comment sets the terms of the argument.