NanoGPT Slowrun achieves 10x data efficiency, enabling powerful models with significantly less data. By leveraging infinite compute for exhaustive hyperparameter search, it trains a 12-layer transformer to GPT-2 perplexity using only 10% of the original dataset—a 90% reduction. This performance is 3x better than the average data-efficient method and 40% more efficient than the runner-up on this list, #6's dithering approach in terms of per-point optimization.

Comments on "NanoGPT Slowrun: 10x Data Efficiency with Infinite Compute"
Create a free account or sign in to join the discussion.
Sign in to join the conversation