Outrageously Small Neural Networks: Emergent Basic Reasoning at 6,616 tok/SEC [pdf] huggingface.co 6 points by Anon84 1 day ago
gdiamos 1 day ago Data is doing more of the work than it used to. Every source in our mixture is a curated artifact built with large modelsTraining a model this small on them is distillationWhen models of this size were last studied seriously such corpora did not exist
gdiamos 1 day ago blog: https://gregdiamos.com/2026/09/07/outrageously-small-neural-...X discussion: https://x.com/GregoryDiamos/status/2096873745420075020?s=20I added some of the main points to the thread so they are easier to read.
gdiamos 1 day ago The loss does not saturate. Across a 4.91B-token run, smoothed training loss falls monotonically within each curriculum phase and is still descending at the end
Data is doing more of the work than it used to. Every source in our mixture is a curated artifact built with large models
Training a model this small on them is distillation
When models of this size were last studied seriously such corpora did not exist
blog: https://gregdiamos.com/2026/09/07/outrageously-small-neural-...
X discussion: https://x.com/GregoryDiamos/status/2096873745420075020?s=20
I added some of the main points to the thread so they are easier to read.
The loss does not saturate. Across a 4.91B-token run, smoothed training loss falls monotonically within each curriculum phase and is still descending at the end