gdiamos 1 day ago

Data is doing more of the work than it used to. Every source in our mixture is a curated artifact built with large models

Training a model this small on them is distillation

When models of this size were last studied seriously such corpora did not exist

gdiamos 1 day ago

The loss does not saturate. Across a 4.91B-token run, smoothed training loss falls monotonically within each curriculum phase and is still descending at the end