htrp 19 minutes ago

> Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.

> Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training.

Early access, no weights no tech details, just a sign up here for info

wg0 9 minutes ago

Suppose I inherited a data center spanning several hundred acres full of GPUs and free electricity.

Where do I get the data?

I mean, this many models. They have to start somewhere.

  • Hamuko 8 minutes ago

    Get data from Claude. That's what the Chinese (allegedly) do.

  • altcognito 4 minutes ago

    If you ask a model, they will generally tell you where to get data. Modern frontier models have the large advantage of having tens if not hundreds of millions of users providing use cases to train against to improve their responses.