Learning to solve hard problems in RL for LLMs by never giving up mnoukhov.github.io 80 points by natolambert 12 hours ago https://arxiv.org/abs/2609.13443
aswegs8 13 minutes ago Seems like persistent models like OpenAI's highly persistent internal model can become really effective over time. Those are the ones that drove most of the HF-OAI incident.
Seems like persistent models like OpenAI's highly persistent internal model can become really effective over time. Those are the ones that drove most of the HF-OAI incident.