akshay_akula 17 minutes ago

Evals on actual research workflows is the right direction, most agent benches are toy tasks.

rubslopes 57 minutes ago

I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.

vatsachak 13 minutes ago

Damn. These things aren't AGI... but I don't care.

Luna is good enough for me to give a parser spec and have it write one.