points by simonw 1 week ago Here's a pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
nikcub 1 week ago Real test here would be using tinker to fine tune a tinker model to generate pelicans
m3kw9 1 week ago How is this still a valid test if there is a good chance they will now specifically train on it?
dozerly 1 week ago I’m afraid you’re going to have to start randomizing your benchmarks somehow. I’m sure these models are trained on this problem by now. tyre 1 week ago If they did, they launched early! If they didn’t, their training data contains a bunch of poorly executed pelicans on bikes by other models. tstrimple 1 week ago Wait. Based on the results of the test linked above you think this model might have been trained to produce it? Did you look at the results?! argee 1 week ago That’s the problem. If all (or zero) models were trained on it, it would be fine as a benchmark. 0-_-0 1 week ago It's not a benchmark, it's a meme OrangeMusic 1 week ago To be fair, the "they're trained on this benchmark" response is also a meme.
tyre 1 week ago If they did, they launched early! If they didn’t, their training data contains a bunch of poorly executed pelicans on bikes by other models.
tstrimple 1 week ago Wait. Based on the results of the test linked above you think this model might have been trained to produce it? Did you look at the results?! argee 1 week ago That’s the problem. If all (or zero) models were trained on it, it would be fine as a benchmark.
argee 1 week ago That’s the problem. If all (or zero) models were trained on it, it would be fine as a benchmark.
0-_-0 1 week ago It's not a benchmark, it's a meme OrangeMusic 1 week ago To be fair, the "they're trained on this benchmark" response is also a meme.
Real test here would be using tinker to fine tune a tinker model to generate pelicans
well it's good they didn't train on the test!
How is this still a valid test if there is a good chance they will now specifically train on it?
Gist is returning 403
I’m afraid you’re going to have to start randomizing your benchmarks somehow. I’m sure these models are trained on this problem by now.
If they did, they launched early! If they didn’t, their training data contains a bunch of poorly executed pelicans on bikes by other models.
Wait. Based on the results of the test linked above you think this model might have been trained to produce it? Did you look at the results?!
That’s the problem. If all (or zero) models were trained on it, it would be fine as a benchmark.
It's not a benchmark, it's a meme
To be fair, the "they're trained on this benchmark" response is also a meme.