nikcub 1 week ago

Real test here would be using tinker to fine tune a tinker model to generate pelicans

calny 1 week ago

well it's good they didn't train on the test!

m3kw9 1 week ago

How is this still a valid test if there is a good chance they will now specifically train on it?

el_io 1 week ago

Gist is returning 403

dozerly 1 week ago

I’m afraid you’re going to have to start randomizing your benchmarks somehow. I’m sure these models are trained on this problem by now.

  • tyre 1 week ago

    If they did, they launched early! If they didn’t, their training data contains a bunch of poorly executed pelicans on bikes by other models.

  • tstrimple 1 week ago

    Wait. Based on the results of the test linked above you think this model might have been trained to produce it? Did you look at the results?!

    • argee 1 week ago

      That’s the problem. If all (or zero) models were trained on it, it would be fine as a benchmark.

  • 0-_-0 1 week ago

    It's not a benchmark, it's a meme

    • OrangeMusic 1 week ago

      To be fair, the "they're trained on this benchmark" response is also a meme.