logoalt Hacker News

segmondy • yesterday at 9:37 PM • 2 replies • view on HN

A lot of people claim to have made better than jev, there's a jev benchmark, I have tried many of those models and they eventually end up failing, a non trivial task which doesn't seem like much but reminds me of the svg pelican bench is games, have one of these decision/classifier models play a game, hook it up to the input, most of the ones that are supposedly on jev level end up playing a terrible game, showing that they are very narrow. Cloudflare doesn't compare to the top open bench alternatives, I just finished downloading it and will compare it to jev for non trivial tasks tonight.


Replies

SebastianSosa • yesterday at 10:47 PM

Public benchmarks are easy to cheat, if I am typesafe I would also release a public benchmark to distract otherwise competent people in overfitting to a benchmark instead of making something actually useful. Diogo very much is against public benchmarks ;)

verdverm • yesterday at 10:49 PM

watching Jev play Pokemon demonstrated this too, more hype than meat

- I'd like a potion, are you sure, no, repeat

- in and out of doors on loop

- sisyphean effort in the cave

- jev-ish level grinding

It was impressive, beat pokemon for less than $2, but not all that interesting. People asking how different Math.Random plays pokemon would be, and at the other end, regular llms playing games.