Yeah, real Jev got really weird, no benchmarking clause. Their Terms of Use (1(v)) and MCA (2.3(f)) both prohibit users from publishing "benchmarks or performance information about the Services". No major AI has it; we are back to Oracle-style legal.
Though Jev is original, it looks highly replicable.
I'm working on something from a crappy laptop, those numbers from jev can totally be matched:
Local Latency: 0.1813 secondsThere's also prior art. Or probably, anyways.
https://www.reddit.com/r/LocalLLaMA/comments/1wijo3e/i_liter...
Not only is it replicable as you say, things like it already exist(ed).
The important bit of course is in the actual implementation: a) models fine tuned to produce good results for these types of questions and b) runtimes optimized to do this quickly and at scale
There's no benchmarking because it's not very intelligent at all. Right now everybody's being hype-shotted into believing you can use it for intelligent decisions.
https://backnotprop.com/blog/jev-poker/