why wouldn't they benchmark the accuracy against jev too?
Interestingly, I've seen better performance and similar cost to whats on JevBench
They say they are only benchmarking public models in the blog.
Also, section 2.3: https://typesafe.ai/legal/mca
Interestingly, I've seen better performance and similar cost to whats on JevBench