My gut is that the market for "ai as a tool" aka Claude Code/Codex/computer use/etc is a significantly bigger one than "models behind the scenes of some service". I've seen people use evals a lot in the latter case and very little in the former case (outside of people's whose job is basically to review the new releases). I haven't personally met anyone with something like "here are a bunch of tickets + a snapshot repo checkout, please try to solve them all" eval approach.
Though honestly I'm also surprised by the hype around Jev from a POV of "wait, are so many people just building on these by using them for classification tasks vs something more multi-step or generative?"