This implies you didn’t run any sort of evaluations? It is not realistic for any sort of production use case to do this.
Agree. Switching models with a keypress is for developer coding.
Jev and this decisions api are mostly useful for inference at scale in a workload where cost and latency matter… and that’s where evals become crucial. Could coding tools use it? Sure, but that’s probably a special case.
My gut is that the market for "ai as a tool" aka Claude Code/Codex/computer use/etc is a significantly bigger one than "models behind the scenes of some service". I've seen people use evals a lot in the latter case and very little in the former case (outside of people's whose job is basically to review the new releases). I haven't personally met anyone with something like "here are a bunch of tickets + a snapshot repo checkout, please try to solve them all" eval approach.
Though honestly I'm also surprised by the hype around Jev from a POV of "wait, are so many people just building on these by using them for classification tasks vs something more multi-step or generative?"