logoalt Hacker News

johnfn • today at 12:06 AM • 2 replies • view on HN

This implies you didn’t run any sort of evaluations? It is not realistic for any sort of production use case to do this.


Replies

majormajor • today at 3:28 AM

My gut is that the market for "ai as a tool" aka Claude Code/Codex/computer use/etc is a significantly bigger one than "models behind the scenes of some service". I've seen people use evals a lot in the latter case and very little in the former case (outside of people's whose job is basically to review the new releases). I haven't personally met anyone with something like "here are a bunch of tickets + a snapshot repo checkout, please try to solve them all" eval approach.

Though honestly I'm also surprised by the hype around Jev from a POV of "wait, are so many people just building on these by using them for classification tasks vs something more multi-step or generative?"

shermantanktop • today at 12:45 AM

Agree. Switching models with a keypress is for developer coding.

Jev and this decisions api are mostly useful for inference at scale in a workload where cost and latency matter… and that’s where evals become crucial. Could coding tools use it? Sure, but that’s probably a special case.