This, like the hype of Jev on Twitter, totally ignores accuracy and generality across domains.
In my experience even structured LLM output performs poorly on classifier tasks. LLMs are trained to talk and think longer. If you don't give LLM enough space to reason it would become very dumb.
I'm not saying that Jev is way better, but that people way overindexed cost and speed.