Now, can you do it in <200ms for 45 questions at once, have 0% malformed output, and any kind of meaningful benchmark? We’ll wait!
The principle is this.
Latency-calibration charts or it didn't happen. (LLMs are not optimized for calibration.)
I get dishonest vibes from this post? Jev claims to be cheaper/more efficient, and the post claims just to achieve the same functionality.
Nothing I hate more than bullshit articles claiming X in Y lines of code, only to use libraries abstracting hundreds of thousands of lines of code.
[flagged]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
I'm so sick of seeing these people who "made Jev in 25 lines of Python" or whatever the flavor of the day is. Do you people seriously think that Qwen3-0.6B-Q8_0.gguf is frontier intelligence? If you want to argue that Jev is NOT frontier intelligence, then go make that argument. Don't try to pretend that Qwen3-0.6B-Q8_0.gguf is frontier intelligence. That's retarded.
I'm surprised something like Jev came out "so late", but the hype has been ridiculous. Yes, it's a good idea. No, it only helps when fast and cheap are important and I guarantee existing labs will have this figured out in a matter of days.
Add visual understanding, add reasoning and bring down the size to run on my computer. That's when it will be interesting.
So many people that don't understand the tech jumped on the hype train because "it cannot hallucinate" and else. It's crazy.