logoalt Hacker News

Foobar8568 • yesterday at 9:02 PM • 1 reply • view on HN

Well, according to claude and Jevbench, Qwen 3.6 35b with ninfer on a RTX 5090@480W is like 3-5 time slower but 10%-15% better performance on the public set, I could see prefill > 15k for 700-800decode. Latency against what and which hardware? I don't really get jev...


Replies

alex7o • yesterday at 9:14 PM

Look I can convince my boss to pay for jev, but I won't convince him to run our prod stuff on a rented vast.ai 5090. And the pricing wouldn't be worth it. If you have ideas I would be glad to hear them

➕ show 1 reply