logoalt Hacker News

OtherShrezzingyesterday at 12:05 PM2 repliesview on HN

> A model you can run on a loptop is simply not going to work as well as it's needed for programming

The models you can run on a high-spec laptop today are approximately where frontier models were 12-18mo ago (albeit at a lower tok/s rate). If you scan back through hn comments from that era, you’ll find plenty of people saying “this is powerful enough to massively increase my productivity”.


Replies

anon373839yesterday at 1:05 PM

> albeit at a lower tok/s rate

Not always! I get 80-100 tok/s from Qwen 3.6 35B-A3B on a MacBook Pro thanks to MTP. With long contexts that dips to around 50-60. However, prefill is much slower than API models. So it becomes really, really, really critical to not have cache misses.

drivebyhootingyesterday at 5:10 PM

Slower is meaningfully dumber when you’re time bounded and need all the inference time compute you can get.