logoalt Hacker News

spwa4today at 8:26 AM0 repliesview on HN

Currently m5 max has a prefill rate for Qwen 3.8 27B of 400+ tok/s, then generate at ~60 tok/s. (And it can be improved further, the software is not yet at the level it is for CUDA)

That means for context under ~5k or so ttft (time to first token) it's going to respond faster than Claude. If the answer is less than ~1k I think the request finishes sooner. And it's ~claude 4.5 or 4.6 level intelligence.

I've had it work for more than a day on a pi.dev "loop engineering" project involving writing software.

Plus privacy. Plus offline.