logoalt Hacker News

bigyabaiyesterday at 4:59 PM3 repliesview on HN

Not if it's 5-10x slower than a remote inference server. Mac prefill latency is exhausting.


Replies

spwa4today at 8:26 AM

Currently m5 max has a prefill rate for Qwen 3.8 27B of 400+ tok/s, then generate at ~60 tok/s. (And it can be improved further, the software is not yet at the level it is for CUDA)

That means for context under ~5k or so ttft (time to first token) it's going to respond faster than Claude. If the answer is less than ~1k I think the request finishes sooner. And it's ~claude 4.5 or 4.6 level intelligence.

I've had it work for more than a day on a pi.dev "loop engineering" project involving writing software.

Plus privacy. Plus offline.

jamiek88today at 3:48 AM

M5 changed that a lot though - it could still be better but 4x improvement made it cross the frustratingly slow barrier for me.

sanderjdyesterday at 7:05 PM

Oh tell me more about prefill latency.

show 1 reply