logoalt Hacker News

pwythonyesterday at 1:38 PM5 repliesview on HN

I was already rolling around the idea of a 128GB M5 Max MBP. Now this!

A 4-bit MLX quant with 128k window should fit perfectly, in the 50-70 tok/s range.


Replies

Eric_WVGGyesterday at 3:11 PM

Just out of curiosity, why run "local-local" when you could just set up a Mini or Studio at home and query it over http? [edit] whole conversation about this in another thread https://news.ycombinator.com/item?id=49433413

I’m personally considering retiring my MBP for a Studio + 15" Air whenever this MBP ages out.

show 2 replies
rdsubhasyesterday at 8:46 PM

How do you folks code at 40-50 tps? With an extremely lightweight harness (pi) and just 8k system and tools context, and ~40tps on qwen 3.8 27B 4-bit on low thinking mode, it still takes me nearly 30-45 mins for a basic coding session...

Does it work? yeah... But I'd pick a subscription anyday...

show 1 reply
sscaryterryyesterday at 1:41 PM

I have a 128GB M5 Max, and it sucks at this stage. 50-70 tok/s might be something...

show 1 reply
irthomasthomasyesterday at 2:01 PM

IDK, prefill speed is a bigger concern for most wokflows, like agent coding, and I heard that this is quite low on macs?

show 1 reply