logoalt Hacker News

mirekrusinyesterday at 7:30 PM1 replyview on HN

Speculative decoding also works on Mac, 64G is more than what I have, m5 max should handle up to ~40 t/s with optimized setup (and with a lot of vram you can get great wins on concurrency – that harness can take advantage of for single user task as well), but agree memory bandwidth in mac or spark is still too slow, next gen for both will be great hardware to have for sure.


Replies

smcleodyesterday at 8:38 PM

I get around 70tk/s on the m5 max, with 5bit AWQ / oQ5 slowing only to around 40tk/s at higher context.