logoalt Hacker News

jumploopstoday at 12:31 AM2 repliesview on HN

My biggest problem with running local LLMs on my M4 Max/128GB RAM is the prefill latency.

I've since acquired two DGX Sparks, and it feels so much snappier.


Replies

shell0xtoday at 1:40 AM

Would you mind sharing your local Mac setup and which models you currently use and whether it’s GGUF or MLX? I’ve the hardware same specs.

show 1 reply
c0rruptbytestoday at 12:45 AM

m5 max really fixed pp with the better matmul support, im sure the m5 ultra will be even crazier

the sparks have much slower memory bandwidth is the trade off

show 1 reply