logoalt Hacker News

ASalazarMX • yesterday at 9:53 PM • 1 reply • view on HN

Funny, I'm running local LLMs in a modest iMac M4, they are slow indeed but not utterly slow. A dedicated system should be way faster I'd guess.


Replies

zamadatix • today at 12:38 AM

The M4 (soldered wide LPDDR5X + unified architecture + GPU with neural cores) is also closer in design to a dedicated GPU than an Epyc server with DDR5 or the components in the article.

But yes, a dedicated consumer GPU can still typically outpace the M4 quite well (if things fit in VRAM) and both will have their socks blown clean off by a hyperscaler GPU cluster. For some hard numbers I can run a ~24 GB model on my 5090 (about as fast as you can get on a single consumer class device) about 3x-4x faster than on my M4 mini and I'd still consider that pretty slow to running models 10x the size in the cloud.