The M4 (soldered wide LPDDR5X + unified architecture + GPU with neural cores) is also closer in design to a dedicated GPU than an Epyc server with DDR5 or the components in the article.
But yes, a dedicated consumer GPU can still typically outpace the M4 quite well (if things fit in VRAM) and both will have their socks blown clean off by a hyperscaler GPU cluster. For some hard numbers I can run a ~24 GB model on my 5090 (about as fast as you can get on a single consumer class device) about 3x-4x faster than on my M4 mini and I'd still consider that pretty slow to running models 10x the size in the cloud.
The M4 (soldered wide LPDDR5X + unified architecture + GPU with neural cores) is also closer in design to a dedicated GPU than an Epyc server with DDR5 or the components in the article.
But yes, a dedicated consumer GPU can still typically outpace the M4 quite well (if things fit in VRAM) and both will have their socks blown clean off by a hyperscaler GPU cluster. For some hard numbers I can run a ~24 GB model on my 5090 (about as fast as you can get on a single consumer class device) about 3x-4x faster than on my M4 mini and I'd still consider that pretty slow to running models 10x the size in the cloud.