logoalt Hacker News

Lwrlessyesterday at 11:22 PM1 replyview on HN

Thank you, found some benchmarks and they look really promising. To my understanding those Neural Accelerators are like AMX but for GPUs. With those accelerators the GPU performance on a M5 Max in LLM inferencing would totally be on par with a 5090, that's quite impressive!


Replies

bigyabaitoday at 3:20 AM

Which inference benchmark are you looking at? The prefill speeds should be comparable on some LLMs, but the 5090 has much a higher theoretical max decode speed.

show 1 reply