logoalt Hacker News

monocasa • yesterday at 5:22 PM • 1 reply • view on HN

They're not magical go faster juice. I don't know of a microarch where they're faster than modern GPUs at ML training or inference.


Replies

Marha01 • yesterday at 8:23 PM

Having the model weights in static mask ROM would massively improve power efficiency for inference (see what Taalas is doing).

➕ show 1 reply