logoalt Hacker News

barbacoayesterday at 1:38 PM2 repliesview on HN

They are saying that AMD's new Epyc Venice CPU has 16 memory channels allowing up to 1.6Tb/s of bandwidth. Which is higher bandwidth than most non-HBM GPUs.

So full CPU local AI inference may become viable option in coming years.


Replies

lallysinghyesterday at 3:01 PM

This is essentially guaranteed. There are lots of useful smaller models that we should be able to run locally. Over time they'll be more and more capable and require less API usage.

show 1 reply
adgjlsfhk1yesterday at 8:14 PM

the GPU competition is using 16 gpus, so the actual comparison is that the CPU has <1/10th the bandwidth