logoalt Hacker News

formvoltronyesterday at 3:18 PM2 repliesview on HN

and get high token bandwidth?


Replies

numpad0yesterday at 7:25 PM

Not badly so because MoE models(identifiable by "CoolName-xxxB-AxxB" naming scheme) have bunch of branches in the middle that only one out of all gets non-zero values. Each of branches aka "Experts" as well as top/bottom parts are significantly smaller than the whole, and so CPU emulation of CUDA operations mixed with GPU taking as much as possible become not so out of question, unlike for dense models("CoolName-xxxB" without "-AxxB")

colordropsyesterday at 6:00 PM

Similar to a spark, which isn't blazing fast but usable.