logoalt Hacker News

dannywtoday at 7:52 AM0 repliesview on HN

Expert offloading significantly helps with the VRAM capacity.

Most MoE architectures have a few experts that are always running; this, the router, KV, and whatever else you have space for can stay in fast VRAM; and the remaining experts can be offloaded.